Explores limitations of conventional DEA models including bias, serial correlation, and bounded efficiency scores, with solutions like bootstrapping.
Key Takeaways
- Conventional DEA models tend to overestimate technical efficiency due to sample bias.
- Serial correlation in DEA complicates second-stage regression analysis of efficiency scores.
- DEA’s bounded efficiency scores require special statistical treatment to avoid misleading conclusions.
- Bootstrapping methods provide a practical solution to address bias and serial correlation in DEA.
- Stochastic frontier analysis inherently accounts for statistical properties unlike DEA.
What the video covers
- Introduction to parametric and non-parametric approaches for measuring technical efficiency, focusing on DEA.
- DEA constructs a frontier from sample data, which may not represent the true population frontier, causing bias.
- The estimated DEA frontier is always below the true frontier, leading to an upward bias in technical efficiency scores.
- Unlike stochastic frontier analysis, DEA lacks statistical properties to adjust for sample bias or confidence intervals.
- Serial correlation arises in DEA when technical efficiency scores are used in second-stage regressions with environmental variables.
- Serial correlation is due to peer observations influencing each other's efficiency scores, complicating inference.
- Technical efficiency scores are bounded between zero and one, with some observations always scoring one on the frontier.
- Bootstrapping and jackknife techniques are proposed to correct bias, serial correlation, and bounded score issues.
- Simar and Wilson algorithms are highlighted for correcting serial correlation and bias in DEA models.
- The session emphasizes caution in interpreting DEA results due to these inherent limitations.
Chapters
- 00:00Introduction and overview of efficiency measurement approaches
- 01:05Constructing DEA frontier and sample bias
- 01:44Bias in technical efficiency estimation
- 03:12Example illustrating DEA frontier bias
- 04:46Output-oriented technical efficiency and bias implications
- 05:56Effect of additional observations on DEA frontier
- 07:19Comparing DEA with stochastic frontier analysis
- 09:50Serial correlation issue in second-stage DEA analysis
- 12:02Bounded nature of DEA efficiency scores
- 12:34Bootstrapping and jackknife methods for bias correction
Full Transcript — Download SRT & Markdown
Speaker A
Hi. Welcome back to the course Applied Production Analysis Using MATLAB. So, so far we discussed two lines of literature.
Speaker A
One is the parametric approach and the other one is the non-parametric DEA or free disposal hull approach.
Speaker A
For measuring technical efficiency or productivity in the latter case. So, in this session, we will revisit our whole idea of data envelopment analysis and see what are the mainly criticized limitations of DEA as a non-parametric approach.
Speaker A
So, when it comes to DEA, what we are doing is we are taking data and considering that as the population and constructing a frontier based on that, and we are estimating the technical efficiency against that frontier.
Speaker A
So, unless you have the data for or access to the entire data, we are not going to have the true frontier.
Speaker A
So, that can create a kind of bias. This is very similar to the asymptotic bias you can conceptualize in a regression model, but here in this context it becomes even more complicated because we are talking about a technical efficiency score.
Speaker A
And building upon that, we have another issue that needs to be addressed, especially when you are going for an analysis that in the first stage you are estimating the technical efficiency and in the second stage you are making this
Speaker A
technical efficiency a function of Z variable, like the set of explanatory variables that we have already seen in the context of stochastic frontier analysis. In this session, we'll revisit our DEA framework we followed so far and
Speaker A
see how DEA models are prone to bias in the first instance. In the direction.
Speaker A
So, bias when I mentioned it is basically the bias in the estimation of efficiency score or technical efficiency scores.
Speaker A
So, we start with a very simple example. Say these are the—I'm putting more because I'm considering these as the population.
Speaker A
Consider the case output input. Suppose these are the population and out of that as a researcher you are able to get observations, say, okay.
Speaker A
See, this is somewhat half of the population. So, now following a DEA framework, we'll be constructing a frontier.
Speaker A
Basically, I'm constructing the VRS frontier. This is going to be my VRS frontier. Okay? But in reality or this frontier I call as F hat, which is basically the estimator counterpart of the actual frontier.
Speaker A
But in reality, when you have access to the entire data, which is not going to be the case for most of us, the true frontier is going to be something like this.
Speaker A
Same. After that is same. You're fortunate enough to have that observation. Okay. This is like a very simple example that we had. And here I'll call this as the actual frontier.
Speaker A
Okay. Now let's see what will happen when you don't have access to the original data, like entire data.
Speaker A
And when you have a framework of sample data, what will happen to the technical efficiency score?
Speaker A
So for the time being, we consider the technical efficiency output oriented. Say with this observation, we can take one simple case which is going to be the case for if you are having say X A Y A.
Speaker A
This is basically X A observation. You see, sorry, Y A. This is going to be Y A star under estimator frontier and this is going to be the Y A double star. So now with the simple logic we can understand that
Speaker A
Y A divided by Y A star is going to be greater than or equal to Y A divided by Y A double star. Same is going to be the case when you are approaching this problem from an input-oriented approach.
Speaker A
So, what is happening here? Since you are having access to only the data like sample data, the estimated frontier is going to be something below the true frontier.
Speaker A
And it's going to be always below the true frontier. As I mentioned, having a new observation never pulled the frontier down.
Speaker A
If not, it can just make a case that it can pull, push the frontier. So, by adding any observation over here, say I had this observation. It does not pull the—But, if I had an observation here, it
Speaker A
would have pulled my inner frontier up. That is the case, right? So, having an additional observation does not keep any frontier down. So, it will surely, having more observation, it always decreases the technical efficiency of other observation. So, we
Speaker A
can say that technical efficiency is a non-increasing function of number of observations that we are having so and so.
Speaker A
And there is a huge possibility that your estimated frontier is not close to the actual frontier. That is highly likely. Even if you miss out one observation, that gives us a room for doubting our estimated frontier and see what
Speaker A
whether it is the technical efficiency that we estimated is against actual frontier or not.
Speaker A
So, this gives us a summary that conventional DEA approach basically overestimates E or we can say upward bias.
Speaker A
As in upward bias when it comes to the technical efficiency estimation. That is the main criticism that we are going to encounter when you are using DEA as a tool.
Speaker A
But why not? It's not a problem for stochastic frontier analysis. Basically, as you know, stochastic frontier analysis is basically a frontier or MLE estimation-based approach, so it itself will take care of the property that the data is not the
Speaker A
actual data because by using the data we generally get a confidence interval and we always be skeptical about the fact that this data is not representative of our true population and based on the standard errors and so on. So, we
Speaker A
make some adjustment or make some caution while making an inference about the data. But unfortunately, DEA being a non-parametric approach does not have any statistical properties and it does not talk about the population or how much deviation is happening for the
Speaker A
individual observation from the potential or true frontier that we are unable to observe.
Speaker A
We revisit this concept how to account for that so and so. Now, I quickly take you through another issue that we are going to encounter. Basically, that is the serial correlation issue.
Speaker A
Serial correlation is not going to be an issue when you are estimating technical efficiency for the purpose of ranking or just for getting an overall trend of technical efficiency so and so.
Speaker A
It becomes an issue the moment you are going for a second stage regression where you are trying to conceptualize technical efficiency as a function of some environmental variable. How is it happening?
Speaker A
So here you can see efficiency of this observation is estimated again this particular observation which is coming as the peer.
Speaker A
Same for this observation, efficiency of this observation is estimated again these two observations, which is coming as its own peer. So these peers will face some, if these peers face some external condition which is out of the control of the model,
Speaker A
they may have more output or having them more output will keep a case where this observation may get a lower efficiency score.
Speaker A
Them having a lower output may give a case that this observation is going to have a higher efficiency score so and so. So that means this is more of a relative performance evaluation framework that we are following. So that relative nature
Speaker A
comes with a complication that the error term that you are going to include in this model these are likely to be correlated with each other which we are going to call as the serial correlation or it is very similar to
Speaker A
what we claim autocorrelation in a time series framework or in even some literature or textbook call serial correlation interchangeably with the autocorrelation. Okay.
Speaker A
Another issue that we are going to face, another subset of these two issues that we are going to face in the context of DEA. So, this is basically the first issue, the second issue. And another third issue I would say this is
Speaker A
basically the bounded nature of technical efficiency score. The moment you are estimating technical efficiency score using conventional DEA, surely some of them are going to be on the frontier. So, they will get a technical efficiency score one.
Speaker A
But these observations are coming as technical efficiency score one.
Speaker A
If they were in the population, they wouldn't have performed that well. So, these are basically spuriously efficient also.
Speaker A
Okay? So, this spuriously efficient issue will also create a case where you have a lot of observation with value one. The moment you go for the second stage analysis that I was referring, it will create a problem.
Speaker A
So, the bootstrapping is going to be a solution for uh these two issues, mainly these two issues that I'm referring to, plus the serial plus the bounded nature.
Speaker A
So, and so. So, in this case, how we are going to proceed? This is going to be very technical, at least for few of you as compared to the earlier module.
Speaker A
So, how are we going to proceed? Initially, I'm going to introduce you to the very basic bootstrap algorithm that is basically a bootstrapping that built upon simply some uh sampling techniques.
Speaker A
Something very familiar to you like jackknife technique. So, initially in the first instance, just to understand the philosophy of bootstrapping, we'll try the jackknife or we uh see what is the jackknife approach.
Speaker A
Then later, if you proceed, uh basically the bootstrapping as a method it was proposed by Tibshirani, Effron and Tibshirani. So, we are not going to the very technical details of the bootstrapping as an approach.
Speaker A
We'll be using mostly the Simar and Wilson algorithm, which is basically the 2000 Simar Wilson 2007 algorithm, which is most most like the latest full-fledged development, I would say.
Speaker A
This algorithm is based on their earlier works, which starts from 1998, then two works on 2000, and then later one on 2007.
Speaker A
Uh so, as I mentioned, we'll start with the very basic fundamental philosophy of bootstrapping.
Speaker A
And then we see how to apply simple bootstrapping in the context of DEA. The very foundational bootstrapping may not work in the context of DEA because it does not follow the data generating process that we expect the algorithm to follow.
Speaker A
So, the Simar and Wilson is basically a case that we are going to discuss in detail, which has two algorithms. First algorithm will correct only for the serial correlation, which is seemingly uh less complicated. And the second algorithm will correct for both bias in
Speaker A
the estimation and the serial correlation issue in the context of uh DEA score that we'll see in detail.
Speaker A
Thank you.
Topics:Data Envelopment AnalysisDEA limitationstechnical efficiencybias in DEAserial correlationbootstrapping DEAstochastic frontier analysisefficiency scoreSimar and Wilsonproduction analysis











