Overview of Simar–Wilson algorithms for bias correction and serial correlation in DEA efficiency estimation with statistical foundations.
Key Takeaways
- Jackknife procedure is insufficient for bias correction in DEA as it ignores the data generating process.
- Simar and Wilson algorithms provide a statistically grounded method to correct bias and serial correlation in DEA efficiency scores.
- The IID assumption is critical for applying these algorithms, limiting their use to cross-sectional data.
- Algorithm 2 of Simar and Wilson first corrects bias, then serial correlation, offering a comprehensive correction approach.
- Smoothness and continuity of the production function are essential assumptions for the validity of these methods.
What the video covers
- The video discusses limitations of the jackknife procedure in bias correction for DEA efficiency scores, emphasizing its failure to account for the data generating process.
- It introduces the Simar and Wilson algorithm (2007) as a method that incorporates the original data generating process in efficiency estimation.
- Two main issues in conventional DEA are highlighted: bias in efficiency estimation and serial correlation.
- Simar and Wilson propose two algorithms: Algorithm 1 corrects only serial correlation, while Algorithm 2 corrects bias first and then serial correlation.
- Key assumptions underlying these algorithms include the independent and identically distributed (IID) nature of inputs, outputs, and efficiency determinants.
- The production possibility set is defined with a smooth and continuous production function, particularly in an output-oriented DEA context.
- The video explains the importance of the Farrell distance function (delta) as a measure of efficiency.
- Exogeneity and separability conditions are discussed as foundational statistical assumptions for the bootstrap approach in DEA.
- The residuals in the model are assumed to be normally distributed with zero mean, with truncation ensuring efficiency scores are greater than or equal to one.
- The video clarifies that the Simar and Wilson algorithms are not directly applicable to panel or time series data due to the IID assumption.
Chapters
- 00:00Introduction and critique of jackknife procedure in DEA
- 01:53Limitations of jackknife and data generating process
- 03:45Introduction to Simar and Wilson algorithm
- 06:00Overview of Simar and Wilson's two algorithms
- 07:22Assumptions underlying Simar and Wilson algorithms
- 08:53Statistical foundations: IID assumption and joint distribution
- 10:44Smooth and continuous production function assumption
- 12:28Exogeneity, separability, and residual distribution assumptions
Full Transcript — Download SRT & Markdown
Speaker A
Hi, welcome back to the course Applied Production Analysis using MATLAB. In the last session, we had a very simple example of the jackknife procedure and how to do the bias correction involved in the estimation of DEA efficiency score. As I mentioned, if you see the latest literature, none of them are using jackknife as a tool for correcting the bias. The main criticism of jackknife as a tool for correcting bias in the estimation of efficiency score under DEA is basically it does not account for the data generating process.
Speaker A
mentioned if you see the latest literature none of them are using jack knife as a tool for correcting the bias. The main criticism jack knife as a tool for correcting bias in the estimation of efficiency score under DEA is basically it does not
Speaker A
What it worked on, we have the population and from that population we have taken a sample. Say this is the population. This is the sample that we are having. Within that, they create subsamples and then this gives you a—this is basically ideally something that you get, a delta star. Basically, delta we use this as a, in place of theta in this context, and basically this delta is basically the Farrell distance function or one by delta will give you the Shephard distance function or efficiency score.
Speaker A
subsamples and then this get this give you a this is basically ideally something that you get a delta stars. Basically delta we use this as a uh in place of fee in this context and basically this delta is basically
Speaker A
So my point was, if you have the population, we get delta, but since we have only a sample, we get delta hat. But to see the difference between this delta and delta star, what we are doing is we are keeping delta star as the delta, and from each sample you get delta double hat. And then you see how much deviation is happening between delta double hat and delta hat to get an approximation of the bias or deviation between delta hat and delta.
Speaker A
So my point was if you have the population we get delta but since we have only sample we get delta hat but to see the difference between uh this delta and delta star what we are doing we are keeping delta
Speaker A
So that was like the philosophy remains the same even for our jackknife procedure. But the problem is jackknife just worked on the sampling procedure, not on the procedure involved in determination of efficiency score or we call it as the data generating process.
Speaker A
So uh that was like the philosophy remains the same even for our jack knife procedure.
Speaker A
So data generating process is basically, from a very layman's perspective, it is basically the way the parameters are being determined. So here the parameter that we are referring to is basically the delta or the efficiency scores, and the approach is basically the efficiency scores are determined as a relationship between outputs and inputs, right? How inputs are converted into output, how well, but that relationship is basically determined by another set of variables, say the size of the firm or the R&D and so on.
Speaker A
So data generating process is basically like from a very layman's perspective it is basically the way the parameters are being uh determined. So here the parameter that we are referring it is basically the delta or the uh efficiency scores and
Speaker A
Having said that, now the algorithm that we are going to discuss is the Simar and Wilson algorithm. So it is basically the 2007 development. It's an algorithm that claims that, or first of its kind attempt to follow the original data generating process involved in the determination of efficiencies post.
Speaker A
the R& R&D and so and so. Having said that, now the algorithm that we are going to discuss that is Simar and Wilson algorithm. So it is basically the 2007 development.
Speaker A
So in this session, we see the statistical foundations of Simar and Wilson algorithm. And when it comes to Simar and Wilson algorithm, as I mentioned, there are two issues in the context of DEA or conventional DEA.
Speaker A
and Wilson algorithm as I mentioned there are two issues uh in the context of DEA or conventional DA.
Speaker A
One is basically two issues. One is basically the bias in the estimation of efficiency score and the other one is basically the serial correlation.
Speaker A
Okay. Serial correlation. These are the two issue that we saw in the context of conventional DA efficiency estimation.
Speaker A
Okay. Serial correlation. These are the two issues that we saw in the context of conventional DEA efficiency estimation.
Speaker A
I hope you are getting the context. So there are two algorithms that account for the bias in the efficiency of estimation the serial correlation and uh here correction of serial correlation is bit more straightforward. So algorithm one does only the serial correlation
Speaker A
As I mentioned, this Simar and Wilson algorithm has two different algorithms. Basically, we have algorithm one which corrects only for the serial correlation, and algorithm two basically corrects for bias then it corrects for the serial correlation.
Speaker A
Okay. And algorithm two has two steps. The first step is correct the bias and the second step it corrects the serial correlation.
Speaker A
I hope you are getting the context. So there are two algorithms that account for the bias in the efficiency estimation and the serial correlation, and here correction of serial correlation is a bit more straightforward. So algorithm one does only the serial correlation correction, and serial correlation correction comes into the picture when you have a second stage regression.
Speaker A
these two algorithm. So the first assumption is basically the IID our input output and the determinance of efficiency scores are expected to be independently and identically distributed.
Speaker A
Okay. And algorithm two has two steps. The first step is correct the bias and the second step it corrects the serial correlation.
Speaker A
even in the conventional DA that is going to be very complicated we don't have any time series framework or framework that uses conventional DE or whatever I'm talking about panel data if it was a panel data each observation its
Speaker A
So before going into the details of these two algorithms, I would like to take you through the assumptions involved in the, or assumptions or the statistical foundation that Simar and Wilson has kept in the back end for implementing these two algorithms. So the first assumption is basically the IID. Our input, output, and the determinants of efficiency scores are expected to be independently and identically distributed.
Speaker A
Okay. So that is the independent uh property and identically distributed. We are expecting all of them to come from the same pool of population. So both all of them will have all observation n that we are considering will have same uh
Speaker A
So the independent property basically, from an applied point of view, the first condition for independence I would say it should be cross-section data. If it is time series data or so, no, time series is not going to be a case even in the conventional DEA. That is going to be very complicated. We don't have any time series framework or framework that uses conventional DEA or whatever. I'm talking about panel data. If it was panel data, each observation, its own lags are coming or its own earlier period observations are coming. So that will create an issue. So we cannot use the Simar and Wilson algorithm directly to an observation which has multiple time periods.
Speaker A
So here Sn is basically the realization of identically and independently distributed random variables. And here the joint distribution can be defined of f of xy z. And in this context we have something called p. Basically it is a u
Speaker A
Okay. So that is the independent property and identically distributed. We are expecting all of them to come from the same pool of population. So all of the observations n that we are considering will have the same parental distribution, like each observation is coming from a pool of the same population. So each of them will have the same probability density function that gives them the identical characteristics.
Speaker A
is basically the uh frontier or the production function that we are going to consider.
Speaker A
So here Sn is basically the realization of identically and independently distributed random variables. And here the joint distribution can be defined as f of x, y, z. And in this context, we have something called p. Basically, it is a real number and here it takes the dimension of p + q, and p is basically the number of inputs and q is basically the number of outputs. And then p is basically defined as the production possibility set and the boundary of that is basically the frontier or the production function that we are going to consider.
Speaker A
determined by the uh or the efficiency of the firm is being determined by the zed variable. And there is a very important condition that f of xy given zed should be different from symbol f of x. Right? the or the
Speaker A
Now, over and above the input variable, we are plugging in the environmental variable. We call it as the z variable, and here z are the determinants of efficiency. So given inputs, how much the firm could produce is basically determined by the, or the efficiency of the firm is being determined by the z variable. And there is a very important condition that f of x, y given z should be different from symbol f of x. Right? The conditional probability of our input-output values should be conditional upon the environmental variable and should be different from their independent distribution. Right?
Speaker A
And this gives a fact that if f of x y = to f of x and y given value zed, this means zed has nothing to do with the values of x and y. That gives us a hint
Speaker A
And this gives a fact that if f of x, y equals to f of x and y given value z, this means z has nothing to do with the values of x and y. That gives us a hint that these variables that you included are not being the determinant of efficiency score. Having said that, there is no second stage estimation that we can or second stage regression model that we can do meaningfully in this context.
Speaker A
The second assumption that we are having smooth and continuous production function. We expect our production function that is basically boundary of pro boundary of our uh P set that is production possibility set and that has to be continuous and smooth in the sense
Speaker A
The second assumption that we are having is a smooth and continuous production function. We expect our production function, that is basically the boundary of our production possibility set, and that has to be continuous and smooth in the sense that this, say for delta i for the given values of z, the function s that we are defining should be a smooth and continuous function. There should not be any discontinuity. And here this function is something we are going to keep on using. We are defining delta i as a function of this z i, and basically this is our matrix vector. Z i is basically a matrix plus epsilon i, which is going to be greater than or equal to 1 in the context of output-oriented approach. Here what we are considering is basically the output-oriented case, and delta as I mentioned is basically the Farrell distance function.
Speaker A
use. We are defining delta i as a function of this uh z i and basically this is our matrix vector. Z i is basically a matrix plus epsilon i which is going to be greater than equal to 1
Speaker A
Okay, here the function that implies the relationship between our environmental variables and efficiency should be smooth and continuous where we can specifically mention that for that to hap—
Speaker A
Okay, here the function that implies the relationship between our uh environmental variables and efficiency should be smooth and continuous where uh we can specifically mention that for that to happen the epsilon I and that should also follow a
Speaker A
continuous iid uh random value not necessarily iid that smoothness can happen even if it is not independently distributed. Okay. And we assume one more assumption here that comes uh that is it's like an extension of our assumption to epsilon i and z i should
Speaker A
be independent. So that basically hints us the fact that there should not be any endogenity case in the context of explanatory variable. So that will if there if x epsilon i and z is a correlator that will violate the uh
Speaker A
exogenity assumption. So here the assumption one and two that we are mentioning it implies the something called separability condition. So here this is something that you may come across in the context of later developments in the field of bootstrap
Speaker A
DEA model. from a very direct sense separability condition in this context mean we have X Y and Z sometime some variable say R&D expenditure sometime in some sense R&D expenditure may become an input right it can happen so then can we use
Speaker A
R&D expenditure as a determinant of efficiency so this This is same the same framework what we had it in the kohi uh 1995 model. So what they did they included determinance of effic inefficiency also within the model. Uh
Speaker A
there were not much complication but in this context if zeds are not separable from the other inputs or if there is no very restricted or very tight classification of inputs and explanatory variable it becomes very complicated. So separability condition says that or
Speaker A
separability condition ensures that the variable that you are included in the model as explanatory variable zed over here are not a variable that satisfy the condition for being included in the model as an input. Okay.
Speaker A
And here the P that we are uh specifying it is basically a subset of entire sample space sample space and the coariate Z affect the production through the dependency between Y and E is Z. So if you have given input how much output
Speaker A
can be uh produced will be determined or the value Y in this context will be determined by the Z variable. Moving ahead uh we expect residual in the model. This is basically this to follow some set of assumption. The very first assumption
Speaker A
that we are keeping it should follow a normal distribution and uh with zero mean and sigma epsilon square as the variance. Here this has to follow or it has a left trangation at 1 minus s uh z i bit. This is slightly uh
Speaker A
complicated for at least for few of you to comprehend what will happen. This has to be greater than or equal to zero always. Right? So this is what something the deterministic component that we are having. Now we are talking
Speaker A
about the this is the deterministic component and now we are talking about the epsilon I and since it is feral efficiency distance function it should be always greater than or equal to 1 and if this epsilon is something coming from any
Speaker A
range between infinity to minus infinity to plus infinity what will happen the moment you plug in those values this value for the given this values it may violate the condition greater than or equal to 1.
Speaker A
Say if epsilon is not less than 1 minus s z i beta this will become less than one right so that condition is very essential in the context of bootstrap da uh simar wson algorithm that we are referring and also
Speaker A
uh we assume as I mentioned epsilon to be normally distributed with zero mean and the truncation ensures that the par distance function that you're estimating are greater than or equal to one and not necessarily it has to be normally
Speaker A
distributed. You can conceptualize other distributions also. But the model that we are going to discuss especially in the SimR and Wilson 2007 algorithm it is basically a assumption that model that based on normality assumption and they are claiming that the same model can be
Speaker A
extended to semiparametric uh alternatives is existing in the literature. Moving ahead uh we have convexity and boundedness of the production technology.
Speaker A
So this is very familiar to you. This is something we have already defined. uh so our P is basically a closed and convex uh set and then boundary of that is basically the production function as I mentioned it's a uh it's a um space or a
Speaker A
it has a domain of only the real numbers positive real number okay and both x and y will take only positive values and uh here we are claiming that the input requirement said that we are having basically the convex and the output
Speaker A
corresp respon yx is basically boundary. So this satisfy basically the something very similar to what and conceptualize in the context of one input one output.
Speaker A
So it is basically a closed bounded and convex set. Bounded from above not from the input side bounded from above closed and convex set. Okay.
Speaker A
So now uh another assumption we have already seen there is no free production. uh this is something an extension of our um our what I would say uh predisposibility.
Speaker A
If you keep on uh reducing your input that means if you read the point zero you will lose the output also. Here uh this ensure that if input values are zero then output will also be uh zero and if
Speaker A
there is a case that u output not equal to zero then x and y are not basically uh considered as a set of part of the production possibility set. So all production requires some inputs and no production without input
Speaker A
is possible and uh we have another assumption called uh like this is basically the there is no free length.
Speaker A
Uh if you are keeping input zero your output is also going to be zero.
Speaker A
Moving ahead we have strong disposability of inputs and output. This is same predisposivity of input and predisposability of output. We are particularly keeping it as strong such a way that uh this model does not take into account undesirable output
Speaker A
uh which has a weak disposability at some case that we have already discussed. And for here suppose you have x tilda as a value which is greater than uh x tilda is greater than x and y tilda less than y. So if x and y are part of
Speaker A
the uh production possibilities x tilda y and x y tilda should also be part of our uh production possibilities. So this is basically both predisposibility of input and predisposability output kept into one condition. Okay.
Speaker A
So having more uh input remain PC because this is not bounded from the right side. The technology set is bounded from the above. Even uh you can go to the toward the right side and get get the same level of out output what we
Speaker A
are getting using disposing more input and uh now moving ahead our frontier should be continuous and positive. Okay, for that the production possibility set that we are considering P say here if I draw this manner so say we have one point here say X and
Speaker A
Y and if I take a uh if I take a case where theta greater than 1 that means 1 by theta become so this basically 1x theta x y this it will not belong to p if theta is
Speaker A
greater than one that means it will be a point uh toward the this side okay it is not going to be feasible and uh we expect our f of x y z to be strictly positive the same condition applies in the context of
Speaker A
output also uh if theta is greater than one. So it is going to be somewhere above right. Uh so here we are following a theta. This is basically not delta.
Speaker A
This is basically theta. So we cannot have a point over uh above above the technology frontier. This is basically the it ensures the continuous uh contin continuous function or the production positivity basically um or from their density basically continuous
Speaker A
and moving ahead we have a strict positive uh assumption that our f of x y given the zed should be continuous toward the interior of our production possibility set everywhere you should be able to define a value of uh x and y
Speaker A
there should not be any discontinuity and it is positive there should be a positive probability near the frontier and this ensures the consistency of our estimator.
Speaker A
Another assumption that we are having the final assumption I would say it is basically the differentiability of the distance function uh that means our f of given f of x the data value that you are getting given p. So here P is basically
Speaker A
the u the production posibility set that we are getting given the production technology the efficiency estimate that we are getting it should be uh differentiable in both the argument basically it should be differentiable against X and Y and here that ensures or
Speaker A
this this assumption ensure that the frontier is smooth extension of what we have already assumed and it should be assumed to uh uh this allow us to have an asytoic analysis and also and in the applied mathematical literature you will
Speaker A
see something called lipstick uh continuity the the continuity that we are imposing or assuming in this context going beyond that or beyond uh twice differentiability and other conditions.
Speaker A
So just to summarize we have eight assumptions involved in the bootstrap correction algorithms of simulation.
Speaker A
So as you can see uh the initial assumption we can say that uh up to this IAD and smooth and continuous production function and we can see the truncation of these are basically the new uh assumption that we are seeing from the u second stage
Speaker A
regression point of view and the later assumptions are basically u uh something we have already seen the properties of production possibility set and these assumption will be kept in the back end and what Simar and Wilson algorithm does
Speaker A
keeping these set of assumptions or the statistical properties in the back end they try to simulate the original data generating process involved in the uh determination or estimation of efficiency score and that simulation is not simply simulating by resambling. what they do. These set
Speaker A
of assumption gives us say strong foundation that once you have the explanatory variable we should be or the zed variable we should be able to explain or we should be able to uh define the efficiency score and given that
Speaker A
efficiency score given the values of input we should be able to see how much the output the firm should have produced. So it is basically a reverse of or reverse of the process.
Speaker A
So earlier we were seeing X and Y efficiency and that it regressed against a set of explanatory variables and get the quotient. But that's not what we uh do here in the bootstrapping. We do the we start from the very root given the
Speaker A
environmental factors or the determinance of efficiency that determines the level of output. Oh no that determines the level of efficiency and that determines given the input how much of a could produce that is the interconnection. So keeping these
Speaker A
statistical properties or assumptions in the back end sim will simulate the original generating process and see how much deviation in the in the sample value and the simulated value and that gives you an approximation of efficiency that would have happened in the uh
Speaker A
actual process. That means effic bias in the efficiency that we are estimated by using the sample as compared to the population parameter. We'll see the two algorithms in the next session. Thank you.
Topics:Data Envelopment AnalysisDEA efficiencySimar Wilson algorithmbias correctionserial correlationjackknife procedurebootstrap DEAproduction functionFarrell distance functionstatistical assumptions











