**Simar–Wilson Bootstrap Procedures: Algorithms #1 and #2 — Transcript & Summary | SozAI**
Source: https://sozai.app/transcript/simar-wilson-bootstrap-algorithms/

Detailed explanation of Simar and Wilson bootstrap algorithms for bias correction in efficiency score estimation using MATLAB.

## Key Takeaways

- Simar and Wilson bootstrap procedures address bias and serial correlation in efficiency score estimation.
- The method involves two-stage regression with truncated regression in the second stage.
- Bootstrap iterations simulate the data generating process to produce bias-corrected efficiency scores and confidence intervals.
- Spuriously efficient observations are identified and excluded to improve estimation accuracy.
- Both CRS and VRS models can be used, with implications for the extent of bootstrap correction.

## What the video covers

- Introduction to assumptions underlying Simar and Wilson bootstrap procedures in efficiency score estimation.
- Explanation of the framework using Q inputs, P outputs, and choice between CRS and VRS models.
- Estimation of output-oriented technical efficiency scores using data envelopment analysis (DEA).
- Discussion on serial correlation in efficiency scores and its impact on regression coefficient estimation.
- Description of Algorithm 1 which corrects serial correlation by estimating beta hats and sigma epsilon hats using truncated regression.
- Explanation of the bootstrap procedure involving multiple iterations (typically 1,500) to draw epsilon values from truncated normal distributions.
- Generation of pseudo efficiency scores (delta i star) and re-estimation of parameters to build confidence intervals.
- Introduction to Algorithm 2 which builds upon Algorithm 1 to further refine bias correction and frontier estimation.
- Handling of spuriously efficient observations by excluding them from regression to avoid bias.
- Mention of both output-oriented and input-oriented counterparts of the procedure with similar serial correlation considerations.

## Chapters

1. 00:00 Introduction and Assumptions of Simar and Wilson Algorithm
2. 02:04 Framework: Inputs, Outputs, and CRS vs VRS Models
3. 03:35 Estimation of Beta Hats from Efficiency Scores
4. 05:07 Excluding Spuriously Efficient Observations and Regression Setup
5. 06:41 Bootstrap Procedure and Monte Carlo Iterations
6. 08:01 Generating Pseudo Efficiency Scores and Confidence Intervals
7. 12:16 Overview of Algorithm 2 and Further Bias Correction
8. 14:42 Output-Oriented vs Input-Oriented Counterparts

Answers

## Questions about this video

What is the main purpose of the Simar and Wilson bootstrap procedure?

The procedure aims to correct bias and serial correlation in efficiency score estimation by simulating the data generating process and providing bias-corrected scores with confidence intervals.

Why are spuriously efficient observations excluded in the regression step?

Observations with efficiency scores equal to one may appear efficient due to sample limitations but might not be truly efficient in the population, so excluding them avoids bias in estimating regression parameters.

How many bootstrap iterations are typically recommended for the Simar and Wilson algorithm?

Typically, around 1,500 iterations are recommended based on Monte Carlo simulations to adequately correct for serial correlation and produce reliable confidence intervals.

## Full Transcript — Download SRT & Markdown

00:14

Speaker A

Hi, welcome back to the course Applied Production Analysis using MATLAB. So, in the last session, we saw the set of assumptions involved in the implementation of Simar and Wilson algorithm.

00:30

Speaker A

Basically, those assumptions are a set of assumptions that we already had in the context of estimation of the efficiency score, say convexity, disposability, the bounded nature of our technology set, or the closed nature, and so on.

00:50

Speaker A

Having these assumptions, as I mentioned already, what Simar and Wilson does is they try to simulate the original data generating process and see how much bias the efficiency estimation could have when you are using the efficiency score, efficiency when you're estimating efficiency scores using a sample.

01:10

Speaker A

So, here for implementing these procedures, or the motivation of this procedure, is basically a case where you have efficiency scores, and these efficiency scores are expressed as a function of some set of variables. We consider it as Z variable.

01:31

Speaker A

This is very much in line with what Coelli 1995 model that we saw in the context of stochastic frontier analysis, but the context was different. There, they were using those Z variables to determine the mean of our truncated normal distribution or whatever. But here, we are exploiting that nature for correcting the bias involved in the estimation of efficiency score and also the serial correlation issue.

01:49

Speaker A

So, what we do in the first instance, we estimate output-oriented technical efficiency using this data envelopment analysis tool. And hereafter, we are going to call our fee in place of, we are using delta star fee scores itself, but just to keep it consistent with the algorithm that we see in the Simar and Wilson, we are considering it as the delta star and it is basically the Farrell distance function and one by that will give you the efficiency score or the Shephard distance function.

02:04

Speaker A

And here we are following a framework. We have Q input and P output and we are keeping the CRS VRS over here with this constraint. Not necessarily you need a VRS constraint for implementing bootstrap correction. Having said that, you can think whether VRS or CRS model will have more correction when you are doing a Simar and Wilson bootstrap correction procedure.

02:23

Speaker A

So, what we are discussing is basically the Simar and Wilson 2007 algorithm. As I mentioned, there are two algorithms in this context. The first algorithm is seemingly simple. What it does, it knows that as a researcher, we have X, Y, and the determinants of efficiency, that is Z variable. So, in the first instance, you are estimating the efficiency score and regressing it against the Z variable.

02:36

Speaker A

Since these efficiency scores are serially correlated, that can create a complication or that can impact the BLUE properties of our coefficients of Z variable involved in the second stage regression.

02:57

Speaker A

So, algorithm one does some corrections in the context of serial correlation. So, what it does using the same LP program that I mentioned in the last slide, they estimate delta hat i's. Basically, that is something we are getting by using the input-output values. And here you can see we have something called D's hat.

03:11

Speaker A

This is basically the data generating process. Or we can even represent it as the production possibility set that we are getting in the context of sample.

03:25

Speaker A

Okay. So, here using the original data that we are having, basically that is a sample data, you estimate delta i hat, and we are putting hat because it is not going to be the actual values of delta, but it is going to be an estimated counterpart of delta i's that you could have gotten in the context of population.

03:35

Speaker A

Now, what we do by using these delta i hats, we estimate beta hat. Beta hat is basically a vector.

03:53

Speaker A

And along with that, you get a sigma hat of epsilon. That is basically the variance of the error terms involved in the functional form, the second-stage regression model from a truncated regression.

04:14

Speaker A

And here this point is very tricky. For doing that estimation, we use only observations for which in the first stage where you are estimating efficiency scores are getting a value greater than one. As I mentioned, these are output-oriented distance functions that we are referring to. If in our output-oriented distance function, if the observation is getting delta i equal to one, that means one by one will give you an efficiency score of one.

04:27

Speaker A

If it is greater than one, it will get an efficiency score between zero and one, but not one.

04:45

Speaker A

Why are we doing such censoring? It is happening because Simar and Wilson algorithm acknowledge the fact that some observations are there in the sample when you're estimating efficiency score, they're getting efficiency score one.

05:00

Speaker A

And they are becoming efficient by the virtue of the procedure that they might not be efficient when you are having more observations or when you put this observation in the population. And they call observation with value delta i hat as basically spuriously efficient. Or they may or may not be efficient when you are having a projection towards the population production frontier or the actual true frontier, I would say.

05:07

Speaker A

So, that's why we are removing them from the regression where you are trying to estimate beta hats, that is the vector, and the sigma hat epsilon as an estimate. So, this is the population regression function, and this is basically the sample regression function. You get beta hats here. Okay? Now, once you have these betas and beta hats and sigma epsilon hats, now we are going to get another set of values that is basically beta hat star and epsilon hat, sorry, sigma epsilon hat star. How are we going to get it? So, we'll be running one procedure for L times. That is basically L1, I would say. This is basically the loop one. Here, there is only one loop, that's why we are not putting any subscript. So, here we run some procedure that we are going to discuss for L times. So, generally, they do it for 1,500 times. In the Monte Carlo simulation that is done as a part of Simar and Wilson algorithm, it says that or claims that 1,500 as an iteration or the number of iterations works decently for the correction of serial correlation. How are we getting those values?

05:17

Speaker A

In this loop, in the first instance, we draw epsilons from a distribution of N0 sigma epsilon hat squared. This sigma epsilon hat squared is coming from the last step.

05:33

Speaker A

So, it is going to be a distribution. We will have a distribution of this sort with a zero mean and sigma epsilon hat squared as the variance.

05:55

Speaker A

From this distribution, not the entire distribution, basically, it will have a truncation at 1 minus, say, this is like C, I call it as 1 minus Z i theta hat, and from this distribution, this is very similar to what truncation that we discussed in the context of SFA. Basically, it can take only truncation, but here truncation is not at zero. If you put a truncation at zero, it is going to be very complicated.

06:08

Speaker A

We put a truncation at 1 minus Z i beta in such a way that when you plug in these epsilon scores that you're drawing from this distribution to this function that we're getting here, this function, it should not go less than one, okay? And then what we do for step three, step two in the loop that we're referring, for each observation, you get a pseudo efficiency score. So, basically, this delta i star is basically the pseudo efficiency score that you're estimating by plugging in the epsilon i that we're drawing in the earlier step.

06:18

Speaker A

Then we use maximum likelihood estimation. Instead of delta i hat that we had in the earlier context, we use delta i star for estimating beta hat star and sigma of epsilon star.

06:41

Speaker A

Now we can use these values in A, that is basically each iteration, you get beta hat star and sigma hat star. So, each iteration, you store them, and then use these values to get a confidence interval.

07:00

Speaker A

spuriously efficient. Or they may or may not be efficient when you are having a projection towards the uh population production frontier or the actual the true frontier, I would say.

07:18

Speaker A

So, that's why we are removing them from the regression where you are trying to estimate beta hats, that is the vector, and the sigma hat epsilon as an estimate. So, this is the population regression function, and this is basically the

07:36

Speaker A

uh sample regression function. You get beta hats here. Okay? Now, once you have these uh betas and beta hats and sigma epsilon hats Now, we are going to get another set of values that is basically beta hat star and

08:09

Speaker A

epsilon hat. Sorry, sigma eps- sigma of epsilon hat star. How are we going to get it? So, we'll be running one procedure for L times. That is basically L1, I would say. This is basically the loop one. Here, there is only one loop,

08:29

Speaker A

that's why we are not putting any subscript. So, here we run some procedure that we are going to discuss for L times. So, generally, we they do it for 1,000 500 times. In the Monday Carlo simulation that uh done as

08:44

Speaker A

a part of Simarin Wilson algorithm, says that or claims that uh 1,500 as an iteration or the number of iteration works decently for the correction of serial correlation. How are we getting those values?

09:01

Speaker A

In this loop, in the first instance we draw epsilons from a distribution of N0 sigma epsilon hat square. This sigma epsilon hat square is coming from the last step.

09:18

Speaker A

So, it is going to be a distribution. We will have a distribution of this sort with a zero mean and uh sigma epsilon hat square as the variance.

09:29

Speaker A

From this distribution we not the entire distribution, basically, it will have a truncation at 1 minus Say, this is like C I call it as 1 minus Z I theta hat, and from this distribution, this is very similar

09:48

Speaker A

to what truncation that we discussed in the context of uh SFA. Basically, it can take only truncation, but here truncation is not at zero. If you put a Z truncation at zero, it is going to be very complicated.

10:01

Speaker A

We put a truncation at 1 minus Z I beta such a way that when you plug in these epsilons scores that you're drawing from this distribution to the uh this function that we're getting here this function.

10:22

Speaker A

It should not go less than one, okay? And then what we do for step three step two in the loop uh that we're referring, for each observation, you get a pseudo efficiency score. So, basically, the this delta I star is

10:42

Speaker A

basically the pseudo efficiency score that you're estimating by plugging in the epsilon I that we're drawing in the earlier step.

10:54

Speaker A

Then we use maximum likelihood estimation instead of delta I hat that we had in the earlier context, we use delta I star for estimating beta uh hat star and sigma of epsilon star.

11:12

Speaker A

Now we can use these values in A, that is basically each each iteration, you get uh beta hat star and sigma hat star. So, each iteration, you store them, and then use these values to get a confidence interval of our

11:38

Speaker A

beta values. Okay, so that's how we are uh getting rid of the serial correlation issue that we may face in the context of the efficiency scores we used in the second stage regression. So, this uh basically at the end you'll be getting A

11:55

Speaker A

as a result. It will have uh the vector of beta and value of sigma for L times or 1,500 times, and get using that you get a uh standard error of that or whatever, and you can use those variables

12:16

Speaker A

or that information or variance of that to construct a confidence interval. Moving ahead quickly, the algorithm two, as I mentioned, algorithm two of Simar and Wilson is basically building upon the algorithm one.

12:33

Speaker A

Here, along with the serial correlation issue that we are facing in the context of DEA efficiency estimation or the second stage regression involved in the uh non-parametric DEA followed by a non-parametric DEA estimation efficiency estimation, we try to correct the bias involved in

12:56

Speaker A

the estimation itself. So, in the first instance, when I say it is upon the algorithm one, we get a feel that, okay, algorithm one is done in the first instance, and um then something else is done. No, that's not how we do.

13:13

Speaker A

In the first instance or the first step, we correct the bias, and these bias-corrected efficiency scores are feed into the algorithm one in the second step, and there these serial correlation issues get, uh, corrected.

13:31

Speaker A

So, as I mentioned, the first step of algorithm, uh, or first loop, I would say, we have two loops, loop one and loop two.

13:46

Speaker A

So, loop one corrects for bias and loop two basically corrects for serial correlation. Okay. So, let's see how it works in the context of loop one.

14:07

Speaker A

So, as we did in the context of algorithm one, using the, uh data given the, um, production possibility set, we estimate the delta score. And we consider any observation which is coming, uh, greater than one only in the observation.

14:28

Speaker A

And any observation getting value one, we consider them as the spuriously efficient observation. Okay.

14:34

Speaker A

So, if you have n observation, EM will be the number of observation that go into estimate uh, or used for estimation of beta hat and epsilon, uh, variance of epsilon, that is basically the estimate of, uh, variance of epsilon, that is

14:52

Speaker A

basically sigma epsilon hat. Okay. That is basically done using this regression, okay? And here, once we have that values of uh, beta hat and delta hat, now we go for a uh, bias correction. That's the crux of the bias correction.

15:20

Speaker A

As I mentioned, from a distribution which has a zero uh, mean and variance of sigma epsilon hat square.

15:35

Speaker A

Plus one triangulation at one minus ZI beta. We draw one epsilon I and the same way what we did for algorithm one, we plug in that epsilon I for each observation and and their X ZI values. So, that

15:52

Speaker A

gives you a new sigma I that we are calling it as a sigma I star.

16:05

Speaker A

So, now once we have this sigma I star, what we can do? We can plug in this into the original data that we are having. This is basically a pseudo efficiency score.

16:15

Speaker A

Once you have this thing, what we can do? We can plug in that to the output values and to move like to extract out the actual efficiency score, we multiply that with the uh ZI hat that we got in the first stage

16:34

Speaker A

and divided by delta I star for each observation. So, here input remain the same, input values remain the same, but at the end, since it is an output-oriented approach, we'll be getting XI star and YI star. Basically, a

16:52

Speaker A

pseudo data that you are getting for the purpose of simulating the original data generating process or getting a bias correction.

17:01

Speaker A

Now, what we do? We plug in these values, basically you know, XI star and YI star.

17:10

Speaker A

Oh, no. Not XI star. So, using this uh XI stars and YI stars, we estimate a new frontier that is basically D hat star or following the data generating process that manner, we estimate DI Sorry D's D hat star and then we see what is the

17:33

Speaker A

efficiency of our actual input-output bundle again that from this. And we store that as a delta I hat star as the distance function.

17:47

Speaker A

Then, what we do? We collect We get a efficient like the bias score and the bias is basically uh defined as the difference between delta double hat and delta I hat and using that difference following the formula, we get a

18:10

Speaker A

bias correction and and we take a uh new refined delta I S star bias corrected efficiency score.

18:21

Speaker A

So now what we do? We go for the step two of So this is basically the bias correction procedure.

18:27

Speaker A

Now we go for the step two of algorithm two. Basically, once we are going through that, we get an idea. This is basically the same implementation what we saw in the context of our uh serial correlation correction.

18:41

Speaker A

So here what we do, we use uh delta I double star. This is basically the bias corrected distance function scores.

18:59

Speaker A

And then using that, we get an estimate of beta double hat and along with that you'll be getting a estimate of sigma epsilon double hat. Okay. So only slight difference instead of estimated efficiency score estimated distance function scores, we are getting

19:19

Speaker A

the bias corrected distance function scores over here. Then, what we do, we use maximum likelihood estimation to get a value of delta I double star and then for that actually uh we follow the more or less same procedure what we did in the context of

19:39

Speaker A

last uh Simar Wilson algorithm one. And once we have that, we get a uh C as the matrix with doing a loop. Basically, it will be done in a loop for L2 times. L2 here remains 1,500.

19:58

Speaker A

And L1 here, as I mentioned, here we are doing this one for L1 times and here generally people do it for 1,000 times.

20:10

Speaker A

Right? And the Monte Carlo simulation say that the kind of bias correction or the procedure or the data generation process involved, uh the 1,000 iteration does decent correction in the context of uh non-parametric DEA first stage bias correction and 1,500 in

20:31

Speaker A

the context of second stage of our second loop of our uh algorithm two or the first loop of our or first or only one loop of our algorithm.

20:45

Speaker A

So, uh this is how Simar Wilson algorithm two corrects for the bias as well as the serial correlation issue.

20:54

Speaker A

I would encourage you to go through the reference. This is basically the classic paper by Simar and Wilson published in the year 2007, which is widely cited not only in the context of efficiency estimation even in the context of

21:13

Speaker A

productivity analysis based on DEA that is basically the Malmquist productivity index people use Simar and Wilson algorithm.

21:21

Speaker A

So the main criticism was the earlier models before Simar and Wilson algorithm follow a very naive bootstrap procedure simply by re-sampling or without acknowledging the true data generating process.

21:40

Speaker A

What Simar and Wilson algorithm does, it tries to follow the original data generating process so the statistical process involved in determination of efficiency score as much as possible and try to simulate that then get an estimate of bias in the first

21:56

Speaker A

stage of efficiency estimation and try to correct the serial correlation natural get an estimate of confidence interval and so and so or a statistical inference in the second stage which is free from the serial correlation limitation.

22:14

Speaker A

As I mentioned there are two algorithms and if you see algorithm two is building upon algorithm one and the first loop of algorithm one is basically the bias correction. So suppose you don't have a second stage regression or we don't you

22:29

Speaker A

are not interested to see what are the determinants of efficiency you just see you just want to get an bias corrected estimate of your efficiency score first loop of algorithm two works.

22:43

Speaker A

You can stop with that. The one very big limitation of our model from an applied point of view not a theoretical point of view for implementing Simar and Wilson 2007 algorithm bias correction, you need Z variable. Simply having X and Y does

23:09

Speaker A

not help you to uh do a bias correction. You need at least one variable which is not a candidate to be included in the production framework as an input or there is a strict separability between X and Y X and Z.

23:26

Speaker A

And we need that variable to be included in the model or algorithm to simulate the original data generating process. So, that is one data constraint.

23:36

Speaker A

And you can uh use that. So, here uh whatever we discussed is basically the output-oriented counterpart. We can have an input-oriented counterpart. In input-oriented counterpart, uh the serial correlation component remain more or less same.

23:54

Speaker A

But, the slight difference here instead of uh delta I had, say we'll be having theta I star, which will take a value basically theta I star in the context of um our input-oriented technical efficiency less than or equal to one. So, accordingly we

24:12

Speaker A

will have to change the truncation. So, it will not be a truncation at uh 1 minus theta ZI. It This truncation will not work in this context. You may have to take a truncation at minus of uh theta ZI or minus of C in the

24:30

Speaker A

uh if I consider this as a uh single value. Over and above step 3.3 may change slightly. Here, instead of this uh creating a pseudo uh sample by modifying the output, we'll be multiplying that or modifying our input

24:54

Speaker A

values and getting a pseudo sample of X star and Y star. Here, Y star remain the same as Y.

25:04

Speaker A

X star will get adjusted accordingly. And the model works that way. Moving ahead, we can have a uh panel framework over here. But, as I mentioned, the original model by Simar and Wilson is not very compatible with the

25:21

Speaker A

panel framework. It is very uh confined to cross-section data. If you have a panel data, what you can do in the Z actually, you get one more variable. The time will also become a uh determinant of efficiency. It is like

25:39

Speaker A

very uh clear-cut uh separability condition being satisfied, and you can use that as a new variable.

25:47

Speaker A

But, the way you create the frontier, it becomes very complicated. Uh you may have to go for a contemporaneous frontier framework or sequential frontier frameworks, so on.

25:58

Speaker A

So, two what things and then Zelenyuk I remember has come up with a new paper uh after uh 2007 developments in the bootstrap algorithm of Simar and Wilson 2016. So, they do a two uh double bootstrap procedure for correcting the

26:21

Speaker A

correcting the bias and serial correlation issue in the DEA framework for a panel data. You can refer to that.

26:28

Speaker A

So, here the main reference, as I mentioned, it remains Simar and Wilson 2007 estimation inference in two-stage semi-parametric models of production process, but published in general econometrics. Okay, thank you.

Topics: Simar and Wilson bootstrap procedure efficiency score estimation data envelopment analysis serial correlation correction truncated regression bias correction Monte Carlo simulation production frontier Applied Production Analysis

Study this video

- [Flashcards from this video](https://sozai.app/tools/ai-flashcard-generator/?from=transcript&id=76479)
- [Quiz on this video](https://sozai.app/tools/quiz-generator/?from=transcript&id=76479)

Free, made by AI from this transcript. No signup.


---
This is the markdown twin of https://sozai.app/transcript/simar-wilson-bootstrap-algorithms/ — the same content, without the markup.
Published by SozAI (https://sozai.app). Reuse and quotation are allowed with attribution and a link back.
Machine-readable index: https://sozai.app/llms.txt · data API: https://sozai.app/api/
