Skip to content
Get App

Corrected ordinary least squares (COLS) and Corrected Mean Absolute Deviation (CMAD)

Introduction to Corrected Ordinary Least Squares (COLS) and Corrected Mean Absolute Deviation (CMAD) for production function estimation and technical efficiency measurement.

Key Takeaways

  • COLS adjusts the OLS intercept to create a deterministic production frontier enveloping all observations.
  • Technical efficiency can be derived from the deviation of actual output from the corrected frontier.
  • CMAD offers a robust alternative to COLS by using median regression to mitigate outlier effects.
  • Parametric approaches assume a known functional form and treat deviations as inefficiency without random noise.
  • Understanding these models is essential for empirical production function estimation and efficiency analysis.

What the video covers

  • The video covers two key concepts: Corrected Ordinary Least Squares (COLS) and Corrected Mean Absolute Deviation (CMAD) models in applied production analysis.
  • It explains the parametric approach to production function estimation, assuming a deterministic frontier without random noise.
  • The production function is conceptualized as the maximum output given inputs and technology, with deviations attributed solely to inefficiency.
  • Standard OLS regression lines do not satisfy the production frontier property because they do not envelop all observations below the line.
  • COLS corrects the OLS intercept by shifting it upwards by the maximum positive deviation to ensure the frontier envelops all data points.
  • Technical efficiency is measured as the ratio of actual output to the frontier output, expressed as an exponential function of the inefficiency term.
  • CMAD uses quantile regression (median regression) instead of OLS to reduce sensitivity to outliers while estimating the production frontier.
  • The video highlights the difference between COLS and CMAD, with CMAD being a robust alternative less affected by outliers.
  • The session emphasizes the foundational role of these models in applied production analysis using MATLAB.
  • Future sessions will address stochastic frontier models incorporating random noise.

Answers

Questions about this video

What is the main purpose of Corrected Ordinary Least Squares (COLS) in production analysis?

COLS adjusts the standard OLS regression line by shifting the intercept to ensure the production frontier envelops all observations, allowing deviations to be interpreted as inefficiency.

How does Corrected Mean Absolute Deviation (CMAD) differ from COLS?

CMAD uses quantile regression (median regression) instead of OLS to estimate the production frontier, making it less sensitive to outliers compared to COLS.

Why can't the standard OLS regression line be used directly as a production function?

Because the OLS regression line represents an average relationship and does not envelop all data points below it, it cannot represent the maximum possible output or production frontier.

Full Transcript — Download SRT & Markdown

00:14
Speaker A
Hi. Welcome back to the course Applied Production Analysis Using MATLAB. Today's session, we'll be covering two concepts.
00:25
Speaker A
One is the corrected ordinary least square model, and the other one is the corrected mean absolute deviation or CMAD model.
00:35
Speaker A
To be frank, with this session, we are entering into the hardcore applied part of this course.
00:45
Speaker A
So far, whatever we discussed are basically the foundational things, especially how we construct an empirical production function given the input-output data.
00:57
Speaker A
And once you have the production function, how to measure technical efficiency or inefficiency. As I mentioned, there are two lines of literature when it comes to applied production analysis.
01:13
Speaker A
The very first is the parametric approach, and the second one is the non-parametric approach. As the name suggests, the parametric approach a priori assumes a functional form, Cobb-Douglas or translog in our case.
01:29
Speaker A
And it estimates the production function and sees how individual observations are deviating from their potential outcome to conceptualize the inefficiency involved in the production process.
01:46
Speaker A
Moving ahead, the parametric approach can also take two variations. Whatever deviation is happening, we consider it as a case of inefficiency.
01:58
Speaker A
And sometimes we give room for random noise or factors other than inefficiency to play a role in the deviation.
02:10
Speaker A
That takes the case of stochastic frontier analysis. So, whatever we are discussing today are very basic concepts that come under the parametric approach that consider any deviation happening from the frontier as a case of inefficiency, or we
02:27
Speaker A
consider our frontier to be a deterministic frontier. There's no random noise involved. We'll see how it can incorporate random noise and take the form of a stochastic frontier at a later stage.
02:43
Speaker A
So, as I mentioned, we'll be covering the concept related to the deterministic frontier. As I explained, it takes the form of a case where we have a production function and whatever deviation is happening for individual observation is purely due
03:01
Speaker A
to the inefficiency. And here, the production function will have a component intercept, the coefficient associated with inputs, and then a random error. Here, we try to conceptualize the error term as the inefficiency or deviation. But our standard error term that we get from an
03:24
Speaker A
OLS estimation is not sufficient because OLS will take a form of this case. Say, this one is the estimated frontier.
03:34
Speaker A
You can see, sorry, not estimated frontier, estimated regression line. And here you can see given Y and X, some of them are deviating from the average volume as above, and some of them are deviating toward below the line that we are
03:54
Speaker A
conceptualizing. So, having said that, we can't consider this function that we estimated as a production function. This is basically the regression line because the production function needs the characteristic that it should envelop the entire observation below.
04:13
Speaker A
Our production function is defined as the maximum possible output given the level of input and technology.
04:19
Speaker A
So, in this session, we'll be using concepts or foundation from OLS itself and see how this OLS estimate can be corrected or adjusted in such a way that it can be used as a production function.
04:39
Speaker A
So, for a simple case, let us consider the case of a trans-Cobb-Douglas production function.
04:46
Speaker A
We can conceptualize log of Y_i and we have n inputs over the, so that can be conceptualized as beta_n log x_i minus U_i.
04:59
Speaker A
So, this minus the U_i is basically the deviation from the production function or the potential output.
05:10
Speaker A
And here we can have only minus values for the error term. We can conceptualize that way because we don't have any deviation which is letting an observation lie above the frontier because if that was possible, then of course our line should
05:32
Speaker A
have passed through that point because given the technology, if there is no randomness involved in the outcome, that the firm could have well achieved.
05:46
Speaker A
That would have been the potential output. So, the production function should pass through that point.
05:54
Speaker A
So, we conceptualize the production function in such a way that their observations are lying below, and few of them or at least one of them will be on the line.
06:03
Speaker A
Then, we can say that that line satisfies the condition of a production function. So, we conceptualize our production function in a Cobb-Douglas form over here, and here we want a regression model with a non-positive disturbance.
06:24
Speaker A
That is the minus U_i. So, how to achieve that? The very basic model that you can conceptualize is like a corrected OLS. We'll see that.
06:37
Speaker A
And once you have the corrected OLS, how to conceptualize technical efficiency? What we can do, we can take the actual outcome.
06:48
Speaker A
That is basically the frontier. And how much it is deviating from the frontier. So, we conceptualize in a form f of x_i beta.
07:01
Speaker A
That is basically the frontier part and minus of U_i is the inefficiency part. So, since you're taking it in the log, so we consider it as f of x_i beta_i exponential of minus U_i.
07:18
Speaker A
And then f of x_i beta is basically the point on the frontier, the deterministic frontier.
07:25
Speaker A
And the ratio of actual output that the firm is getting, that is basically the numerator divided by the f of x_i beta will give us an estimate that is exponential of minus U_i. Basically, this will turn as a
07:42
Speaker A
technical efficiency of firm i. So, how to get these points? For that, we need to start with the OLS, and we correct that OLS through a method that we are discussing today.
08:01
Speaker A
So, OLS is an approach. It was basically proposed by Winston by taking concepts from Farrell et al.
08:13
Speaker A
To conceptualize how to measure efficiency from a deterministic approach point of view. So, in the first approach, like the first step following the OLS model, what we do, we estimate the intercept and slope coefficients involved in the production
08:38
Speaker A
function using OLS. Here, the slopes are expected to be unbiased and consistent estimates, but the intercept, since our error term is expected to have a component that is inefficiency, a systematic component that is inefficiency, or it takes a one-sided
08:57
Speaker A
distribution, we cannot assume our intercept to be unbiased or which is usable for our production function point of view.
09:16
Speaker A
So, what we do, we correct this biased estimate of intercept. So, what the procedure, what we are doing in a diagrammatic point of view, so these are the data points.
09:37
Speaker A
I say this is the regression line that we are getting using OLS. So, we need to make this regression line or the OLS line satisfy the properties of a production function.
09:52
Speaker A
So, for that, what we can do, we can pick the maximum deviation that is happening, maximum positive deviation that is happening from the OLS, so that will basically be the max of U_i.
10:09
Speaker A
And here this one is the beta zero. And what we do, we try to shift our OLS line in such a way that the new intercept that you are getting beta zero hat plus max of U_i hat or U_i.
10:35
Speaker A
So, that will cover the entire observation below the new line that you are getting.
10:45
Speaker A
So, that line that you are getting by following corrected OLS will satisfy the characteristics of a production frontier.
10:53
Speaker A
And here, that's what we are doing in the second step. In the second step, the biased OLS intercept from a production point of view is shifted up or corrected by taking the maximum value of U_i
11:12
Speaker A
and adjusting our intercept in such a way that we get a new line which satisfies the property of a production frontier where entire observations will lie below that.
11:25
Speaker A
Once you have that, we can estimate the technical efficiencies of firms the way we are doing.
11:34
Speaker A
Here, we are defining beta hat zero star as the new intercept that is basically the corrected intercept, which is defined as beta zero hat plus maximum of U_i hat.
11:53
Speaker A
And now, by doing this modification, we can define a new error term or new deviation that is basically the deviation from the corrected OLS line as U_i hat minus maximum of U_i.
12:11
Speaker A
And once you have minus of UI hat star, exponential of minus of UI hat star will give us a measure of technical efficiency.
12:26
Speaker A
Okay? Though it is a very straightforward way of estimating production function from a deterministic point of view, it is not free from limitations.
12:41
Speaker A
The very first limitation as I mentioned, we trust the slope coefficient being estimated using OLS. What we did, we use the same slopes, but only our intercept being shifted.
12:55
Speaker A
So, with that, what we are doing, we are doing a parallel shift where the underlying principle is that our technology remains same for our COLS as well.
13:10
Speaker A
Alternatively, we can say that whatever slopes being got or the technology that we got in the context of OLS is being embedded in the context of COLS as well.
13:22
Speaker A
The second limitation, the line that we are getting not necessarily it will Say this was the OLS that we are having.
13:38
Speaker A
And say this one is the observation that we are having. Since it is a parallel shift this new frontier does not cover or envelop the entire observation as tight as possible.
13:53
Speaker A
Right? And that makes us the case that sometime our estimate will be very sensitive to the outlier.
14:00
Speaker A
There might be an outlier and that outlier will sh- Say the outlier from the positive point of view.
14:08
Speaker A
Say one observation is producing too much because of some unknown reason to us. And this is being included in the sample. So that will shift the frontier that COLS frontier too much above.
14:22
Speaker A
As a result the target output of the non-outlier firms will get estimated too high and it will result in a biased estimator or unreasonable estimation of inefficiency or technical efficiency of the firm.
14:43
Speaker A
So from a very econometric point of view we have seen there is an alternative to our OLS that is absolute deviation approaches.
14:54
Speaker A
That is which is less sensitive to the outlier. Basically it consider absolute deviation and instead of mean sometime it may take the uh deviation from median which is called quantile regression.
15:09
Speaker A
So that makes our estimation less sensitive to the outlier. So that's what corrected mean absolute deviation approach or CMAD does.
15:22
Speaker A
The terminology is a bit confusing. Uh in the standard literature that uh we have kept it as a reference also, they use it as a CMED that corrected mean absolute deviation, but uh from an empirical point of view, instead
15:35
Speaker A
of mean, we'll be taking deviation from median or we consider it as a quantile regression from the estimation point of view that we are going to cover in the next session.
15:47
Speaker A
So, it can be considered as an alternative to corrected ordinary least square. What we do in the first step, mean or median absolute deviation is being uh used or median regression is used for estimating the slope and intercept
16:06
Speaker A
coefficient. In the same way what we did for corrected OLS, the slope remains the same.
16:13
Speaker A
And we do a parallel shift such a way that the intercept is connect corrected.
16:20
Speaker A
And or modi- uh correct modified such a way that uh the entire observations are laying below that. So, the new intercept will become beta zero hat double star just to make it different from what we had it in the context of corrected OLS,
16:35
Speaker A
which is basically the beta zero hat that we estimated using quantile regression plus maximum of UI hat.
16:45
Speaker A
And similarly, what we had it in the context of corrected OLS, we can define minus of UI double star as UI hat minus maximum of UI hat.
16:58
Speaker A
And technical efficiency of the individual observation under CMED can be defined as exponential of minus UI double UI hat double star.
17:11
Speaker A
So, this is what it does uh in the context of corrected OLS. Philosophically, it is very much in line with what we did in the case of corrected OLS.
17:24
Speaker A
But, we can see since it is using median in place of mean, it might be less sensitive to the outliers.
17:34
Speaker A
But, it carries the uh significant limitation or criticism of COLS that what we are doing, we are doing a parallel shift. As a result, the best practice technology remain the same.
17:49
Speaker A
And uh that might be a very uh strong assumption to have. And also uh CMAD also not necessarily cover the entire observations as tightly as possible.
18:04
Speaker A
So, that's all from uh this session. So, just to summarize, in this session, we uh familiarize ourselves with two very popular deterministic frontier approaches that comes under the parametric approach.
18:23
Speaker A
And our fundamental assumption was whatever deviation is happening for an individual observation from the frontier, it's just happening because of the inefficiency.
18:32
Speaker A
And how to get this production frontier in this context, we started with a standard Cobb-Douglas production function. We can do the same with the uh translog production function as well.
18:43
Speaker A
And uh then what we did, we estimated the production function using OLS. But, we were not very happy with that OLS estimate because we cannot consider that as a uh production function, and we need to shift such a way that entire
19:00
Speaker A
observations are lying below that newly estimated uh COLS or CMAD line such a way that we have only one one-sided or negative deviation that is happening for the entire observation or zero observation zero deviation happening for the observation with
19:20
Speaker A
efficiency one. And how CM CMAD and COLS are different? COLS uses OLS and in place of OLS to avoid issue of uh sensitivity to outliers, we use quantile regression.
19:39
Speaker A
So, in next session uh we'll cover the technical aspects of how to estimate COLS and CMAD using MATLAB.
19:50
Speaker A
So, these are the references that we are using. Uh the book by Kumbhakar and Lovell stochastic frontier analysis and Kumbhakar et al.
20:04
Speaker A
book that is a practitioner's guide to stochastic frontier analysis using Stata. Thank you.
Topics:Corrected Ordinary Least SquaresCOLSCorrected Mean Absolute DeviationCMADProduction FunctionTechnical EfficiencyParametric ApproachDeterministic FrontierQuantile RegressionApplied Production Analysis

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries and speaker detection. 30 minutes free.

Or transcribe another YouTube video here →