Skip to content
Get App

Stochastic Frontier Analysis: Estimation

This video explains stochastic frontier analysis estimation using maximum likelihood estimation as an alternative to OLS, focusing on econometric concepts and model assumptions.

Key Takeaways

  • MLE provides a robust alternative to OLS for estimating stochastic frontier models where error terms are not normally distributed.
  • IID assumptions are critical for formulating the joint likelihood function in regression analysis.
  • Stochastic frontier models include a one-sided inefficiency term to capture production inefficiency.
  • Log-likelihood maximization is used to estimate model parameters due to the non-linear nature of the likelihood function.
  • The ratio of variances (lambda) helps interpret the relative contribution of inefficiency and noise in the model.

What the video covers

  • Introduction to stochastic frontier analysis and its estimation using MATLAB.
  • Explanation of maximum likelihood estimation (MLE) as an alternative to ordinary least squares (OLS).
  • Discussion on the assumptions of independent and identically distributed (IID) variables in regression models.
  • Derivation of the likelihood and log-likelihood functions for normally distributed errors.
  • Estimation of unknown parameters beta_1, beta_2, and sigma squared by maximizing the likelihood function.
  • Explanation of the non-linear nature of the likelihood function and the use of log transformation for simplification.
  • Introduction to the inefficiency component in stochastic frontier models and its one-sided distribution.
  • Discussion on the joint distribution of noise and inefficiency components and the role of parameters like lambda.
  • Overview of model testing, efficiency estimation, and extensions to panel data and different inefficiency distributions.
  • Explanation of how stochastic frontier analysis separates noise from inefficiency in production analysis.

Answers

Questions about this video

Why is maximum likelihood estimation preferred over ordinary least squares in stochastic frontier analysis?

MLE is preferred because the error terms in stochastic frontier analysis are not normally distributed, violating OLS assumptions. MLE accounts for this by maximizing the likelihood of observing the data given the model parameters.

What does IID mean and why is it important in this context?

IID stands for independently and identically distributed. It means each observation is independent of others and drawn from the same distribution, allowing the joint probability to be expressed as a product of individual probabilities, which is essential for MLE.

What role does the inefficiency component play in stochastic frontier models?

The inefficiency component captures deviations from the frontier that represent inefficiency in production. It is modeled as a one-sided distribution, typically non-negative, to separate inefficiency effects from random noise.

Full Transcript — Download SRT & Markdown

00:05
Speaker A
[music] [music] Hi. Welcome back to the course of light production analysis using MATLAB. In the last session, we diagrammatically introduced the concept of stochastic frontier analysis.
00:28
Speaker A
In this session, we are going to see the same model from an estimation point of view.
00:34
Speaker A
In the first few minutes, I'll be spending time to familiarize you with the maximum likelihood estimation, which is going to be the backbone of the estimation involved in the stochastic frontier analysis. Also, we see how the MLE
00:52
Speaker A
that we do in the case of stochastic frontier analysis. Just a disclaimer, this course, this session is going to be very lengthy.
01:02
Speaker A
So, those who are familiar with the concept of maximum likelihood estimation may skip to the main component of maximum likelihood estimation involved in the stochastic frontier analysis.
01:17
Speaker A
Even if you have a background in econometrics, I would like to request you go through this content so that you can refresh your concept of maximum likelihood estimation and get an idea of how we use the same tool
01:31
Speaker A
as against OLS in the context of stochastic frontier analysis. So, basically, maximum likelihood estimation is placed as a substitute for ordinary least squares estimation.
01:47
Speaker A
So, say we start with a functional form Y_i equal to beta zero plus beta two X_i plus U_i.
01:55
Speaker A
Not necessarily inputs and outputs. This is basically a general framework that we are discussing over here.
02:01
Speaker A
So, here what OLS does, OLS tries to get a value for beta one and beta two. We call it beta one hat and beta two hat such a way that the sum of U_i hat squared is
02:19
Speaker A
minimum. You know, sum of U_i hat squared is minimized. This is what our OLS does.
02:33
Speaker A
What maximum likelihood estimation does, philosophically it remains the same. It tries to get the values of beta one hat, beta two hat, and over and above we have sigma hat squared such a way that the moment you plug in
02:54
Speaker A
this beta one, beta two, and beta n, we have, say here we do have only beta two and sigma hat squared.
03:04
Speaker A
The moment you plug in these values into this model, the likelihood of observing Y_i will be maximized or the likelihood function of the Y_i is maximized by getting this value. So, here instead of dash, we use hat, we use the tilde.
03:25
Speaker A
Basically, it's an alternative to our ordinary least squares estimation and we try to get the estimate such a way that the likelihood of observing Y_i is maximized by plugging in the value of beta one, beta two, and the variance estimate
03:45
Speaker A
for the given value of X_i. So, why do we need an alternative estimate over here? As I mentioned, our error terms are not normally distributed in the context of stochastic frontier analysis, so we need an alternative tool for
04:00
Speaker A
estimating the coefficient. That is the motivation that we are having over here.
04:06
Speaker A
So, how are we proceeding? So, here in the simple case, this is not the stochastic frontier we are talking about. We are simply talking about the two-variable regression model. We have Y_i as the dependent variable and X_i as
04:20
Speaker A
the independent variable. Here we assume Y_i to be independently and identically distributed with normal mean that is basically beta 1 plus beta 2 X_i and sigma squared as the variance.
04:34
Speaker A
This concept IID is very theoretically involving. So, independently and identically distributed. So, the independently distributed component in our production framework or this Y_i is coming from the fact that each observation Y_i is non-dependent on each other.
04:53
Speaker A
Or as we can say, having information about Y_i of one observation, say Y_j, does not give you any hint about Y_k, that is another observation that we are having.
05:06
Speaker A
In the context of time series data, this property may get violated because having information about Y_t, you'll be able to get information on Y_t plus 2 or T plus 1 will be depending on the
05:22
Speaker A
Y_t values. So, here that is the concept of independently distributed. Identically distributed are coming from the fact that these observations are drawn from the same population.
05:37
Speaker A
So, each of the observations has the same population density function or probability distribution in the back end, the same probability distribution function for each observation to be included in the sample.
05:51
Speaker A
So, because of that, we can call all of them identically distributed. So, these are identically distributed.
05:58
Speaker A
It has a lot of implication if it is not independently distributed. We cannot conceptualize the probabilities and joint probability as a product of these distributions. That is the implication, as I mentioned.
06:16
Speaker A
And identically distributed comes and plays a very significant role because of the identically distributed nature of these Y_i, we are able to have a single functional form for the probability distribution of all these observations.
06:33
Speaker A
So, as I mentioned, we have n observations Y_i, i equals 1 up to n. And then, since all of them are independently distributed or all of them are identically distributed, we can conceptualize their probability density as one or similar. That is what we are
06:53
Speaker A
conceptualizing it as f of, and f of Y_1, Y_2 up to Y_N given the values of beta_1 plus beta_2 X_i and the values of sigma squared.
07:04
Speaker A
And since all of them are independently distributed, we can take the product of these independent density functions for getting the joint distribution, which is going to be f of Y_1 given beta_1 plus beta_2 X_i sigma squared into f of Y_2 up to f of
07:24
Speaker A
Y_N. So, these are like this character we could conceptualize from the identically distributed nature and this character or this joint distribution or joint density function we could get from the fact that these are independently distributed.
07:41
Speaker A
Then, here we are assuming this to be normally distributed. So, it has a 1 by sigma root 2 pi exponential of 1 by 2 Y minus the mean that is basically Y minus beta 1 minus beta 2 X squared sigma
07:55
Speaker A
square as the functional form. And then the moment you take the product n times product of this distribution, we get a functional form that is basically 1 by sigma to the n root of 2 pi to the n exponential of 1 by 2 sum of
08:12
Speaker A
Y minus beta 1 minus beta 2 X squared divided by sigma squared. So, this sum is coming from the product that you are taking because you have to do it n times.
08:26
Speaker A
So, as I mentioned, here this is the joint density of the independently and identically distributed Y_i, i equals 1 up to n.
08:38
Speaker A
And now what MLE does, it tries to get the estimate that is beta 1 tilde, beta 2 tilde, and sigma squared tilde for the population parameter that is basically beta 1, beta 2, and sigma squared. So, our MLE involves the
08:56
Speaker A
estimation of unknown parameters that is basically beta 1, beta 2, sigma squared and we get an estimate of that which is represented tilde over here such a way that the probability of observing Y for the given values of X are
09:19
Speaker A
maximized. How do we do that? We have this functional form. This is not that easy to solve because it is in a non-linear form. So, we take the log of that. So, log of ln lf that is basically the log likelihood function
09:33
Speaker A
minus n log sigma and log sigma. And here n by 2 log 2 pi plus 1 by 2 sum of Y minus beta 1 minus beta 2 X squared divided by sigma squared.
09:51
Speaker A
And the sigma is not that easy to conceptualize. We take the square of that, that is basically the variance. So, it becomes 1 minus n by 2 log of sigma squared and then we have this as a function of
10:06
Speaker A
beta 1, beta 2, and sigma squared. So, this entire log likelihood function will become the function of beta 1, beta 2, and sigma squared. So, our objective is to, or the objective of maximum likelihood function is to get
10:20
Speaker A
the estimate of beta 1, beta 2, and sigma squared. These are the unknown parameters such a way that the moment you plug in these values, that is the estimate of these values into the model, this likelihood function is being
10:34
Speaker A
maximized. So, how do we do that? As I mentioned, this is basically the function of beta 1 and beta 2 and sigma squared. So, we can take the first difference of that and by taking difference that to zero, you
10:47
Speaker A
get an idea. So
11:06
Speaker A
OLS does the same estimate. Okay. But, that's not the case in our uh stochastic frontier analysis.
11:16
Speaker A
In our stochastic frontier analysis, as I mentioned, we do have two components for our return that is basically the maximum likelihood estimation that we are going to do here.
11:27
Speaker A
So, here we do have two components for our return. One is v i which is basically the i.i.d. normal with zero mean and sigma v squared variance.
11:36
Speaker A
But our UI is creating the problem or that is the unique component that we bring in in the context of stochastic frontier analysis.
11:45
Speaker A
UI is distributed independently identically uh but with a case that it has only one-sided distribution. That is basically only the positive sided distribution.
11:57
Speaker A
Here we are considering the case of with zero mean. Of course, we can consider the case with a non-zero mean.
12:03
Speaker A
We'll come to that point as well. With uh so the UI is distributed independently and identically with the one-sided distribution, zero mean and sigma u squared.
12:15
Speaker A
Now it comes the complication. Our VI and UI determines the composite error term. So we need two independent uh density function or two different density function for UI and VI.
12:30
Speaker A
So since VI is normally distributed with zero mean and sigma u sigma v squared, we can conceptualize it as 1 by sigma v square root of 2 pi exponential of minus v squared 2 sigma v squared.
12:46
Speaker A
But our UI is one-sided distributed, so it will have the same almost same uh structure but in order to ensure that the moment you integrate it from zero to infinity f of say u d u this has to be one. So
13:05
Speaker A
since we are doing a truncation at zero in the earlier or in the half it is a coming from a mother distribution which has a uh domain of minus infinity to infinity.
13:18
Speaker A
Right? So this is the domain of the month the actual distribution where we are doing a truncation.
13:24
Speaker A
So in order to ensure this thing uh make sure this property that the moment you are integrating the function from zero to infinity, that is basically the domain, it comes one. So, we get a value of two.
13:40
Speaker A
So, it's a slight modification of our uh, f of v, but here we have two into uh, one by sigma u square root of 2 pi exponential of minus u square 2 sigma u square. So, this is like
13:57
Speaker A
coming from the fact that we have a zero mean for both of them. Surely, we are going to relax this assumption and go for a uh, advanced model which has a non-zero mean for our inefficiency component.
14:12
Speaker A
So now it's a very tricky stuff that we are doing. We assume our uh, u's and error terms to be independent. It's a very strong assumption to have, but in this context of stochastic frontier analysis, we have to assume that our
14:30
Speaker A
noise component is in- independent of the inefficiency component. So, the moment we have such framework, we can consider the joint distribution of u and epsilon as the product of these two distributions that is going to be 2 by 2
14:45
Speaker A
pi sigma u sigma v exponential of minus u square Here, we are conceptualizing u as epsilon plus uh, sorry, v as epsilon plus u, which is doable.
15:02
Speaker A
Uh, so this component become epsilon plus u square divided by sigma v square. So, this becomes a function of u and epsilon.
15:15
Speaker A
And the marginal density of epsilon can be obtained by integrating u out of f of u epsilon, because we need a uh joint density or a marginal density for going forward. So, it is integrated in the domain of u that is basically uh f
15:33
Speaker A
of epsilon equal to integral of 0 to infinity f of u e du. This is going to be the function that we are getting for the joint distribution of or the marginal density of our u and here we bring in
15:56
Speaker A
new parameter that is basically sigma. Sigma is conceptualized as sigma u squared plus sigma v squared square root of the same.
16:06
Speaker A
And alternatively we can conceptualize lambda. Here we do have a lambda. Lambda is basically the ratio of sigma u to sigma v.
16:16
Speaker A
So, this formulation may find a bit tricky, but we are taking taking it from the maximum likelihood framework only, but this one instead of having uh distribution of u, we are considering it as a marginal density of epsilon which is going to be the
16:32
Speaker A
function of both our u and v over here or u and epsilon over here.
16:41
Speaker A
So here simplifying, we can conceptualize f of epsilon as 2 by sigma into phi of uh phi This phi is basically the capital phi and basically the cumulative distribution. Here, this is slightly bold cumulative density function and this
17:03
Speaker A
phi's are basically the probability density function. Uh so, phi of epsilon divided by sigma into the cumulative distribution of 2 minus e lambda sigma.
17:17
Speaker A
It's like a coming from the standard normal distribution. Okay. So, from a non-technical point of view, the lambda provides an indication of the relative contribution of U and V in determining the epsilon that we discussed in the context of
17:36
Speaker A
appropriateness test. And if lambda tends to infinity, that means either sigma U square sigma V square is tending to infinity or sigma U square is tending to zero.
17:50
Speaker A
That means the V component, that is the noise component, dominates in determining uh our composite error term, which says that it can be taken back to a wireless production function and there is no uh technical inefficiency involved in
18:09
Speaker A
the production framework that you are talking of. Alternatively, if lambda tends to infinity, that means U is dominating V and the production function without any noise can be used. Similar to our CULS or CMAD can be used in
18:29
Speaker A
this context. So, here uh it's a slightly technical component. Here, we do have epsilon as a joint function of our U and uh V.
18:44
Speaker A
And our main task is to get a conditional value of U for the given value of epsilon or get a value of uh or get a function that represent our uh UIs. So, this is basically a J L M S
19:03
Speaker A
Jondrow et al. uh decomposition that say that f of UI can be considered as a ratio of f of U epsilon divided f of e which can be used for further analysis that you basically the technical inefficiency analysis and here
19:23
Speaker A
as we did in the case of our graphical representation the technical efficiency of i observation is considered as exponential of minus UI hat.
19:32
Speaker A
And as I mentioned, we do have a package that is basically the frontier here. In the case of frontier, they conceptualize a new terminology that is basically gamma which is very equivalent of the lambda that we conceptualize as a ratio sigma U squared to sigma
19:50
Speaker A
squared that in the total variance of our uh composite error term how much our inefficiency component is contributing.
20:00
Speaker A
So that gives us a uh measure of or the measure of the contribution of our inefficiency component in the composite error term. So higher the gamma much of the variation in epsilon is happening due to the inefficiency component. So that gives us
20:20
Speaker A
a hint of our appropriateness of our stochastic frontier model. So uh moving ahead, we do have varied formulations or varied variety of uh models that comes under the stochastic frontier analysis. It start with the Ignor Lovell and Schmidt
20:42
Speaker A
uh the very early model. Between that after that there are several models. Out of that, the most popular model that has a lot of empirical applications are basically the It is basically a panel and panel data model which accounts for or which is
21:04
Speaker A
capable of handling the unbalanced panel model and here it was like a same functional form y i t equal to x i t beta plus the composite error term and then we take the log of that and so and so
21:24
Speaker A
and here u i t is conceptualized as u i exponential of minus eta t i t minus t here you can see it as a non-zero mean so here you can see that as a mu as the non-zero mean so this is basically a
21:46
Speaker A
not the half normal model that we consider this is basically a truncated normal model so the original model proposed by Battese 1992 is basically a model based on truncated normal distribution and here this eta has also to be
22:08
Speaker A
estimated basically the eta is basically the coefficient associated with the time because it's a panel data we are considering similarly what we had sigma squared also conceptualized as sigma v squared plus sigma u squared or gamma is basically
22:23
Speaker A
conceptualized as ratio of sigma u squared to sigma squared and by modifying the model eta if you keep it as zero that means there is no impact of time in the model or inefficiency so it's basically a time invariant model that was being
22:42
Speaker A
proposed by Battese and Coelli before that that is basically a time invariant model which is going to be a more of a cross-sectional framework so before this model we had a model which is basically based on balanced data framework which is
22:58
Speaker A
basically the model in 1988 and if you impose mu equal to zero this is basically a half normal model model framework. The error terms or inefficiencies assumed to be half normal distributed, which is basically a model by Pitt and Lee.
23:14
Speaker A
And so if you keep T equal to one, this is basically the very original model that I'm referring that you basically Aigner, Lovell, and Schmidt 90 77. So it is basically a reverse order that we are referring to. So in the year
23:28
Speaker A
1977 Aigner, Lovell, and Schmidt proposed the model stochastic frontier analysis. Then Pitt and uh Lee conceptualized a model which is basically a half normal framework model.
23:44
Speaker A
Uh very similar to the Aigner ALS model. Then uh Coelli Battese and Coelli 1988 is basically a model that imposes balanced panel or that goes for a panel data model which is a which has a restriction of balanced panel.
24:02
Speaker A
Then we do have a time invariant model that keeps uh eta equal to zero and the 9092 model.
24:11
Speaker A
Beyond that, there is a model called Battese and Coelli 1995. It is basically a model that incorporate environmental environmental variable that is a determinant of inefficiency also into the model. So because of that reason, we call it as a
24:28
Speaker A
uh inefficiency component model or the error component model here. So here we do Generally what uh the model does, they estimate the efficiency in the first stage and then these efficiency scores are regressed against a set of variable that is
24:51
Speaker A
basically the Z variable that I'm referring or the determinants of efficiency or inefficiency in the second step, which is not theoretically that sound. In order to overcome this uh framework, Battese and Coelli extended the model by Kumbhakar et al. or Schmidt
25:07
Speaker A
and Stevenson model, which introduced inefficiency effect that is basically the UI as a uh inefficiency expressed as a explicit function of the determinants of inefficiency that is basically the Z in the model.
25:22
Speaker A
So, this extends the model uh basically it is a Battese and Coelli is an extension of the earlier model by Kumbhakar et al. 19 91, which uh extends the model with allocative efficiency imposed, remove the first-order profit maximization
25:39
Speaker A
condition, and also which is a accounting for the panel data model. This is the main specification over here. So, UI still half one-way distributed or one-side distributed with normally distributed, but we do have a uh non-zero mean that is basically MIT
25:58
Speaker A
sigma U squared as the variance. And this MIT is basically a function of uh MIT is basically a function of Z or lambda Z in this context.
26:09
Speaker A
Uh or yeah, it is in the matrix form ZIT lam- lambda ZIT delta. So, not lambda delta.
26:18
Speaker A
And uh so, then what we do we plug in these ZIs also in the original estimation that we are doing. We'll be going in details when we are doing the same example using MATLAB.
26:32
Speaker A
And uh here as against the earlier model, suppose you are estimating the 1995 model, you need X data. Say, you have matrix of X. You have Y data that is only one Y variable.
26:48
Speaker A
And along with that you need a determinants of efficiency that is like say few or one or few uh variable that capture the factor or that represent the factor determining inefficiency also into the model.
27:07
Speaker A
So, uh this is something we have already discussed. To see the appropriateness of SFA, we had skewness-based measures of uh test that is basically Schmidt and Lee, which is basically based on M3 and M2 and modified by Coelli. So, now after
27:27
Speaker A
seeing these models, we have something called gamma, which is basically uh annotation being used in frontier, which is represented as a ratio of sigma u squared to the sigma squared.
27:38
Speaker A
And uh using this gamma, we can see uh what is the contribution of inefficiency in the overall variance of the composite error term, and based on that, you can see whether the stochastic frontier analysis that you did was a toy
27:54
Speaker A
of a meaningful practice. Okay. And if gamma equal to zero, that means uh if if gamma equal to zero not rejected, this would indicate that the sigma u squared is zero. There is like there is no it's a specification with parameter
28:09
Speaker A
that can be consistently estimated with OLS itself, and there is no inefficiency uh that we need to bother about in the data that we are talking about.
28:20
Speaker A
To summarize uh these components or the main estimation that we did, basically uh we started with the very basic uh MLE estimation. Using that, we uh used this MLE framework for the stochastic frontier analysis.
28:38
Speaker A
And here the main difference is that here we do have a composite error term.
28:42
Speaker A
So, that composite error term consist of uis and vis. Uis are the inefficiency that is one-sided distribution.
28:51
Speaker A
And VIs are the noise component, the normally distributed zero mean IID variables. And then now we have the joint distribution of the marginal density of f of e, basically the uh It will have the characteristics of both u and v. And our object One What we did,
29:11
Speaker A
we conceptualized our maximum likelihood function based on this f of e. Then uh we estimate our betas and sigma square.
29:21
Speaker A
And then one Once you have this thing, we can get a uh marginal density of f of e. And the moment we have marginal density of f of e, we can get the conditional value of f of u. And that f of u will give a value,
29:37
Speaker A
uh which is used for estimating technical efficiency, which is basically technical efficiency as exponential of minus UIs. Okay. So, this component I have taken mostly from Kumbhakar et al.
29:50
Speaker A
Kumbhakar and Lovell stochastic frontier analysis. Also, as I mentioned, Frontier is a very interesting package that all of you can explore.
29:57
Speaker A
It's basically a package based on Windows framework, which is basically a theoretically sound uh platform for estimating uh stochastic frontier analysis. But it has a limitation. It calculate uh half normal and truncated normal. The models like uh exponential and all are
30:17
Speaker A
yet to be explored in this framework. And the graphical representation that we discussed at some point can be seen in the Coelli et al. intro An introduction to efficiency and productivity and all.
30:28
Speaker A
Thank you. [music] [music]
Topics:Stochastic Frontier AnalysisMaximum Likelihood EstimationOrdinary Least SquaresEconometricsProduction AnalysisIID AssumptionsInefficiency ComponentLog-Likelihood FunctionPanel DataMATLAB

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries and speaker detection. 30 minutes free.

Or transcribe another YouTube video here →