Introduction to Data Envelopment Analysis (DEA) using the CCR model with Korean electricity data to measure technical efficiency nonparametrically.
Key Takeaways
- Nonparametric DEA models like CCR avoid assumptions about functional form and inefficiency distribution.
- CCR model efficiency scores range between 0 and 1, facilitating straightforward interpretation of relative performance.
- Maximum average productivity serves as the benchmark for technical efficiency in DEA.
- DEA can be applied using real-world data and simple MATLAB coding for practical efficiency analysis.
- The CCR model assumes constant returns to scale, which can be relaxed in more advanced DEA models.
What the video covers
- The video introduces nonparametric approaches to measuring technical efficiency, contrasting them with parametric methods reliant on functional form assumptions.
- Focuses on the Charnes, Cooper, and Rhodes (CCR) model, the foundational DEA model that uses linear programming without assuming a specific production function.
- Uses a simple one input, one output case for clarity, applying the model to Korean electricity utilities data.
- Explains the calculation of average productivity as output divided by input and discusses its limitations as a performance measure.
- Introduces the concept of maximum average productivity as the benchmark for technical efficiency.
- Demonstrates how to standardize average productivity scores to lie between zero and one, representing relative technical efficiency.
- Shows how to identify the Decision Making Unit (DMU) with the highest efficiency score (firm 12 in the dataset).
- Illustrates the use of MATLAB code to load data, calculate efficiency scores, and plot results.
- Mentions that the CCR model assumes constant returns to scale and previews future sessions on theoretical foundations and relaxing assumptions.
- Emphasizes the practical interpretation of DEA efficiency scores for policy and consulting purposes.
Chapters
- 00:00Introduction to parametric and nonparametric efficiency measurement
- 01:23Nonparametric approach and CCR model overview
- 02:35Simple one input one output case setup
- 03:40Loading and describing Korean electricity data
- 04:42Calculating average productivity and its limitations
- 05:50Identifying maximum average productivity and benchmarking
- 08:30Standardizing efficiency scores using CCR model
- 11:23Adding efficiency scores to data and plotting results
- 13:29Extending to multiple inputs and shadow prices
- 18:00Summary and preview of theoretical foundations and assumptions
Full Transcript — Download SRT & Markdown
Speaker A
Hi, welcome back to the course Applied Production Analysis using MATLAB. So, so far what we discussed is basically the parametric approaches for measuring technical efficiency.
Speaker A
It has its own advantages. It's very theoretically sound. It's basically building upon the neoclassical framework of production function and estimation thereof.
Speaker A
But it is being criticized basically from the fact that the moment you bring in a functional form and the estimation of technical efficiency based on that function that you are estimating, it is very sensitive to the functional form that you are assuming.
Speaker A
So now, as I mentioned in the introduction session, we take a slight deviation. We are going to estimate the production function without having any specific functional form, and instead of going for any estimation based on econometric modeling, we'll be
Speaker A
identifying the potential output of each observation by using a linear programming problem, and that approach, which is free from any functional form assumption or later any distributional assumption about the inefficiency term, is called the nonparametric approach.
Speaker A
So the very first nonparametric approach that you can see in the literature is basically the Charnes, Cooper, and Rhodes.
Speaker A
So Charnes, Cooper, and Rhodes is basically the very first model that you see in the case of nonparametric approach, and we keep calling them the Data Envelopment Analysis model, and for this session, we'll see the
Speaker A
very original model that you can conceptualize in the context of Data Envelopment Analysis. Here we call it the CCR model, the Charnes, Cooper, and Rhodes model. It is basically a simple model that anyone can conceptualize. We
Speaker A
don't need any understanding of linear programming problems. So here we are going to see this thing from a simple framework. In order to ensure a very smooth transition, we'll consider the one input one output case. And in
Speaker A
order to give a strong motivation, we'll be using a real-life data set. Basically, a data set coming from Korean electricity utilities. It is basically used in the Park and Desh paper that is basically efficiency of conventional fuel power plants in South
Speaker A
Korea. So we'll be using the same Korean electricity utilities data for our analysis, and in order to make it simple, we will use the one input one output case, and the objective of this session is to make you understand what this CCR
Speaker A
model does or what the general framework that we use in Data Envelopment Analysis follows for estimating technical efficiency.
Speaker A
So I load the data as data for I already loaded. I have already kept the Korean electricity data over here.
Speaker A
So we can define it as readtable, readtable. So it comes in the quotes since the data is already there in our drive or the current directory, it comes like this.
Speaker A
So this data consists of 30 observations from Korean electricity utilities. We do have an output variable, and in a real setup, we do have capacity, labor, and fuel as the input variables.
Speaker A
So in order to make a simple example, what I'm going to try is we are going to estimate the average productivity, that is basically the ratio of output to input by using the one input one output case. So here it is going to be average
Speaker A
productivity. We will be writing it as average productivity is basically data.output divided by each observation's individual elements divided. So we are putting a dot data.capacity.
Speaker A
This is the average productivity of capacity that you can conceptualize. So the straightforward question: can we use this average productivity as a measure of performance? Of course, we can do, but what is the problem? As a consultant or a policy person, you get this average
Speaker A
productivity scores and give these scores for the firms or the plants to have a meaningful interpretation.
Speaker A
One thing, this average productivity is sensitive to the unit that you are referring to. Here you can see some of them are coming less than one and some of them are coming greater than one. So it has a problem for a layman to interpret
Speaker A
it. Of course, we can see which plant is operating at the optimal scale. We call them the technically optimal scale of operation, but it does not give you a very straightforward interpretation in terms of performance. So say firm 7
Speaker A
is getting 1 as the average productivity and firm 2 is getting 8.782 as the average productivity. So we can get an idea, okay, firm 7 is doing better than firm 2, but it does not say how less each firm is performing
Speaker A
or underperforming in terms of their technically optimal production scale or technically optimal scale of operation.
Speaker A
In order to do that, we conceptualize a new terminology that is basically the maximum of average productivity. Max AP is basically conceptualized as the max of average productivity. So here you can see it has 1.494974 as the maximum
Speaker A
average productivity, currently here in the 12th observation at the moment. Okay. So firm 12 is getting 1.4374, that is the maximum average productivity.
Speaker A
So now we need to consider this as the ultimate outcome or the technically optimal scale, and now we need to see individual observations how they are deviating from this potential outcome or the potential average productivity. In order to do that, we can
Speaker A
we can here we can see the firm that already we saw by scrolling it up. So here we can find the firms with average productivity that is, say, I need firm at maximum average productivity, basically data
Speaker A
.firm in bracket, you can keep the row that as AP equal to equal to max of AP.
Speaker A
So this will tell you the firm 12 is the DMU with the maximum average productivity. So this code is basically firm at max is the new variable that you are creating. What we need to get is data.firm firm ID basically with the
Speaker A
condition that AP equal to maximum of AP. We'll revisit this firm at a later stage. So what this CCR does is it basically tries to get the measure between zero and one. Here what we can do is we can get a
Speaker A
standardized value that is basically the average productivity of each observation divided by the maximum average productivity. So with that, what happens is firm 12 will get value one, since that is the maximum. All other firms will get a value less than one unless they have an
Speaker A
average productivity equivalent to that of firm 12, and something above zero because it's a nonzero value that we are getting productivity. So with that, we get a series that we are going to see here which has a characteristic that lies between
Speaker A
zero and one. Also, it says how far we are deviating from the technically optimal production scale. So we can define T_CCR as a new variable that is basically the CCR technical efficiency, Charnes, Cooper, or model efficiency estimation, but in the
Speaker A
simple one input one output case. So it is basically AP of individual observation that has to be coming from the data.AP divided by max AP. So each observation will be divided by the maximum of AP data.
Speaker A
It is better we use the suggestion data.AP calculated here divided by max AP.
Speaker A
So here we have already calculated the average productivity, but we didn't attach that to the table. So that is the last step I skipped. So for that, to avoid that, we can data so that it gets attached
Speaker A
to the table. That was the error showing. So here you can see it has added to the table. So now we can use it as a code length.
Speaker A
So now you can see here the TE CCR that I calculated, as I claimed, all of them are coming less than one, and over and above the observation with highest average productivity in our data, that is Korean electricity utility data, you
Speaker A
get a value one for that, and all other observations are compared against this potential outcome or the technically optimal production scale.
Speaker A
So now we have a code to plot it. I'll just plot it using a bar. We are using a firm ID as the categorical variable and TCCR as the values. And then Y label we are putting TCCR and X label we are
Speaker A
keeping it as firm, and this is basically the technical efficiency CCR by firm that we calculated.
Speaker A
So this graph is being generated. You can see here. So here you can see as we saw firm 12 has the highest technical efficiency CC here which comes volume one. Of course, firm 13 also coming very close to that
Speaker A
followed by firm 11 or firm 10 and 11 and 10 and firm six has the lowest average productivity when you are considering it is a very hypothetical case we are considering we are considering only one input and uh one
Speaker A
output. So it might not be the true efficiency when you do for the advanced model but here we are considering one uh input case. So here what did this uh CCR model did? I'll try to generalize it into M input and
Speaker A
N output. So in the initial stage uh here we had only one input and one output. So here this ratio was very easy for us to calculate. Ideally when you have multiple inputs you should get a weightage such a way that that represent
Speaker A
the importance of each input uh in the production process. Generally we take the price as a vector for that. But in reality we'll not be getting price vector. So what we do ideally we should use the input prices.
Speaker A
Same manner. uh in the last example we discussed it basically had only one output. Sometime we may have multiple outputs and also we may have to account for the importance of these multiple outputs in the production framework. For
Speaker A
that that ideally we should have considered the uh output prices also in the production framework or estimation of technical efficiency. But in reality we'll not be able to get the input prices. So this m input that we are
Speaker A
conceptualizing basically that is basically a vector of m inputs. So I can write it as input uh price basically v_sub_1 t v_sub_2 t up to v mt and here uh the first subscript is basically for the corresponding input
Speaker A
that we are referring one up to m and the t subscript is basically for the uh the decision making unit that we are referring to. Also we should have an output price that is basically considered as UT which is basically U1
Speaker A
up to U n we have n uh outputs and each of them are for the individual observation T that we are referring. So it is basically a positive real value function but in real reality we will not be able to have an access to the input
Speaker A
output prices to see how important these each individual inputs or output to the individual firm that we are referring to. So instead of that what we use we use the shadow prices that is coming from the linear programming problem
Speaker A
point of view. So we use the shadow prices of inputs and output that plug into a ratio that is given uh for the estimation of average productivity. So average productivity of the teeth observation that T stands for the DMU
Speaker A
that we are referring to is estimated as a ratio of sum of R1 up to N output that we are having VRT YRT divided by sum of I 1 up to M UI XIT.
Speaker A
But how do you decide the shadow price? The model decides the shadow prices such a way that the moment you plug in these shadow prices into the formula for in their observation including t none of them should get an average productivity
Speaker A
greater than one that satisfy the condition for each values to decide the shadow prices and all shadow prices are expected to be greater than or equal to zero. So this is the formula or this is the framework that is being used in the
Speaker A
CCR model. And here the shadow prices are designed such a way that the moment you have a shadow particular shadow prices or the weightage that you are giving for individual inputs and output for endear observation. The moment you plug in
Speaker A
these values to the uh other observation or include in this observation for all observation, none of these observations should get a value greater than one or the average productivity that you estimated using the shadow prices should be less than or equal to one and all of
Speaker A
these shadow prices are greater than or equal to zero. So this is how the CCR model being conceptualized. It looks a bit technical but the moment uh we change this ratio function into a linearized form by taking u some modification that we'll be going
Speaker A
to do in the upcoming sessions. It can be converted into a linear programming problem and the primal of the problem may look very complicated but you can convert that into LDL and we'll be using that linear programming problem to
Speaker A
identify the shadow prices and the potential outputs of each or potential input or the actual minimal input the individual observation could have used in the production process which will be used for identifying the uh potential outcome and the deviation from that
Speaker A
potential outcome is considered as the uh inefficiency or that gives you a measure of technical efficiency of each individual observation. So just to summarize this is basically the starting of the uh our nonparametric model what we started we started with a actual data
Speaker A
set which is basically Korean electricity utility data we consider one input one output case that were we consider output and capacity as a value we calculated the average productivity of endear observation we attached that to this thing here by
Speaker A
attaching for the attaching we use data as a um starting point for that. Then uh these values of course it can be a measure of performance but it has a problem that it lies not in a standardized manner. In order to
Speaker A
standardize it what we did we took the maximum average productivity and divided the endear observation with the uh that maximum value. So we get a series that lies between 0 and one. One representing the maximum average productivity observation or the maximum technical
Speaker A
efficiency in this context. And what we did here basically the basic CCR model which we plotted in a graph and we saw which observation is coming uh inefficient and efficient. And this is basically the formula for estimating efficiency using the CCR framework. It
Speaker A
involves a shadow price vector of inputs and output. And we decide the shadow price of input and output such a way that the moment you plug in these shadow prices, we try to maximize this value.
Speaker A
But such a way that the moment you try to plug in the shadow prices into the u the ratio that we represented over here, none of the observation should get a value greater than or equal to one. And
Speaker A
this is the basic CCR model. And in the upcoming session we see the very uh theoretical foundation of data development analysis. Basically this model is purely based on the uh C CS assumption constant return to circuit assumption and we relax this assumptions
Speaker A
and see the upcoming models and see how it being implemented in the empirical setup. Thank you.
Topics:Data Envelopment AnalysisDEACharnes Cooper RhodesCCR modeltechnical efficiencynonparametric approachproduction functionKorean electricity utilitiesMATLABefficiency measurement











