Arne Duncan’s position on the use of value-added test scores seems to have changed in the last two years. In August, 2010, he seems to be saying that test score gains are a measure of teacher effectiveness. In May, 2011, he seems to want to use value-added measures as part of teacher evaluation, and in February, 2012 he seems to be against the whole idea.
August, 2010: Effectiveness = gains on tests (value-added)?
“U.S. Secretary of Education Arne Duncan said Monday that parents have a right to know if their children's teachers are effective, endorsing the public release of information about how well individual teachers fare at raising their students' test scores.”
May, 2011: Use both observations and value-added measures.
“Together with you (teachers), I want to develop a system of evaluation that draws on meaningful observations and input from your peers, as well as a sophisticated assessment that measures individual student growth, creativity, and critical thinking. States, with the help of teachers, are now developing better assessments so you will have useful information to guide instruction and show the positive impact you are having on our children.”
Feb, 2012: Don’t use value-added measures?
“Teacher evaluation should never, ever be based on test scores.”
A web-based destination for aggregated news and commentary related to public school education in Kentucky and related topics.
Wednesday, February 08, 2012
Is Duncan Moving the Target
Friday, November 19, 2010
The Safe Handling of Value-Added
As with all fire arms, safe handling is essential. Poorly designed and in the wrong hands, value-added assessment is a shotgun. It may hit what its aiming at, but its likely to hit some other things too. But used within strict limits, we may be reasonably assured of it usefulness.
Like writing portfolios, value-added assessment does a pretty good job at the extremes. In the same way that portfolios could separate the distinguished writer from the novice while poorly differentiating performance among the apprentices, value-added can identify the great teachers and the terrible teachers but can't tell high average from low avarage. It should not be relied upon for important decisions among those scoring in the average range.
A balanced look at evaluating teachers through value-added assessment was offered by the Brookings Institute in a new report this week.
As a principal, I always wanted as much data about the performance of the school as I could get. The presence of reliable data is very useful for improving the quality of the decisions that must be made. I wanted accurate information about everything from monthly budget reports to the time individual buses were arriving at the school. But I did not want to beat anyone over the head with the information. I wanted to solve problems and make the school better.
Similarly, I wanted as much information as I could get about the performance of our students and teachers. I wanted to make sure we were giving our students everything they needed to be successful. Thus, I am naturally attracted to interim testing data and the concept of value-added assessment. But to be of use, all data counted upon for high-stakes decisions must be reliable. Otherwise, it stands a good chance of becoming counterproductive.
How value-added systems are constructed and how they are used matters.
The Brookings folks argue that the use of value-added data to predict future teacher performance is consistent with predictive data used in other fields.
Over at Education Week's Teacher Beat blog Stephen Sawchuk sees the findings as adding a contrasting view in the debate over value-added assessments, which have been criticized as an incomplete way to evaluate teachers.
While an imperfect measure of teacher effectiveness, the correlation of year-to-year value-added estimates of teacher effectiveness is similar to predictive measures for informing high-stakes decisions in other fields, like the SAT test, Brookings says.
The evaluation of teachers based on the contribution they make to the learning of their students, value-added, is an increasingly popular but controversial education reform policy. The report attempts to clarify four areas of confusion about value-added.
The first is between value-added information and the uses to which it can be put. One can, for example, be in favor of an evaluation system that includes value-added information without endorsing the release to the public of value-added data on individual teachers.
The second is between the consequences for teachers vs. those for students of classifying and misclassifying teachers as effective or ineffective — the interests of students are not always perfectly congruent with those of teachers.
The third is between the reliability of value-added measures of teacher performance and the standards for evaluations in other fields — value-added scores for individual teachers turn out to be about as reliable as performance assessments used elsewhere for high stakes decisions.
The fourth is between the reliability of teacher evaluation systems that include value-added vs. those that do not — ignoring value-added typically lowers the reliability of personnel decisions about teachers.
We conclude that value-added data has an important role to play in teacher evaluation systems, but that there is much to be learned about how best to use value-added information in human resource decisions...
Critics of value-added methods have raised concerns about the statistical validity, reliability, and corruptibility of value-added measures. We believe the correct response to these concerns is to improve value-added measures continually and to use them wisely, not to discard or ignore the data...
Saturday, February 06, 2010
How Value Added looks in Ohio
Twelve central Ohio schools are among the worst 5 percent statewide.
Their academic struggles mean they are eligible to receive federal money to help them transform or start over. A list of these schools was released yesterday by the Ohio Department of Education.
Six Columbus City Schools buildings are on the list of the worst-off, as are four in Cleveland and 16 in Cincinnati. Several charter schools -- six of them in central Ohio -- also made the "top" rung on the list.
"No one is going to like the fact that they're on this list," said Mark Real, who heads the Columbus-based nonprofit KidsOhio, which studies education issues. He's been monitoring stimulus-related spending and improvement programs. "But this is not just a 'label and leave it' approach. These schools are in for some pretty intensive care."
KidsOhio.org’s new analysis of state education data shows that many of the state’s lowest-rated schools, both Ohio 8 schools and charter schools, rise to near the middle among schools statewide when ranked according to the state’s own “value-added” measure of annual educational progress.

KidsOhio.org ranked Ohio’s public schools - both traditional district schools and charter schools - according to their Performance Index scores (measuring student achievement on state tests in a given year) and their Value-Added Gain scores (the state’s measure of students’ academic progress from year to year). The analysis covered the 2,688 Ohio schools that received scores for both measures in 2008, and the data revealed that:
Ohio 8 district schools ranked, on average, 2,199 out of 2,688 schools on test scores, but jumped 663 places to an average ranking of 1,536 on student progress. Charter schools taken as a group had an average ranking of 2,288 on test scores, but an average ranking of 1,362 on student progress, an increase of 926 places.The analysis also identifies further common ground between traditional district schools in the Ohio 8 (Akron, Canton, Cincinnati, Cleveland, Columbus, Dayton, Toledo and Youngstown) and public charter schools: both groups of schools serve high percentages of economically disadvantaged students and students with special needs.In addition, KidsOhio.org ranked all of Ohio’s 610 traditional school districts on the two measures. Each of the Ohio 8 districts ranked higher on educational progress than on absolute test scores. Graphics showing each of the Ohio 8 districts’ ranks on the two measures are available here. You can download KidsOhio.org’s full report on the analysis here. For data on all 2,688 ranked schools, click here; and for data on all 610 school districts, click here.
Saturday, November 07, 2009
Build a Value-Added Assessment System for P-12 Before Moving on to Teacher Prep Institutions
Recently, there has been increased talk of holding teacher preparation institutions accountable for the performance of Kentucky teachers - in effect holding colleges accountable for student achievement in the state's public schools.
It's a good motivation, but in practice, there's a lot wrong with the idea.
It simply takes all of the problems associated with teacher-accountability-by-standardized-test-score, a central tenant of the Obama Administration, and multiplies that unfairness many times over. While it may feel good to teachers to know that they are not the only ones tied to an unfair system, it does not fix what's broken.
But it may become a better idea, if those who are enthusiastic for change slow down long enough to create a system that addresses fairness by trying to address the many technical problems, and that might actually work to an acceptable degree. That means building Kentucky's new accountability system from the bottom up:
- Curriculum standards, first
- Assessments built on those standards
- Lessons taught on those standards (in that order)
- Assessment results fed into a value-added accountability system that is sensitive to the great variability in children and one that establishes an individual baseline for each child and controls (to the degree possible) for that demographic variability.
- At this point the public should consider a cost benefit ratio: the investment, in relation to the reliability and validity of the system.
- Professionals should gauge the practical limitations of social science research, both quantitative and qualitative. Unlike the natural sciences, our variables refuse to hold still; which bears heavily on the precision of the system.
- Individual student progress is measured from each individual student's established baseline
- Individual student achievement data is collected over the student's entire academic career
- The teacher accountability system should be quantitative and qualitative.
- The principal accountability system should be quantitative and qualitative
- There should be a planned review of the accountability system after collecting about three years of data (barring some unforeseen data catastrophe), with any significant adjustments to the accountability formula made at that time.
- Then, and only then, a teacher preparation institute accountability system should be ready to track the performance of teachers who graduated from each institute, based on that value-added system; and that system should be quantitative and qualitative
The Century Foundation recently outlined their Eight Reasons Not to Tie Teacher Pay to Standardized Test Results. In a nutshell,
Reason #1: Tying test scores to teacher compensation suggests that teachers are holding back on using their experience, expertise, and time because they are not being paid for the extra effort.
Reason # 2: The standardized tests in most states are lousy and so are the standards they are designed to measure.
Reason #3: The idea of compensating teachers individually in order to differentiate their performance from their school colleagues defeats a principal tenet of good instruction—that teachers need to learn from one another to solve difficult pedagogical challenges.
Reason #4: Most teachers do not teach a grade or subject that is subject to standardized testing.
Reason # 5: Even reliable standardized tests are valid only when they are used for their intended purposes.
Reason #6: A key assumption of using test scores to judge teachers is that students are randomly assigned, first, to schools, and, second, to classes. Neither is true.
Reason #7: State data systems are in their infancy. It turns out that it is harder, is more expensive, and takes longer for states to produce reliable, accurate, and secure longitudinal data on students and teachers than widely assumed.
Reason #8: The rationale for tying tests to compensation is not clear.
The non-profit, non-partisan Century Foundation argues that No Child Left Behind has narrowed instruction too much already, that one does not need a standardized test to identify the worst and best teachers, and no system could be constructed with sufficient precision to withstand the inevitable court challenges.
At the heart of the argument in favor of tying pay to test scores is the idea that it will improve practice. But that can only work if the economy provides the anticipated financial incentives. In this recession,
"if teacher compensation does not keep up with inflation because of poor student performance, then teachers will . . . what? Work harder? Dig deeper? Stay longer? There is no evidence that such measures improve instructional practices or student outcomes."
Secretary Duncan is correct when he catalogues the weaknesses in the present system of preparing, recruiting, mentoring, retaining, inspiring, retraining, promoting, and dismissing teachers. but this is an idea that is way ahead of just about everything it would need to have even a chance of working fairly and reliably, if at all.
Sunday, March 29, 2009
Massachusetts Goes to a Value-Added Assessment
Mitchell Chester continues to innovate in Massachusetts. This week he introduced a type of value-added assessment system - the kind Kentucky should now be considering.Oh by the way... he's the guy the Kentucky Board of Education should have hired instead of flirting with Barbara Erwin.
This from the Boston Globe:
MALDEN—Each fall when the state releases MCAS scores, principals often blame a dip in scores on the students, tactfully arguing that the class in question was perhaps not as superb a group of mathematicians or voracious readers as their predecessors.
State education leaders plan to inject a reality check this fall into the "good class
vs bad class" debate by tracking the performance of individual students as they advance from one grade to the next. The new measurement could shed light on who is falling short -- teacher or pupil -- and lead to fundamental changes in the way students are taught.Mitchell D. Chester, the state commissioner of elementary and secondary education, said yesterday that the new analysis will make it harder for local school leaders to be dismissive of poor test scores.
"It takes away a lot of the excuse-making," Chester said at yesterday's meeting of the state Board of Elementary and Secondary Education where the new system was unveiled.
Under the current system, the state judges a school's success by comparing its MCAS scores at each particular grade level to the scores posted by that grade the year before. The English and math MCAS tests are given in grades 3 through 8 and in grade 10.
Many teachers and adminstrators have chastized the approach as an apples-and-oranges comparison because the variation could simply reflect a class of particularly gifted or challenged students. The problem can be especially acute at small schools, where there are only a few dozen students at each grade level and the performance of a handful of students can create dramatic shifts.
Using the new tool, the state will augment that analysis by examining the performance of individual students or classes of students over the period of several years, starting in the third-grade.
The examination of current and past scores will allow them to predict students' likelihood for improvement in the future and assess whether they are on track to meet expectations. If a number of a students at a particular school are exceeding the statistical predictions, that could indicate that the administrators and teachers there have identified promising teaching methods. If a number of students are falling short of predictions, that would indicate there could be a problem...
Wednesday, October 01, 2008
Salvaging Accountability
George W. Bush rode to the White House pledging high standards for all students. He’ll leave Washington with the nation’s public education system focused on teaching basic skills to disadvantaged student populations, with the United States lagging in international comparisons of educational attainment, and with his signature education law plagued by so many problems and mired in so much controversy that it has put at serious risk two decades of work to improve public schooling by making educators accountable for their students’ success.
The most important thing Barack Obama or John McCain could do quickly to salvage the accountability movement is change the way that the federal No Child Left Behind Act judges schools. Not by abandoning NCLB’s focus on students’ meeting standards, a move that would be unwise on both policy and political grounds, but by making the law a more legitimate report card of school performance, one that provides a fair and accurate gauge of educators’ contribution to their students’ achievement. Since its inception, NCLB has instead held schools responsible for factors they can’t control and perversely encouraged states to set standards low....
Friday, March 14, 2008
The value-added idea
Now some folks won't like this idea. Those who feel it's wrong to suggest out loud that some students may never achieve proficiency (due to poverty, neglect, abuse...) may cringe at the notion. But a school's "success" under the present system is impacted by the students who attend. The trick is to fairly adjust for the differences in student population and measure what the school brings to the students; the value the school adds.
Is it OK for a "good school" to cruise along while gains are made by talented students in spite of their teachers?
Is it OK to sanction a low performing school that is outperforming expectation?
With those ideas in mind, the folks at the Center for Educational Research in Appalachia recently conducted a regression analysis to determine the strength of the relationship between poverty and the 2007 district-level CATS index. Here's what they came up with:
The difference allows a peek at those districts that are performing as expected - as well as those districts that are "over-performing" or "under-performing" relative to the socio-economic status of the student population.
Finally the districts were ranked by their differences into five equal quintiles of 35 school districts each - and placed on a color coded map of Kentucky.
Districts with differences of -1.7 to + .5 were considered to be performing about as predicted and are colored yellow. Fully acceptable; not great.
Differences from +.5 to +5.3 are light green. Very Good.
Districts exceeding their predicted score by + 5.4 to a whopping + 16.3 are dark green. Terrific.
On the other hand, districts that failed to meet their predicted scores are orange (-.5 to -1.7) and Red (-5.0 to -14.0). These represent the coasters and low performers.
The MacDaddys of the process are those high-performing districts that also exceed their predicted outcomes (The heavily supported Ft Thomas, Anchorage...)
The value-added idea would expand this assessment to include a number of other factors that would be tracked in a stable system over time. Dropouts? Graduation rate? Attendance? ...whatever Kentuckians will accept as reasonable measures of school performance.
A New Approach to Accountability in Kentucky
If I had a magic wand, testing would exist to inform better instruction and provide parents with annual information about the progress of their child. But accountability would look something like this:
'Effectiveness index' considers demographics in measuring campuses
How good is your kid's school?
It seems like a simple question, but factors like student backgrounds can make clear-cut answers elusive.
What if you could find a way to evaluate schools that doesn't penalize them for their students' language barriers, lack of parental involvement or other social factors that impact learning?With virtually no fanfare, the Dallas Independent School District does exactly that. Every year, it produces a "School Effectiveness Index," a rating that levels the playing field between schools, no matter where their students come from or what they lack in life.
DISD has been calculating the scores since 1992 but has never publicized the ratings, even though parents are hungry for information showing how their neighborhood schools stack up. The Dallas Morning News recently collected nine years of effectiveness scores from the district and has made them available at dallasnews.com/disdblog.
One of the district's top researchers, who helped develop the ratings, cautioned parents not to view the scores in isolation. Many other factors also are important, he said, such as personal interactions with teachers.
"From a parent standpoint, I would be far less concerned about the school [ratings] than I would be about what my kids are telling me about their teacher," said Robert Mendro, the district's director of evaluation. "A kid will tell them about which teacher requires homework, which teacher is tough."
The News recently shared the effectiveness scores with about 40 parents, educators and community members. Some found the scores helpful in a general sense.
"Rather than choosing the highest-ranking school in a measure of best, I would look for consistency in the top 10 percent of performing schools," said parent Randy Hazlett, who has a child at Townview Magnet Center. "Still, the rating system will reward those schools which accept students from low-performing feeder schools and significantly boost their test scores."
The scores turn conventional wisdom about which schools are "good" on its head. For example, the most effective school in the district last year, according to the ratings, was Rusk Middle School. By way of comparison, the Talented and Gifted magnet school – widely lauded as one of the best public schools in America – ranked No. 29.
That happens because the ratings compare schools by isolating the impact teachers have on student achievement. The ratings measure how far a student advances with his teacher in one year.
Through a complex statistical analysis, the district isolates student characteristics that affect learning, such as family income, ethnicity and English mastery. A school's score is designed to eliminate advantages campuses gain from the social and demographical characteristics of their students.
That's a completely different approach from the state's accountability system, which looks at how many students pass annual standardized exams. Low-income students tend to score lower on such tests.The district's "value-added" ratings focus on how much students learn in a year, not whether they pass certain tests, such as the TAKS, Dr. Mendro said.
Such calculations are a growing trend in education as researchers try to determine how far schools push their kids along the learning curve...
Sunday, August 05, 2007
What Testing Guru Bill Sanders Really Meant About Multiple Measures

Tennessee is the state most strongly identified with value-added assessment. Its system dates back to 1992, when value-added was implemented as an integral part of a comprehensive education reform measure. Using a complex statistical method developed by Dr. William Sanders, then a statistician at the University of Tennessee , the Tennessee Value-Added Assessment System (TVAAS) provides:
Data to the public on the performance of districts and schools, and data for appropriate administrators on the performance of teachers;
Information to teachers, parents and the public on how schools are doing in helping each child make academic gains each year;
Information to school administrators to help identify weaknesses in even the strongest schools.
TVAAS is a statistical methodology that begins with testing each student in each grade in a number of subjects. Through 1997, Tennessee tested second through eighth grades in Reading , Math, Language, Science, and Social Studies. Tennessee began testing grades three through eight in 1998.
The TVAAS statistical model aggregates student growth increases using a design that accommodates missing data. Because of a philosophical belief that schools should insure that all students progress at equivalent rates, no matter their disadvantages, the model does not include other data on students.
In a report by the Council of Chief State School Officers, Tennessee's 8% increase in math and science scores was linked to TVAAS. In addition, Tennessee is one of the few states that have shown improvement on the National Assessment of Education Progress since TVAAS was implemented in 1992.
The state has both rewards, aid, and sanctions linked to its school rating system.There is no specific value-added teacher evaluation as part of Tennessee 's accountability system, but school administrators have access to teacher level data that can be used to improve instruction. Value-added scores can be used for up to 8 percent of a teacher's evaluation.
The incentive funds are only available to schools, not to teachers.
Although teachers and administrators were suspicious at first, they are now finding they can actually use TVAAS to improve teaching, something that no other accountability system has afforded.
Friday, April 20, 2007
Value Added Assessment
From Tennessee a new website giving parents value-added information about schools in the state. Could raise some interesting questions...This scatterplot shows 2006 data on school performance and poverty in Tennesee. It's all over the place suggesting that those accountability folks are onto something...