Showing posts with label Skip Kifer. Show all posts
Showing posts with label Skip Kifer. Show all posts

Monday, November 07, 2011

Emphasis on testing makes less sense than scores


The hottest temperature ever recorded in the United States was 52 degrees in Las Vegas, Nev., according to Wikipedia. Prospect Creek, Alaska, recorded the coldest temperature of minus 63 degrees. I took a college entrance examination first in the warmest city and scored 500. In the coldest city, I scored 20, a whopping loss of 480 points.

The Herald-Leader's Sept. 29 editorial, cited some statistics about differences in test scores between elementary and high school:

"In math, the drop from elementary to high school was a whopping 27 points (from 73 percent to 46 percent proficient or higher) and in science 30 points (71 percent to 41 percent)."

The editorial continued: "There may be a reasonable explanation, but it appears that students are losing a lot of ground in math and science once they reach high school."

Yes, there may be. The Herald-Leader's numbers may be little better than the ones I cited above. The issue is the scale upon which the numbers are based. The 52 degrees is temperature on the Celsius scale; the comparable value on the Fahrenheit scale is 125 degrees. The first test was the SAT; the second was the ACT.

The Herald-Leader's differences — which should be expressed in percents, not points — are from scores on different scales. There has been no attempt to vertically equate the tests (fancy words for making tests scores comparable). While there are straightforward rules for translating a temperature in Celsius to Fahrenheit, the process is much more complicated for test scores. Without equating the scales and making sure that cut-points to determine proficiency are comparable, comparisons like those in the editorial are not meaningful.

With the increased use of tests, one would expect increased sophistication about the proper use of scores and how to interpret results. That is not the case. At every level, one gets ridiculous test score use and interpretations.

No Child Left Behind used test scores in reading and mathematics to make judgments about schools.

Kentucky's Council on Postsecondary Education requires placement in courses based on one test score.

Very short tests are given to eighth graders in Kentucky and interpreted to mean a student is or is not "on track" to go to college.

Fayette County schools use a short test to track a sixth-grader into a less rigorous mathematics course.
Perhaps school officials should look more closely at international test results. None of the high-scoring countries track as early or as much as does the United States.

A friend of mine asked me where I thought this recent, silly, unquestioned emphasis on testing came from. I don't know.

But it was not so long ago that a major purpose of testing was "in addition to," not instead of. A score on a college entrance examination was used in addition to the high school record for admission to college. A score on a placement test was used in addition to previous work in the subject matter to make placement decisions. A standardized test was given to get additional information about how students were progressing.


So the real problem with the editorial is not that it misuses test results and misinterprets test score differences. That's the coin of the realm. The problem is that it has been captured by this inexplicable infatuation with testing. Most fads in education fade away. Let's hope.

Monday, June 20, 2011

Accountability Testing Undermines Teacher Authority, Narrows Curriculum

By Skip Kifer

Stephen Jay Gould esteemed paleontologist, MacArthur Foundation prizewinner, and award-winning author, dedicated his book "The Panda's Thumb" to three of his elementary school teachers. He was particularly grateful to his fifth grade teacher who recognized and accommodated "youngsters with a developing passion for science." Ms. Ponti, who provided materials and books, also set aside time each week for the group to sit in the back of the room and talk about science.

Fayette Advocates for Balance in the Classroom (FayetteABC), a group of parents concerned about test-driven instruction in our schools, might rightly wonder whether something comparable to Gould's experience could happen here.

Those who manage the local school system view testing through an accountability lens, apparently believing only higher test scores mean better schools. FayetteABC notices, on the other hand, more time being spent on testing and test preparation and therefore less time and fewer opportunities to respond to the legitimate interests of children.

They are right. Research shows that an emphasis on accountability testing undermines the authority of teachers, leads to students dropping out of school earlier, and narrows the curriculum (what is tested is what is taught). In addition, studies of results on comparable assessments do not replicate score gains found on Kentucky's assessments; that is, students learn to take a particular test but do not master the instructional material upon which the test is based. And, perhaps even more problematic, because of the imprecision of educational measurements, tens of thousands of Kentucky kids each year when given the labels of Novice, Apprentice, Proficient, or Distinguished are given incorrect ones; e.g., students are labeled proficient when they are not or are not labeled proficient when they are.

It is good that FayetteABC has sounded the alarm about these excesses because even more accountability testing is forthcoming. At the Federal level, the re-authorization of the Elementary and Secondary School Act (No Child Left Behind) will surely contain more testing because of a misguided attempt to use test scores to measure teacher performance. At the state level Senate Bill 1 is, at best, just more testing. In addition, Kentucky is part of both consortia that will build new assessments for what are called the Common Core Standards. Those tests are expected to do so many things that inevitably there will be just more testing. Finally, Fayette county schools, beginning in kindergarten, do still more of their own testing and preparing for tests, over and above what is done to meet federal and state requirements.

Informed citizens in a democratic society have an obligation to make schools their own. More than anything FayetteABC wants parents, teachers, and administrators to engage in a conversation about effects of test-driven instruction upon Fayette County children. Although state and federal laws mandate testing, they do not mandate the effect that this testing currently has on classroom instruction. FayetteABC presented their concerns to the Fayette County Board of Education, asking that body to endorse a more balanced approach to instruction in our schools, particularly in the current search for a new superintendent.

FayetteABC has an online petition signed by nearly 400 county residents from at least 14 different schools and 16 zip codes that continues to grow, speaking to the scope of their concerns in Fayette County. They have produced an informative website that I would encourage interested and concerned parents to peruse:
http://fayetteabc.web.officelive.com/.

Almost anyone you talk to has a favorite teacher who greatly influenced him or her. In all of the years I have heard persons talk about that teacher, the testimonials were of the Gould variety. I have never heard someone say his or her most influential teacher increased a test score.

Wednesday, May 25, 2011

FayetteABC Makes It's Case

On Monday, Dr Erik Myrup, representing FayetteABC spent about 12 minutes addressing the Fayette County Board of Education. Myrup discussed the unintended, negative consequences of standardized test preparation on children in Fayette County Public Schools. The board listened without comment or questions.



In his comments, Myrup quoted KSN&C's Skip Kifer on the issue of standardized testing.

Wednesday, March 09, 2011

Measuring the Gap under SB 1

The new SB1 test got a first reading before the Board of Education in February and is scheduled for a second reading in April. We are in a 60-day window set aside for public comment, so let's talk about it.

One of the issues that has worried me is, How should the achievement gap be measured?

This is important because we know that however the state calculates it, teachers will plan their strategies based on whatever focus is likely to provide the best test score results. One approach might induce teachers to focus on a small subset of the kids who are said to be “in the gap.” A different approach would broaden that focus.

KSN&C spoke to KDE testing guy Ken Draut, at the AdvanceEd Conference in December, and learned of KDE’s plans to use ACT benchmarks to measure the gap. Uh oh.

Generally, KDE is looking at what they call a balanced accountability approach.

Achievement score data would come from five days of testing with the new SB1 Test for grades 3-8 built around the new Kentucky Core Achievement Standards, as they are implemented. Scores will be calculated around “proficiency,” meaning that cut scores will be set to determined performance levels. Novice = 0; Apprentice = 0.5; Proficient = 1.0; and Distinguished = 1.5. The + .5 amount given to students scoring in the Distinguished range is thought of as a Bonus, and it will be offset by a negative .5 for each Novice student in the group.

But what about measuring the achievement gap?

Draut told conference attendees that “we really feel like, in the gap world, we got a really innovative model to measure gap.” Draut explained that KDE had three problems with measuring the gap which they have tried to address.

  1. The number of different sub groups, which can number to as many as 45 different goals.
  2. A lot of the kids fall into more than one group. For example, 80 percent of Kentucky’s African American kids are receiving free and reduced lunch. 80 percent of our ELL kids are in the free and reduced lunch group. 70 percent of our special education students receive free and reduced lunch. As a result, we end up counting one student multiple times. So under the present system, a single student might be counted four times. Miss one target and you are likely to miss four.
  3. Comparing “closed gap to group.” Draut said, “You want to close African American to White; American Indian to white… Because of the way testing works, you can end up with a “wavy” pattern.” For example, in JCPS one year, the white kids at Southern Middle School dropped backwards and the African American kids stayed the same, the Courier-Journal reported that Southern Middle was closing the gap. And the opposite can happen (which was our experience at Cassidy) where white kids can go up 8 points and African American kids go up 6 points, but the gap increases.

To address this, KDE plans to present all of the gap students’ data, but will create a new single group of underperforming “gap kids” who would only be counted one time. Instead of comparing the gap kids to the group, it would be measured against the goal of 100 percent proficiency, or what is called “gap to goal.” The gap is to be divided by the number of years schools are given to reach their goal and schools would be awarded points based on the percentage of that goal they were able to close. A school that had a six goals to meet and closed three of them would earn 50 percent of their points.

Growth is to be measured by using a regression of the reading and mathematics scores, the only tests given every year from grades 3-8. It will compare a student’s progress to other students who have been performing similarly. Given a proficient 5th grade student with a scale score of 230, who then earns a score of 240 in 6th grade: the model asks if this level of growth is typical of other Kentucky students, above average, or below average, and awards points accordingly.

KSN&C caught up after his presentation.

KSN&C: Ken, as you may be aware, a number of statisticians…like Skip Kifer, say that the modeling that underlie the statistics of the ACT Benchmarks are a bunch of crap, basically. [chuckles]

Draut: Right.

KSN&C: Are you concerned about that?

Draut: Well, this is how we answered the board the other day: It’s you guys. You guys drive this. If you say the ACT is a bunch of crap, lets’ throw it out…

KSN&C: Well, not the ACT. Just the benckmarks.

Draut: Well, I’m just saying, if you all say it, and then you put something else in, we’ll line right up, because we’re trying to get them ready for you.

KSN&C: OK, but you lost me. Tell me who “you” is. Because you’re saying the board…

Draut: Universities.

KSN&C: Oh.

Draut: You see, we’re driven by the universities. We can’t get our kids into the universities unless we meet your criteria.

KSN&C: So if the universities say, this standard isn’t appropriate, or the metric’s wrong, or something, then that’s going to be a problem for you guys.

Draut: Well, we’ll put in whatever you say, but I tell you, what the issue is, and we’ve said this to several people, tell us what you’d replace it with.

KSN&C: Uh huh.

Draut: Just tell us.

KSN&C: So, the benchmarks are useful, because they are there…But you have to know what they mean or they’re meaningless. And you can’t replace it with the ACT really, because that cuts out middle school and…causes you some other problems.

Draut: Right.

KSN&C: So, then what do I replace it with. I’ve got a bad yardstick, but it’s the best one I’ve got?

Draut: And what are the universities going to accept to get the kids in the door? Because whatever the universities accept, that’s what I’ve got to get my kids ready for.

KSN&C: Are you getting that kind of pushback from the universities?

Draut: No

KSN&C: So the question’s been raised but nobody’s pushing the issue?

Draut: No. It’s kinda like just what you said, tell me what’s in its place?

KSN&C: And nothing comes to mind.

Draut: So now you open up fifty years of research saying, hey, we can tell you it works. It does predict…

KSN&C: Do we know the degree to which those benchmarks are bad, or in what direction they are bad? Or is it that we just don’t know?

Draut: I think that you’d have to do some reading, both the pro and con, when I read, and I’ve heard Skip, but when I read the ___of it, it makes a lot of sense. And when I hear Skip it makes sense, too. I can’t get a sense of which one’s right…But that whole issue is driven by CPE and the universities because if you’re sitting there in the university saying we’re only going to take the kids that make the CPE benchmark, and we’re only going to take the COMPASS, then we say, OK, and we line up with you. But if universities change…and say, you know, we’re not going to use ACT, we’re going to use some new testing, then we’ll realign everything [to that]. ..But I think it would be useful to look at both the pro and the con.

For a few months now, I've been pondering Draut's position that decisions made at CPE should drive the model ultimately adopted by the Kentucky Board of Education. Generally I agree that we can not lower standards and KDE must hit college-ready targets. But I'm much less convinced that CPE ought to dictate how the achievement gap in measured in our elementary and middle schools.

NOTE: It is my understanding that the EXPLORE can predict results on the PLAN test, but not the ACT. The PLAN test can predict performance on the ACT but not performance in college. The ACT can predict performance in college up to a point, and its arguably not the best way, but is made better by the inclusion of other measures.

Which Gap to Close

By Skip Kifer

Both No Child Left Behind (NCLB) and Kentucky's Senate Bill 1 (SB1) refer to achievement gaps and include expectations for closing them. As defined by Kentucky's Senate Bill 168 (SB168) and NCLB, gaps are differences in test scores based on gender, disabilities, limited English proficiency, ethnicity, or socio-economic status. Each school in the Commonwealth - given a set of rules about what is a gap - is expected to minimize differences between those groups, thereby closing it. It is implied that one expects each student's score to increase but those with lower scores are expected to increase by greater amounts. There is a desire for overall improvement in test scores as well as improvement in closing a gap.

As I write, there has been no reauthorization of the Elementary and Secondary School Act (NCLB) so there is no decision about how to define a gap or how to decide whether a gap is closing. To my knowledge, no decision has been made about how the implementation of SB1 will deal with those issues, either. I expect, for reasons discussed below, NCLB definitions to change. My guess is that Kentucky's might also change.

NCLB now defines an achievement gap as a difference between the percent proficient for one group, say girls, versus that of another, say boys. For a school to close the gap, it must reduce the differences between the two percentages. For SB168, Kentucky initially used a complicated average difference to look at closing the gap. That is, a school closed an achievement gap when it reduced, by a set amount, the weighted average between groups of achievement differences across grade levels and content areas.

In what follows, I hope to point out the strengths and weaknesses of both approaches and then suggest a third alternative for consideration. It is so easy for one to mouth the words "closing achievement gaps" without being aware of the technical difficulties of defining the gap and knowing either when it exists or when it has been closed. As a way to discuss the issues, I created data[1] and drew pictures of them.

Figure 1. Six representations of an achievement gap.

Figure 1 contains six pictures of the data. The graphs depict comparisons for one grade level and content area; for example, fourth grade reading. Three pictures (A,C,E) on the left are ways to show shapes, centers and spreads of the data. Three pictures on the right (B,D,F) are ways to show the gaps across levels of the test scores. Pictures A&C and B&D are the same but will be used to describe different features of the data.

Centers, Shapes, Spreads - Averages as Gaps

Figures 1A and 1C compare two groups, one of which is four times larger than the other. The size difference could happen if, for example, one was comparing majority students to minority students. Such differences in size do not affect the ensuing discussion. The groups could be of equal size, too. These dotplots are just detailed histograms that better represent the shapes and spreads of the distributions. A reader should see several things in Figure 1A: the distributions overlap substantially, the shapes are rather similar; the spreads are similar; but, the centers are different. The bottom distribution is shifted to the left indicating lower average performance for Group 2. That average difference could be a measure of the "achievement gap."

Figure 1B is another way to describe the data. This is a particularly good way to view cut-points that are used as the percent proficient goals. The lines I added to the figure are guides to interpreting the data. These curves depict what parts of a score group are at or below certain values. For example, if one follows the lines, one can see that fifty percent of Group 2 students score at or below 35. The comparable number is 40 for Group 1, the higher scoring group. The differences in those percents is the measure of the gap when the cut-point is 40 (i.e., 40 represents the goal, the desired percent, the percent proficient). One's eye can see different achievement gaps as the curves move from about 10 to 70.

Percent Proficient (Cut-points) as Gap Measures

In Kentucky there are three major cut-points, producing four major scoring categories - Novice, Apprentice, Proficient, and Distinguished. NCLB requires at least three categories of performance and that percent proficient be the cut-point for determining gaps.

There are several desirable properties of defining the gap in terms of cut-points.

  1. There are several well-defined, judgmental methods to define the cut-points, i.e. what will be called a proficient performance.
  2. Given the defined cut-points, it is straight-forward to calculate the gap and changes in the gap. This is especially true for summing across grade levels and content areas within a school.
  3. Coupled with a long-term goal of each student being proficient, the gaps are eliminated when the goal is met.
  4. The notions of being proficient in a subject area and having the percent proficient be the indicator of success, are easily conveyed to a broad audience.

There are several undesirable properties as well.

Perhaps the most serious one is depicted in Figure 1D. It shows that if the cut-point is at 40 rather than 50, the gap will be almost double the size. That is, the size of the gap varies according to where a cut-point is placed. Since the methods used to determine cut-points are judgmental, there is no one logical, well-defined place on the scoring scale to place a cut-point. That is a major reason why different states have different percents of students who are proficient.

Another weakness of cut-points as proficiency standards is that if those in the school wished to "game" the system, it is clear how that might be done. A gap can be narrowed by dealing with only a small proportion of the students. One should focus on students in the lower scoring group who are below but not too far below the cut-point. When they are moved to or above the cut-point, the gap is narrowed despite the performance of lowest scoring students. So differences in the percent proficient can be minimized by working with relatively few students.

Conversely, a school could increase dramatically the scores of the lowest scoring students without having an impact on the percent proficient. Imagine moving each student below the cut-point closer to the cut-point. Although the accomplishment would be dramatic, it would have no impact on the percent proficient.

The combination of using cut-points with a rule that each student must be proficient in a certain amount of time, gives a school an impossible task. Figure 1C shows where the cut-points of 1D fall on the score distributions. When the percent proficient is at a score of 50, 90 per cent of students in Group 2 must be moved to or past the cut-off. For Group 1 which is four times greater than Group 2 more than 80 percent of students must be likewise moved. When the cut-point is lower, the task is less onerous, about 70 and 50 percent respectively. I know of no empirical results that show such dramatics effects.

Finally, the whole idea of being proficient may be illusory. Simply placing a label on a test score does not make it true. Tests labeled science, for instance, may be very different kinds of tests. The science portion of Explore, the ACT eighth grade test contains only multiple choice questions and requires an inordinate amount of reading. The National Assessment of Educational Progress (NAEP) eighth grade science contains constructed response and extended constructed response questions and tends to minimize the effects of reading. Whatever proficient may be, it is likely to result in substantially different definitions depending on what science measure is used. And they both are wrong!

Mean Differences as Gap Measures

Just as for cut-points, defining achievement gaps in terms of mean differences have both desirable and undesirable properties. The positive aspects of such a definition include:

  1. Given data that are approximately bell-shaped the mean is a good typical value;
  2. As opposed to a cut-point definition where not all students are affected, the mean takes into account all cases.
  3. An average is a number most persons understand.

But, as I tell my students "never a center without a spread." Figures 1E and 1F show the effects on differences between groups when the spreads differ. The difference between the figures is about 2 1/2 points, a standard deviation of 10 for the first four and between 7 and 8 for the last two. The differences in the cumulative distributions get rapidly "fatter" above the mean of 40 (incidentally, the area between cumulative distributions is equal to the difference between means for the two groups). Minimizing differences when spreads are small may mean something different than when they are large.


Because decreasing mean differences may mean different things depending on the spread of data, it creates interpretation problems across grade levels and content area. Unlike summing percents based on cut-points, there is a question of how one should sum the effects to get an overall school index.

It is possible to "game" the means, although effects may be smaller than what one gets when gaming the cut-point definitions. If one believes, for example, that there are faster and slower learners, then to focus on relatively fast learners in the lowest scoring group could provide bigger gains that focusing on each of the students.

Finally, if it were just a matter of reducing differences between means, there would not necessarily be improvement across the system. So, there should be some specification of an expected amount of improvement.

Effect Sizes and Mastery Learning

An effect size, classically defined, is the mean for a treatment group, minus the control group mean, divided by the control group standard deviation.

This standardizes mean differences making them interpretable in terms of standard deviation units. The general idea can be used in the context of gap differences. For the data I have displayed, Group 1 has a mean of 40 and Group 2 has a mean of 35. Using the larger group's standard deviation of 10, we come up with an effect size of .5, that is, Group 1 performance is on the average 1/2 of a standard deviation higher. That magnitude of effect often would be interpreted as a medium sized.

These effect sizes can be summed over content areas and grade levels in a school to produce a school index. It would take some empirical work to decide how much the index should be reduced in order to say that an achievement gap is closing.

Although effect sizes respond nicely to the question of different spreads they do not help when it comes to different shapes. When Ben Bloom in 1967 outlined the properties of his approach to Learning for Mastery, he recognized the problem of only dealing with average improvement. So his goals included not only influencing average performance but also influencing the spread and shape of performance. The goals are to raise the mean, minimize the variance, and skew the distribution! A desirable outcome, then, is a heavily positively skewed set of higher scores rather than ones that look bell-shaped.

I don't know of anyone who has argued for reducing spreads and creating positive skewness as measures related to closing the achievement gap. Perhaps someone should. It may be worth a look.

Conclusions

If I were to decide what to use as indicators for defining a gap and determining whether it has been closed, I would not use either a method based on cut-points or simple mean differences. I would start with effect sizes and then do some analyzes to determine whether indicators of reducing variation or creating positively skewed outcome data are other possible measures.

What ever measure is chosen, it should be grounded in empirical results. So, there is a major task for the assessment persons in the Kentucky Department of Education to analyze their assessment data and come up with defensible suggestions for measuring a gap, measuring how much it changes, and how much it must change before deciding that the gap has been reduced.


Caveat

I have tried to respond directly to the gap issues without divulging my reluctance to base decisions about what is a good or effective school simply on the basis of test scores. Or, for that matter, whether schools should be held accountable for "gaps" that are based only on test scores. There is what I consider a naive view that backgrounds of students should be ignored when looking at whether schools are effective. At the same time there is an almost religious belief in the efficacy of test scores as the way to determine whether a school is good. Such views defy common experience and ignore research about schools and schooling. Some schools, for example, have relatively small amounts of turnover during a school year; others turnover almost completely. Some schools have huge amount of parental participation; others have virtually none. And, it remains true that the strongest within country correlations with test scores in international studies are based on the background characteristics of students.

The effects of schooling are many, diverse, desirable and undesirable, both short term and long term. Tests get at a small number of similar, desirable, short term effects. NCLB ignores most content areas in judging schools. The Commonwealth's assessment measures fewer than half of its goals. What ever happened to self-sufficiency, effective group membership, and integration of knowledge?

Tests do not get at whether a school produces persons who are thoughtful and reflective. They do not get at whether persons are well-informed. They do not get at how well persons work together or how they well they respect other persons and other points of view. They do not get at whether a school produces good citizens. Good schools do all of the above! Those things are as worth thinking about as is the achievement gap, however defined.

[1] I produced these data. They do, however, mimic those I analyzed for a paper on the gap.

Is the ACT the Sole Criteria?

It all started when Council on Postsecondary Education President Bob King argued in the Herald-Leader that CPE's High School Feedback Report allows the state "to look more deeply into actual performance measured by an external, unbiased resource — the ACT exam."

King suggested that the ACT, by itself, was superior to predictions of college-readiness derived from combinations of data. Is CPE discounting graduation rates and average GPA in favor of a single test?

That drew a response from KSN&C's Skip Kifer, demonstrating the weak relationship ACT musters, and suggesting that CPE should propose placement procedures based on a robust notion of "readiness" that includes more than just the ACT and that did not violate test score use standards. It appeared to Kifer that CPE arbitrarily uses that single test score to determine whether a student is ready for regular course work in Kentucky's public universities.

Too clarify, King stated that CPE does not rely exclusively on the ACT to make college admission or placement judgments, nor does the Council on Postsecondary Education encourage such determinations.

Perhaps King should have reviewed CPE's printed material before saying that.

In last Saturday's H-L, Kifer wrote,

I am puzzled by Bob King's response to my critique of the Council on Postsecondary Education's policy of declaring a student college ready on the basis of a single test score.

King, director of the council, says: "Please allow me to clarify that Kentucky's colleges and universities do not rely exclusively on the ACT to make college admission or placement judgments, nor does the Council on Postsecondary Education encourage such determinations."

Yet when I look on the council's Web site, it says:

"The Kentucky statewide public postsecondary placement policy in English and mathematics applies to any student entering a Kentucky public college or university. The policy is based on your ACT or SAT score and determines what type of English and math classes you will need to take when you enter college."
This is what I found in my search. Notice the date posted. That's the day Kifer's piece ran. Maybe I missed something, but I didn't catch any changes in the language from when I looked at the same material in February.

StatewidePlacementPolicy: Postsecondary Placement Policy does … Postsecondary Placement Policy, please … PLACEMENT POLICY IN http://cpe.ky.gov/nr/rdonlyres/73e9a7b3-84dc-4ec2-8f1b-6a99261b5fb4/0/statewideplacementpolicy.pdf
- 120KB - kdrummond - 3/5/2011 [View duplicates]

And I found this:
The statewide placement policy is applicable to any incoming student entering a Kentucky public postsecondary institution. ACT and SAT standards form the basis of the policy because Kentucky uses the ACT (or equivalent measures) for college admissions and placement decisions.
That language says ACT forms the basis for placement decisions but stops short of saying the ACT is the only determinant. That comes next.

Kentucky Statewide Placement Policy in English
• A student earning an ACT English sub-score of 18 or higher qualifies for placement in a credit-bearing writing course at any Kentucky public postsecondary institution.

Kentucky Statewide Placement Policy in Mathematics
Three levels of readiness are identified for placement in a credit-bearing mathematics course at any Kentucky public postsecondary institution:
• Level 1: A student earning an ACT mathematics sub-score of 19 or higher qualifies for placement in a credit-bearing mathematics course, but this course may not be a requirement for many college majors or lead to subsequent coursework in mathematics. Mathematics for liberal arts is an example of such a course.
• Level 2: A student earning an ACT mathematics sub-score of 22 or higher qualifies for placement in college algebra. College algebra (or placement in more advanced courses) is required for majors such as biology, business, economics, information systems, and technology. College algebra can lead to any major.
• Level 3: A student earning an ACT mathematics sub-score of 27 or higher qualifies for placement in calculus. Calculus is required for majors such as mathematics, physics, chemistry, computer science, engineering, biology, business, and technology.

Kentucky’s statewide public postsecondary placement policy is a guarantee of
placement in credit-bearing coursework to incoming students demonstrating
specified levels of competence.

Monday, February 14, 2011

King Clarifies Stance on ACT and College Admissions

Council on Postsecondary Education President Bob King recently offered his opinions on college admissions saying,

Historically, parents are often directed to focus on graduation rates and average GPAs as evidence of how their high school is performing. Our reports allow parents and educators to look more deeply into actual performance measured by an external, unbiased resource — the ACT exam — now required of all Kentucky students.
That drew a response from KSN&C's Skip Kifer.

Bob King, president of Kentucky's Council on Postsecondary Education applies the council's arbitrary standard of using a single test score to determine whether a student is ready for regular course work in Kentucky's public universities.

He implies a test score is a better predictor of grades in college than is a high school record. He then presents results from one high school that lump higher performing students (those with above-average high school records) with lower performing ones in a misguided approach to justify his position. A test score, however, does not make or break a student's readiness for higher education.

Today, King clairifed his stance in the Herald-Leader.

...Please allow me to clarify that Kentucky's colleges and universities do not rely exclusively on the ACT to make college admission or placement judgments, nor does the Council on Postsecondary Education encourage such determinations.

A student's entire record, including GPA, extracurricular activities, and other placement exams form a portfolio that allows campuses to make informed decisions on admission and placement.

The ACT serves as an important element in this consideration, but more importantly, it serves as an alarm bell in the student's secondary experience about preparation for life after high school.

This might have been a good place to stop. It acknowledges Kifer's concerns and clarifies King's stance. But King then makes allusions to "certain thresholds" in the ACT which serve to warn us if a student is not on track.

It warns of the need to take a deeper look at a student's college readiness if the scores fall below certain thresholds, but it is not the sole determinant when placing students in developmental courses... Far from arbitrary cutoff scores, there is a great deal of data from tens of millions of ACT score results upon which policy makers in Kentucky rely to set the scores used to indicate college readiness in key entry-level courses.

If the thresholds King has in mind are, in fact, a reference to ACT's benchmarks, which are inappropriately modeled, one wonders if King's effort to lay the issue to rest might draw yet another response from Kifer.

We'll see.

Monday, January 31, 2011

ACT: Not the Only Measure of College Readiness

Council on Postsecondary Education honcho Bob King recently argued in the Herald-Leader that CPE's High School Feedback Report allows the state "to look more deeply into actual performance measured by an external, unbiased resource — the ACT exam."

King suggested that the ACT, by itself, was superior to predictions of college-readiness derived from combinations of data. Discounting graduation rates and average GPA in favor of a single test prompted our resident testing expert to retort.

NOTE: H-L seemed to struggle editing Skip's piece, so here's the unadulterated article the paper titled:

If I were to assert that a player who cannot make 56% of his free throws is not "ready" for the NBA, a fan would point out that there is much more to basketball than shooting free throws. An astute fan with a historic prospective would point out that Wilt Chamberlain, Shaquille O'Neal and a bevy of other current players would not be "ready" using that arbitrary standard. One facet of basketball does not make or break a player's "readiness."

Bob King, president of Kentucky's Council on Postsecondary Education, in a recent op-ed piece applies the council's arbitrary standard of using a single test score to determine whether a student is "ready" for regular course work in Kentucky's public universities. He implies a test score is a better predictor of grades in college than is a high school record. He then presents results from one high school that lump higher performing students (those with above average high school records) with lower performing ones in a misguided approach to justify his position. A test score, however, does not make or break a student's "readiness" for higher education.

Decades of research indicate: performance in academic courses in high school is the single best predictor of success in higher education; a combination of the high school record and test scores predict better than the high school record alone; and, how good the prediction is and how the components are combined vary depending on the institution. Although there is a general pattern of the primacy of the high school record, there is no one-fits-all model to predict grades in different courses or different institutions.

In addition to the thoroughly suspect notion of labeling a test score readiness, the council's use of a single score for placement purposes violates standards for the proper use of tests. Those standards include the following:

In educational settings, a decision or characterization that will have major impact on a student should not be made on the basis of a single test score. Other relevant information should be taken into account if it will enhance the overall validity of the decision.

In addition:

When test scores are intended to be used as part of the process for making decisions for educational placement, promotion, or implementation of prescribed educational plans, empirical evidence documenting the relationship among particular test scores, the instructional programs, and desire student outcomes should be provided. When adequate empirical evidence is not available, users should be cautioned to weight test results accordingly in light of other relevant information about the student.

Apparently, the council determines readiness by doing statistical analyses of ACT scores and grades in first year courses without regard to institution. A certain ACT score produces a 50/50 chance of getting certain grades, say C, or better. I could find no information about this or other investigations done by the council. And, although it is possible to present information about how good a model is, I could not find any information of that kind either.

To give a sense of the power of statistical models to predict first year grades I report analyses conducted years ago on University of Kentucky student samples. The question was whether results of KIRIS, the first commonwealth assessment related to school reform, could be used for admission and placement in a university.

If a model exactly predicts grades one can say that the model accounts for 100 % of what could be known. If a model cannot at all predict grades, one can say that 0% is accounted for. One way, then, to talk about the power of a statistical model is determine what percent the model predicts.
The table below gives those percents for different courses at UK for three different statistical models: High school record only, High School record + ACT scores and High School record + KIRIS scores.



The first thing to recognize is that the models are not particularly powerful. They rarely account for 25% of what could be known leaving 75% to be explained. That 75% may be differences in students' study habits, class attendance, interests, any of a thousand other variables or simply things not explained statistically.

The pattern of results, however, is clear. Adding an ACT score to a model containing GPA makes the prediction better but not greatly so. The same is true for KIRIS scores, too. Incidentally, that was without including the KIRIS writing sample.

These are not unusual results. They point, obviously, to gathering more information about a student before making a placement decision. Here is what ACT says:

ACT offers a variety of tools to ensure postsecondary students are quickly and accurately placed in courses appropriate to their skill levels. Assessment tools from ACT offer a highly accurate and cost-effective basis for course placement. By combining students' test scores with information about their high school coursework and their needs, interests, and goals, advisors and faculty members can make placement recommendations with a high degree of validity.

To that, I would add, for obvious reasons, it is desirable for an educational agency to use tests in exemplary ways.

The op-ed piece goes on to exhort parents to ask right questions, asks an undefined "we" to fear international test results, says admissions offices should align themselves with the council's readiness standards, and the still undefined "we" to serve teachers more effectively. Such exhortations would be more convincing if, in the first instance, the council could propose placement procedures based on a robust notion of "readiness" that in addition did not violate test score use standards.

Tuesday, October 12, 2010

Challenging Assertions

What is clear,
and Kentucky has 20 years of experience that shows it,
is that for any test that is adopted
teachers will be coerced into teaching to it.
Despite all the work with national standards,
students may emerge from schools
expected to know only what is on tests.

-- Skip Kifer

In a Monday Herald-Leader Op Ed titled, "New standards alone won't improve education," assessment guru Skip Kifer challenged a couple of ideas offered by Educati0on Commissioner Terry Holliday and CPE head Bob King.

At issue was their Sept.19 column, "Smarter standards; Ky. tackling challenges to meet ambitious goals".

Education Commissioner Terry Holliday and Council on Postsecondary Education President Bob King in their op-ed focus on national education standards and related issues....They say, "It has been common that students who have been earning A's and B's in high school have been placed in remedial courses in college." They say students have not been learning the right stuff in high school and that national standards will solve the problem.

... decades of research has shown the single best predictor of success in college is a student's high school record. Doing well in a rigorous college preparatory curriculum is the key to doing well in higher education. How have or will national standards change that?

Perhaps these hypothetical students were not in college preparatory classes. In that case, one would not expect them to be related to success in college.

Then Kifer disputed international comparisons. Holliday and King stated that, "Among the 40 most industrialized nations, American students perform in the bottom quartile on international math and science examinations."

In the latest 2007 Trends in International Mathematics and Science Study (TIMSS), U.S. fourth graders were ranked 11th of 36 in mathematics. Their eighth grade counterparts were ranked ninth of 48...

Kentucky did not participate as a state in TIMSS. However, the commonwealth's fourth and eighth graders participate in the National Assessment of Educational Progress (NAEP) assessments. Their scores are at about the national average. It is likely then that Kentucky performance on TIMSS would be similar to the U.S. as a whole; they would be in the top one-third in 4th grade and the top 20 percent in 8th grade.

Similarly in science Kifer finds that Kentucky's performance on the TIMSS could be in the top 20 percent: sixth or seventh at fourth grade and ninth or 10th in eighth grade.

Holliday and King stated that, "Since early this summer, teams of teachers and college faculty have been translating sophisticated technical language (of the national standards) into objectives students and their parents can understand." But Kifer wonders,

Suppose another state were doing the same thing. Is there any reason to believe the two groups would come up with the same statements? It would appear the process is a way to turn national standards into state curricula, a problem national standards were meant to solve.

Monday, April 20, 2009

Never a Center Without a Spread

By Skip Kifer

When I ask my classes “who runs faster boys or girls,” someone would, after having decided it is not a trick question, say “boys.” If one is talking about averages or centers of distributions that is the correct answer. If one is talking about how fast people can run, it misses the mark substantially. How many men in the world can run faster than Marion Jones?

Education talk is filled with centers when it should be filled with spreads and, on the best of days, centers, shapes and spreads. Let me give you a couple of examples:

Which of the following two pictures, based on the same data, describe an achievement gap?



An answer is, of course, they both do. But what do they portray and which is the better description? And, are there better descriptions?

The recent entries concerning equity in educational funding in the Commonwealth portray kinds of centers. Here is another picture that contains the data and spreads as well.

I do not see the equity in funding. In fact, spreads, a measure of inequity, may have increased.

Anytime I see reported only averages or only percentages I am immediately suspicious. There are simply better ways to portray data and to make it possible for persons to understand better what the data may be saying.


(Editor's Note: Correction made in text April 21.)

Monday, March 30, 2009

Senate Bill 1 ends the CATS era. What will replace it?

By Skip Kifer

SB1 proposes a huge amount of testing. Some testing is statewide; other testing is not. Some tests are used for accountability purposes; others are not. Some tests are multiple- choice only; others are not. There are formative tests, summative tests, performance assessments, benchmark tests, interim tests, norm-referenced tests, criterion-referenced tests, end of course tests, and writing portfolios as well as program reviews and program audits which are used to measure students, schools and districts.

The Kentucky Department of Education (KDE) is directed to integrate the testing into a new assessment system. It has almost three years to complete the task.

I doubt that a defensible new system can be created, even with that amount of time, without additional resources for KDE and more focused testing.

A first step to building a new assessment is to find out whether what is presently being done works. Major components of the existing system should be evaluated. For instance, EPAs - Explore, Plan and ACT - produces results presumed to help students learn and school personnel make better decisions. Does it? What is the evidence that the information is used? What is the evidence that using the evidence makes a difference? On what? Studies of each major component should provide answers to those or similar questions. If the component is not working, it should be eliminated.

There should be a thorough understanding of the implications of moving from an assessment system focused on schools to one focused on students. CATS relied on different forms of the test to get a better sample of achievement within a school. A student might spend an hour taking a test but with six forms of the test, the assessment contained about six times as much information for each school. Even with that, about 25% of schools were misclassified. Some were said to be progressing when they were not. Others were said to be needing assistance when they were progressing.

The existing assessment had thousands of misclassifications at the student level, too. Proposing to produce longitudinal data on students places more pressure on the assessment and more obstacles to accurate measurement. Creating measures that can be scaled across grades suggests spending more time testing each student. Is that desirable? Is it practicable?

Finally, the biggie! After almost two decades of educational reform with a spotlight on accountability through high-stakes testing, is their evidence that such testing works?

Research results both in Kentucky and nationally are mixed. Test scores go up but there are questions of whether that is mainly because of teaching to the test. Because what is taught and tested narrows, higher test scores may not mean mastery of a content area and may mean that content areas not tested are not emphasized. In addition, states with high-stakes testing do no better, and perhaps worse, on national tests than states without such programs. Reasonable persons would agree that test scores are, at best, a narrow reflection of successful schools. Other aspects are more important. Evidence should be collected, therefore, that indicates whether after about two decades of high-stakes testing Kentucky schools:

a) are better places for kids than they were prior to the reform;
b) nurture talent in ways that it should be nurtured; and
c) insure those who emerge from the school system have commitments to democratic ideals, participation in democratic communities, and tolerance for varied persons and views.

Armed with good evidence, resources and time, I hope KDE creates a system that benefits the Commonwealth and its children.

Thursday, March 26, 2009

To Predict or Not to Predict

By Skip Kifer

A colleague and I were doing a session on formative assessment that began with, would you believe, a formative test on creating and using formative assessments. We generated a lively discussion when going over answers to the test, particularly the answer to this question:

Which of the following is NOT a purpose for using formative assessments?

a. To learn which students are doing well and which are not doing so well.
b. To gather data about what has been taught well or not so well.
c. To predict future performance on a norm referenced test.
d. To provide a basis for future instruction.

C is the correct answer. Some teachers insisted that C was not a correct answer because prediction, too, was a purpose of formative assessments. They used as examples what they have learned recently about “benchmark” assessments.

There is no one definition of formative assessment. And there are those, mainly vendors like ACT, who say that benchmark assessments are formative assessments. They are not.

The crux of the difference between benchmark assessments and formative assessments is the difference between predicting and correcting. Benchmark assessments predict; formative assessments correct.

I give formative tests after teaching an instructional unit. The tests:

a. cover what has been done in the unit;
b. are scored but not graded;
c. are discussed immediately upon completion; and
d. are discussed (hopefully) in earnest by the students.

They are now, therefore, instructional tools. They convey to the student what was important in the unit. The students convey what they learned. My aim is to get most students to know most of the material in the unit. In doing so, I want high scores but small differences, less variance, in the scores. I want each student to know more and the class to be more alike.

The discussion is the first part of correcting. By getting a conversation going, I hope that students better understand what I thought I was teaching. The next step, however, is to look at the students’ scores and figure out what additional experience each student needs in order to master the material. I also peek at the scores to see what I did well and not so well.

Contrast the formative assessment scenario with a benchmark one.

Benchmark tests may or may not be related to what students have been taught. Alignment studies have to be conducted to determine if they are and to what extent.

Since the questions are not released, a teacher does not know what has been done well or poorly. Students have no idea how well they are doing because the questions and the right answers are never discussed. There is no way to tie the results to what a teacher has been teaching because the test, at that time, does not necessarily reflect what has been covered in the curriculum.

What benchmark tests do is predict scores on other tests. That is, benchmark results are used to estimate the extent to which student scores on a first test are replicated by results of a second test. That prediction is strongest when the variation on the tests is largest. On the other hand, the more similar are student scores, the worse the prediction. If each student did exactly the same on the first test, the prediction would be zero. So my attempts to use formative assessments to get high scores with small differences between them fly in the face of strong prediction.

Imagine a student entering a classroom and the teacher saying she can predict where the student will be at the end of the year. The student will be about as much above the mean or below as they now are. That strong prediction would be desirable from a benchmark point of view. It may not be so desirable from the student’s view. But that’s predicting!

Imagine that same student entering a classroom and a teacher saying we are going to work together to master the material in this course. She will use formative assessments to help the student. That would imply no relationship between the status of students upon entering the class and their final status in the class. That is correcting. And, teaching!

Just for fun here is the formative test on formative assessment and an answer key.

A formative assessment on formative assessment…..

1. The person who first distinguished between formative and summative approaches was:

a. Paul Black
b. Benjamin Bloom
c. Michael Scriven
d. Richard Stiggins


2. When the chef tastes the soup, it is _____? When the customer tastes the soup, it is _____? Choose the two words that best complete these thoughts.

a. summative, formative
b. formative, summative
c. objective, formative
d. summative, objective

3. Which of the following is NOT a purpose for using formative assessments?

a. To learn which students are doing well and which are not doing so well.
b. To gather data about what has been taught well or not so well.
c. To predict future performance on a norm-referenced test.
d. To provide a basis for future instruction.

4. Formative assessment activities in the classroom used properly:

a. Provide a basis to rank order students.
b. Predict performance on other tests.
c. Help the teacher know what she has done well or poorly.
d. Add to the precision and usefulness of grades.

5. The Taxonomy of Educational Objectives:

a. Orders outcomes according to how difficult it is to teach them.
b. Is based on learning hierarchies
c. Shows what is desirable
d. Is ordered by cognitive complexity

6. Suppose students in a class were asked to memorize the names of those who came over on the Mayflower. A test question then asked them to write down 10 of those names. At what level of the Taxonomy is the response likely to be?

a. Knowledge
b. Comprehension
c. Application
d. Analysis

7. Suppose the list contains occupations of the Mayflower passengers. Students are asked to write an essay that draws inferences about why certain occupations were present on the ship. At what level of the Taxonomy are the responses likely to fall?

a. Knowledge
b. Comprehension
c. Application
d. Analysis

8. Suppose the list of names of persons on the Mayflower also included their occupations. Students are asked to write an essay that draws inferences about why certain occupations are on the ship. That would represent what Depth of Knowledge?

a. DOK 1
b. DOK 2
c. DOK 3
d. DOK 4


Key to Formative Test

1. C. Scriven, M. (1967). The methodology of evaluation. Tyler, Gagne & Scriven (Eds). Perspectives of curriculum evaluation. (AERA Monograph Series on Curriculum Evaluation, No. 1). Chicago: Rand McNally
2. B. Bob Stake coined this to distinguish between the two.
3. C. Prediction comes from benchmark assessments. A teacher would like all of her students to do well on the formative tests. If so, there would be no correlation (prediction) between the formative test results and some norm referenced test results.
4. C. First of all, did the teacher do well, on what? Secondly, which students did well and which did poorly, on what?
5. D. Bloom called the Taxonomy the most used, least read book in education. The taxonomy orders outcomes according to cognitive complexity.
6. A. Recall would fall under Knowledge, the least complex cognitively of the levels.
7. C or D. A taxonomy purist would probably say D. I have difficulty getting beyond Applications.
8. C or D. I don’t do Depth of Knowledge. I don’t think it adds anything to the taxonomy. There are lots of classification schemes. I am happy being familiar with one.

Tuesday, February 24, 2009

The New Commonwealth Assessment

By Skip Kifer

Lord Mansfield is reputed to have said “Decide promptly, but never give any reasons. Your decisions may be right, but your reasons are sure to be wrong.” I hope his admonition is correct because I get a sinking feeling each time I see another list of what form the Commonwealth’s new assessment should take.

The reasons for the sinking feeling are many. Here I will deal with just three.

The first has to do with the technical aspects of whatever form the new assessment takes. I believe the substantive ideas behind the assessment, not its technical requirements, should drive it. Yet, good technical advice early can save time, money and embarrassment. In the past the Commonwealth has been blessed with unusually good technical advice from panels that were composed of assessment persons who were in the mainstream of the assessment world. They gave dispassionate advice that vendors or persons not in the mainstream could not be expected to give. There should be technical assistance early and often as plans for a new system develop.

There are a number of proposals that make me wonder about the basis for them. To take just one: that formative assessment instruments could or should be provided by vendors. I have taught test and measurement classes for over 30 years and I know that I can teach teachers to make better tests than are presently being used (apart from CATS) in the Commonwealth.

Some very important things go on when one constructs a test. One needs to think seriously about standards, instructional materials, and the learner as the items are written and tests developed. Having gotten the results one has to think seriously about whether students have not learned because they have not been taught well or because the items are no good. Each of these activities is related to effective instruction. None of them come with already made assessments.

I would like to hear the discussion of whether difficulties with the writing portfolio are unique to that form of assessment or whether having assessments in any accountability context corrupts them regardless of their form. Are some kinds of assessments less corruptible than others? Are there practices that are less likely to be gamed by those who seek higher scores at the expense of better learning? Are there ways to gather and report learning outcomes that depend less on massive formal testing procedures? I believe the answer to each of these questions is “yes.” And I think one could give good reasons for answering these questions in the affirmative.

And so, I will enter the fray. Below I am presumptuous enough to come up with a set of principles that should guide the design of the new Commonwealth assessment. I made the mistake of including things we have learned as my reasons for the recommendations. I wonder if there is a Lord Mansfield redux.

New Assessment Principles

After almost two decades of high-stakes assessment in the Commonwealth it is time to step back and decide how what we learned can help us think about the new directions we should take.

1) We learned that assessments can be very time consuming and very costly. We learned that they may not be as cost effective as we hoped.

A new statewide assessment should be:
a. Clear about each of its purposes
b. Less time consuming
c. Less costly

2) We learned that teachers believe they spend far too much time testing and preparing for testing.

A new statewide assessment should
a. Carefully delineate what should be assessed statewide and what should be assessed locally and be controlled by teachers
b. Be restricted to a small but crucial part of the curriculum

3) We learned that parents wish to know how their children’s scores compare with other children’s scores nationwide. We know that off-the-shelf tests have user norms not nationally representative ones. We know that the National Assessment of Educational Progress (NAEP) is the only national test with a nationally representative sample.

A new statewide assessment should
a. Use wisely the NAEP state assessments
b. Link Kentucky assessments to NAEP
c. Provide students with scores that can be compared to national ones

4) We learned that measuring so many content areas at so many grades in all schools is an inefficient way to assess Kentucky’s educational progress.

A new statewide assessment should:
a. Sample schools to track progress over time
b. Accept NAEP scores in mathematics, reading and science as progress indicators

5) We learned that testing requirements of No Child Left Behind (NCLB) have a profound effect on what and how a state assesses.

A new statewide assessment should be
a. Cognizant of new requirements of NCLB
b. Prepared to seek exemptions from some requirements of NCLB

6) We learned that others admired Kentucky’s assessment when it was bold and pioneering

A new statewide assessment should

a. Lead other states and nations in producing useable information about students and schools
b. Emphasize formative and instructionally embedded assessments over summative ones
c. Place students and teachers in the center of schools and assessment

Wednesday, February 11, 2009

Let’s Hear it For Clear Standards

By Skip Kifer:

About 20 years ago the National Council of Teachers of Mathematics (NCTM) produced its first set of standards. About 20 days ago Kentucky’s Republican Senators discovered NCTM’s latest set.

During the interim NCTM standards have informed Kentucky’s and most other state’s assessments, President Clinton’s Voluntary National Test, and the National Assessment of Educational Progress (NAEP). The standards were used both as statements of instructional content and, when refined, as frameworks for assessments.

The standards emerged after the Second International Mathematics Study (SIMS) identified an “underachieving” U.S. mathematics curriculum, declaring instruction to be a “series of one night stands” in a curriculum that instead of “spiraling was a set of concentric circles.” Rather than take on the political task of establishing a national curriculum, NCTM promulgated content standards.

Talk about standards then was understood to be talk about content standards – what should be taught, when. Not now. Increased emphasis on assessment in the country brings into play other sets of standards. As the Kentucky legislature seeks to change the Commonwealth’s assessment, a robust standards discussion means distinguishing one set from another and knowing when each applies.

Clarity is needed.

Here is an example of an NCTM content standard:

Data Analysis and Probability Standard Instructional programs from prekindergarten through grade 12 should enable all students to—

Formulate questions that can be addressed with data and collect, organize, and display relevant data to answer them

Grades 6–8 Expectations In grades 6-8 all student should:

formulate questions, design studies, and collect data about a characteristic shared by two populations or different characteristics within one population;
select, create, and use appropriate graphical representations of data, including histograms, box plots, and scatterplots.

The Lexington Herald-Leader reported a rationale for adopting the new NCTM standards: “we need higher standards,” and we need “more precise and rigorous standards.” It is unclear how restating a content standard would meet such demands. Those of us who have worked with content standards know the issues are not whether they are high or low. Content standards are tricky but on other dimensions: if too broad, they are hard to operationalize; too narrow, and they miss stuff; too many, and they confuse things; too few, we don’t know about since that’s never been done.

Those who call for “higher standards” are probably talking about different standards -proficiency standards. Proficiency standards are based on test scores. Below is a picture of the distribution of KIRIS reading scores with the proficient cut-point for 4th grade Kentucky students.

About 1/3 of the students scored at or above the cut-point of 311 and would be labeled proficient. Clearly one can ask whether this standard is too high or too low. Should one move the cut-point up or down? Interestingly, six years earlier just 10 percent of the students exceeded this cut-point. What might look low one time might earlier looked like high.

There are at least two other sets of standards worth discussing – Opportunity to Learn (OTL) standards and standards for educational and psychological testing. The latter are published by the American Educational Research Association, American Psychological Association, and the National Council on Measurement in Education (AERA/APA/NCTM).

OTL standards reflect whether or not a student has had a chance to learn the material in the test. They are fairness standards. International tests of 8th grade mathematics, for example, typically contain heavy doses of algebra and geometry. Since most of our students are not exposed to substantial amounts of algebra or geometry by the 8th grade, while other systems’ students are, we are at a disadvantage. We lack the opportunity to learn. Hence, we wonder about the validity of international comparisons.

OTL comes into play, too, when tests are used for purposes such as graduation or grade promotion. Have the students had the opportunity to learn the material covered by the test? If not, then the test is not valid for the purpose it was intended.

Using a single test score for important educational decisions is, in any case, considered bad testing practice. That dictate and other proscriptions are in the AERA/APA/NCTM standards volume. Among others, there are standards for properly constructing tests, reasonably interpreting test scores, and providing information about a test and its proper uses. Any major testing endeavor is expected to comply with those standards.

So, what does all of this have to do with Kentucky’s Senate Republicans? Well, as they pursue dramatic changes in assessment practices in the Commonwealth, they may not be as informed as they might be. As a first step to better communication, I welcome them to the standards club and to some clarity!

Kifer and Sanders join Kentucky School News and Commentary

For the first time in my life, I'm actually trying to keep a New Years Resolution. This year I resolved to add new voices to Kentucky School News and Commentary, and today it begins.

KSN&C is pleased to announce that Skip Kifer and Penney Sanders - a couple of kindred spirits - will be the first new voices on board.

Don't expect us to agree with one another. That's not the deal. It may happen, but just about as often I suspect we will diverge on topics as well.

Kifer and Sanders were recruited - not for their compliance to any ideology - but for their experience and insights informed by years in the field. Both are dedicated to improving Kentucky schools and I respect their professionalism.

Edward "Skip" Kifer's expertise is measurement, evaluation and statistical analysis and I believe you will find his commentary surprising, challenging and refreshing. Following a stint in the army, Kifer was educated at the University of Chicago, was a Fulbright Scholar, taught at UCLA, SUNY Buffalo, the University of Stockholm, UK, UofL and Georgetown College. Professor Kifer was an AERA Senior Research Fellow working with NCES and NSF and has an established reputation as a reliable designer of assessments; one who has served on numerous technical panels including the NAEP Design and Analysis committee.

Dr. Penney Sanders is a retired Kentucky educator who served as a math teacher, school counselor, assistant principal and principal. She was Kentucky's first Director of the Office of Educational Accountability, and as far as I can tell, it's most effective. Under Sanders, adherence to KERA was an expectation enforced with panache. Veteran Kentucky educators still remember the good ole days when Sanders stalked the hills; kicking butt and taking names. In her "retirement" she works as a consultant with a special focus on high-risk, low performing schools. She continues to write and pursue a number of interests in education.

And we're not done.

Invitations to a few other contributors have been, and will be, made. They will fall into two categories, "veterans" and "newbies." Kifer and Sanders are examples of the veterans, but I am also hoping to include a few of my masters/doctoral level students who are themselves practicing teachers in Kentucky - to keep us focused on what is really happening in Kentucky's schools.

If all goes well, Kentucky School News and Commentary will add more value to the public debate while remaining realistic about the implications our decisions have on Kentucky's teachers and students.

Sunday, October 19, 2008

Innes smacks Kifer. Kifer smacks back.

We recently had a little dust up over testing between education analyst Richard Innes and Georgetown/UK Professor Skip Kifer.

The Bluegrass Institute's Richard Innes, argues that credibility and stability problems with the CATS are reason enough to dump it. In this argument, the GREAT is the enemy of the GOOD. Innes pounds on the CATS' several flaws; hoping that by seeking perfection the utility of the instrument will be doomed. He thinks there's a better idea. Just throw the CATS out. Maybe, replace it with the ACT. After all, the ACT is now a super test with super powers. Just ask ...the folks at ACT.

Ben Oldham recently wrote, "Since the ACT is administered to all Kentucky juniors, there is a tendency to over-interpret the results as a measure of the success of Kentucky schools."

Innes, assisted that over-interpretation mightily, and decided he would "school" Professor Oldham saying,
Oldham pushes out-of-date thinking that the ACT is only a norm-referenced test. The ACT did start out more or less that way, years ago, but the addition of the benchmark scores, which are empirically developed from actual college student performance to indicate a good probability of college success, provides a criterion-referenced element today, as well.

Well, I'm no testing expert but even I knew that was wrong. Wrong enough, that I began to worry that BGI's testing expert might have some holes in his own preparation.

KSN&C responded,

The problem of over-interpretation has been somewhat exacerbated by the inclusion of benchmark scores in the ACT. But benchmarking does not change the construction of the test nor the norming procedures. It does not turn the ACT into a criterion-referenced exam as Innes tries to suggest...

The National Association for College Admission Counseling Commission — led by William Fitzsimmons, dean of admission and financial aid at Harvard University — issued a report last month the sparked a lot of conversation about how tests like the ACT and SAT were being misused.

This is a recurring theme in the college admissions business but this year's report warned that the present discussion of standardized testing has come to be “dominated by the media, commercial interests, and organizations outside of the college admission office.” Some of those groups have other items on their agenda.

Skip Kifer mentioned the report's warning about over-emphasizing test scores and argued for prudent use of the ACT in selecting students. He also warned folks not to get snookered by ACT officials new claims that without having changed the nature of the test, their scores now tell whether a student "meets expectations" or is "ready" to attend college.

As Kifer pointed out,

The benchmark stuff is statistically indefensible. Hierarchical Linear Modeling(HLM) was invented because people kept confusing at what level to model things and how questions were different at different levels. The fundamental statistical flaw in the benchmark ... is that it ignores institutions. Students are part of institutions and should be modeled that way.

Innes persisted and tried to play it off saying,

I guess Kifer and his compatriots at Georgetown College ... will never get the idea behind the ACT’s Benchmark Scores. It really isn’t hard to understand the Benchmarks.

This is when I should have suspected an academic spittin' contest was on the way and somebody was going to have to put up or shut up. (Actually, in the blogosphere, nobody really shuts up, but you know what I mean.)

Stung by the suggestion, and always the professor, Kifer challenged Innes to "Describe the statistical models used to determine the ACT benchmarks." This question would show whether Innes understood the nature of Kifer's argument and at the same time would disprove his claims.

Innes said he was waiting for some information from ACT (which in my experience is a lot like waiting for Godot) and changed the subject to how lousy the CATS is.

We're still waiting to hear him defend his position that the ACT Benchmarks constitute a valid criterion-reference test and that "the Benchmarks can fairly be considered a real measure of proficiency." That may be like waiting for Godot as well.

KSN&C has taken the position that all social science tests are imperfect. This includes the CATS, the KIRIS, the ACT, the SAT, the CTBS, the NAEP, and the EIEIO.

Tuesday, October 14, 2008

Kifer on the Problem with an ACT-only Assessment

KSN&C has maintained that what's really wrong with Senate Bill 1 and its suggestion of using the ACT as THE Kentucky test, is that using the ACT doesn't fix the problem. As one part of a comprehensive testing program the ACT may have a place. But alone, it fails.

Sunday, Skip Kifer illustrated the point in the Herald-Leader.


Narrow view of student potential
By Skip Kifer
We old guys remember what the ACT did in its heyday. Along with a student's high school record, ACT scores helped admissions offices decide who had the best chance of succeeding in their institutions. Colleges used ACT scores, along with an array of other information, to place students in courses once they arrived on campus.

Last month, a committee representing the nation's admissions officers released a report about overemphasizing test scores and argued for prudent use of the ACT and the Scholastic Achievement Test in selecting students. The committee suggested that some colleges and universities can enroll students without requiring the tests. This is a recurring theme, not a new one.

What is new is that ACT officials, without having changed the nature of the test, say their scores now tell whether a student "meets expectations" or is "ready" to attend college.

Kentucky's high school juniors are required to take the ACT. Their results, released a couple of weeks ago, were interpreted to mean that too many students failed to "meet expectations" or too few were "ready."

What, we old guys wonder, could these assertions mean?

An ACT score is just that: a score on a test. It tells something about me as a student but means little without a context.

My 18 on the ACT means something very different if I am in the top 10 percent of my high school graduating class rather than just in the top half. My 18 means something different if I have taken solid courses rather than choosing an easy way to a diploma.

My 18 means something different if I were to take the test again. My 18 on the ACT means something different depending on the college or university I attend. Among Kentucky's public universities, for example, my 18 would be a relatively low score in one but a relatively high score in another.

My 18 is like Shaquille O'Neal's free-throw percentage. Being able to shoot barely 50 percent from the free-throw line says something about O'Neal. But, does it mean not "meeting expectations" or not being "ready" for the NBA? Hardly. NBA scouts are clever people who would not let one piece of information become a judgment.

ACT says that a person with an English score of 18 has a 50/50 chance of getting a B or higher in a college English course. We old guys say that is throwing the statistical dice.

My left foot is in boiling water and my right foot is in ice water but "on the average" I feel fine.

Of all students with a score of 18, one-half will get a B or better and one-half a C or worse. In which half is my 18? Is it in the ice or boiling water? ACT cannot answer that question. But if I were in the top 10 percent of my class, it is more likely that I will get the B or better.

We old guys also wonder why so much was made of ACT results and so little reported about Kentucky's Advanced Placement results. As opposed to the ACT, AP scores are a direct measure of whether a student can do college work.

Students take AP courses in a variety of subjects. An AP course is so well defined that what Kentucky students experience is comparable to what is experienced by students throughout the country. Each AP course is taught by a capable teacher and is described in detail with course goals, materials and examinations available for scrutiny.

The examination measures what is learned in the course. High-school students who score 3, 4 or 5 on an AP examination can get placement or credit or both in the college or university they attend.

Recent results show increased numbers of Kentucky's students taking AP courses and scoring higher on them. That is, more Kentucky students are doing acceptable college work in high school. It would make sense, therefore, for an educational community wanting more students to attend college to work to expand AP opportunities.

While it is important to provide opportunities to prepare each student to take AP courses, it is unlikely that all students will take them. Even so, AP courses point in a positive direction for schools. They are a model for structuring courses and tying together testing and instruction.

An algebra course, for example, in Fayette County should be the algebra course in Christian County. More important, two algebra courses in one Fayette County school should be the same algebra course.

AP results show, not surprisingly, that students tend to learn what they are taught. It is important, therefore, to define and teach what it is important to learn. Then one can assess and give credit.

Someone might ask the old guys what this has to do with O'Neal. The answer is clear. If you want to know what kind of basketball player he is, you let him play basketball. You don't look only at his free-throw shooting percentage and decide that he does not meet expectations.

Thursday, May 22, 2008

Kifer says EXPLORE and ACT Overstate Claims

Skip Kifer of the Center for Advanced Study of Assessment at Georgetown College released a report this month that calls into question claims made by the ACT program regarding the efficacy of its EPAS system.

Kifer has a history with the ACT... and its a good one. In fact, he has to have been considered a fan. In Mental Measures Yearbook (Kifer, 1985) he lavishly praised test-maker ETS and the technical construction of the ACT. In fact, he found their practices to be worthy of admiration and emulation.

But times change. Since states have begun requiring exams that are specific to state curriculum, ETS has entered the state market and made claims about the efficacy of the EPAS system. It is these claims Kifer examines.

The study was motivated by the introduction of Senate Bill 1 during the 2008 legislative session. The bill included additional reliance on off-the-shelf standardized tests. Since the state assessment already contains such measurements under the EPAS system, a relatively new ACT program, Kifer set out to determine whether the program lived up to its claims.

Kifer looked at the characteristics of EXPLORE and ACT, two components of EPAS taken by 8th, 10th and 11th or 12th grade students, to determine whether or not the claims made for it and them are valid. He also compared ACT predictions to those obtained from other components of the Commonwealth’s assessment.

Kifer concludes that claims for the efficacy of the EPAS are stronger than the findings would suggest.

For example, EXPLORE lacks a representative national comparison group and therefore lacks true national norms. EXPLORE was constructed on the basis of national information, and therefore does not represent the Kentucky curriculum. Kifer also says,

"Although the technical manual suggests that the tests measure 'higher order thinking skills,' there is no dimension for those attributes in the test construction descriptions and no indications of how many questions of that type are on a test. EXPLORE is alleged to be diagnostic but is diagnostic in the limited sense that scores on EXPLORE predict scores on the ACT. Finally, there are validity studies of Kentucky’s assessments and ACT scores that suggest that the Commonwealth’s assessments are comparably effective predictors of performance in higher education."

Kifer concludes that EXPLORE and ACT do not live up to their claims. It also shows that other measures in the Commonwealth’s assessment do about as well and sometimes better than the ACT tests in predicting grade point average in college.

What this means, says Kifer is that,
If a student were certain that he or she was going to attend a public university in Kentucky, it would not be necessary to take the ACT. That student’s CATS score would be sufficiently predictive to take the place of the ACT. A consequence of using CATS scores in a college admissions process might motivate high school students to do their best on the statewide assessments. Finally, not requiring the ACT could save the Commonwealth a substantial amount of money.