Showing posts with label Skip Kifer. Show all posts
Showing posts with label Skip Kifer. Show all posts

Wednesday, March 9, 2011

Which Gap to Close

By Skip Kifer

Both No Child Left Behind (NCLB) and Kentucky's Senate Bill 1 (SB1) refer to achievement gaps and include expectations for closing them. As defined by Kentucky's Senate Bill 168 (SB168) and NCLB, gaps are differences in test scores based on gender, disabilities, limited English proficiency, ethnicity, or socio-economic status. Each school in the Commonwealth - given a set of rules about what is a gap - is expected to minimize differences between those groups, thereby closing it. It is implied that one expects each student's score to increase but those with lower scores are expected to increase by greater amounts. There is a desire for overall improvement in test scores as well as improvement in closing a gap.

As I write, there has been no reauthorization of the Elementary and Secondary School Act (NCLB) so there is no decision about how to define a gap or how to decide whether a gap is closing. To my knowledge, no decision has been made about how the implementation of SB1 will deal with those issues, either. I expect, for reasons discussed below, NCLB definitions to change. My guess is that Kentucky's might also change.

NCLB now defines an achievement gap as a difference between the percent proficient for one group, say girls, versus that of another, say boys. For a school to close the gap, it must reduce the differences between the two percentages. For SB168, Kentucky initially used a complicated average difference to look at closing the gap. That is, a school closed an achievement gap when it reduced, by a set amount, the weighted average between groups of achievement differences across grade levels and content areas.

In what follows, I hope to point out the strengths and weaknesses of both approaches and then suggest a third alternative for consideration. It is so easy for one to mouth the words "closing achievement gaps" without being aware of the technical difficulties of defining the gap and knowing either when it exists or when it has been closed. As a way to discuss the issues, I created data[1] and drew pictures of them.

Figure 1. Six representations of an achievement gap.

Figure 1 contains six pictures of the data. The graphs depict comparisons for one grade level and content area; for example, fourth grade reading. Three pictures (A,C,E) on the left are ways to show shapes, centers and spreads of the data. Three pictures on the right (B,D,F) are ways to show the gaps across levels of the test scores. Pictures A&C and B&D are the same but will be used to describe different features of the data.

Centers, Shapes, Spreads - Averages as Gaps

Figures 1A and 1C compare two groups, one of which is four times larger than the other. The size difference could happen if, for example, one was comparing majority students to minority students. Such differences in size do not affect the ensuing discussion. The groups could be of equal size, too. These dotplots are just detailed histograms that better represent the shapes and spreads of the distributions. A reader should see several things in Figure 1A: the distributions overlap substantially, the shapes are rather similar; the spreads are similar; but, the centers are different. The bottom distribution is shifted to the left indicating lower average performance for Group 2. That average difference could be a measure of the "achievement gap."

Figure 1B is another way to describe the data. This is a particularly good way to view cut-points that are used as the percent proficient goals. The lines I added to the figure are guides to interpreting the data. These curves depict what parts of a score group are at or below certain values. For example, if one follows the lines, one can see that fifty percent of Group 2 students score at or below 35. The comparable number is 40 for Group 1, the higher scoring group. The differences in those percents is the measure of the gap when the cut-point is 40 (i.e., 40 represents the goal, the desired percent, the percent proficient). One's eye can see different achievement gaps as the curves move from about 10 to 70.

Percent Proficient (Cut-points) as Gap Measures

In Kentucky there are three major cut-points, producing four major scoring categories - Novice, Apprentice, Proficient, and Distinguished. NCLB requires at least three categories of performance and that percent proficient be the cut-point for determining gaps.

There are several desirable properties of defining the gap in terms of cut-points.

  1. There are several well-defined, judgmental methods to define the cut-points, i.e. what will be called a proficient performance.
  2. Given the defined cut-points, it is straight-forward to calculate the gap and changes in the gap. This is especially true for summing across grade levels and content areas within a school.
  3. Coupled with a long-term goal of each student being proficient, the gaps are eliminated when the goal is met.
  4. The notions of being proficient in a subject area and having the percent proficient be the indicator of success, are easily conveyed to a broad audience.

There are several undesirable properties as well.

Perhaps the most serious one is depicted in Figure 1D. It shows that if the cut-point is at 40 rather than 50, the gap will be almost double the size. That is, the size of the gap varies according to where a cut-point is placed. Since the methods used to determine cut-points are judgmental, there is no one logical, well-defined place on the scoring scale to place a cut-point. That is a major reason why different states have different percents of students who are proficient.

Another weakness of cut-points as proficiency standards is that if those in the school wished to "game" the system, it is clear how that might be done. A gap can be narrowed by dealing with only a small proportion of the students. One should focus on students in the lower scoring group who are below but not too far below the cut-point. When they are moved to or above the cut-point, the gap is narrowed despite the performance of lowest scoring students. So differences in the percent proficient can be minimized by working with relatively few students.

Conversely, a school could increase dramatically the scores of the lowest scoring students without having an impact on the percent proficient. Imagine moving each student below the cut-point closer to the cut-point. Although the accomplishment would be dramatic, it would have no impact on the percent proficient.

The combination of using cut-points with a rule that each student must be proficient in a certain amount of time, gives a school an impossible task. Figure 1C shows where the cut-points of 1D fall on the score distributions. When the percent proficient is at a score of 50, 90 per cent of students in Group 2 must be moved to or past the cut-off. For Group 1 which is four times greater than Group 2 more than 80 percent of students must be likewise moved. When the cut-point is lower, the task is less onerous, about 70 and 50 percent respectively. I know of no empirical results that show such dramatics effects.

Finally, the whole idea of being proficient may be illusory. Simply placing a label on a test score does not make it true. Tests labeled science, for instance, may be very different kinds of tests. The science portion of Explore, the ACT eighth grade test contains only multiple choice questions and requires an inordinate amount of reading. The National Assessment of Educational Progress (NAEP) eighth grade science contains constructed response and extended constructed response questions and tends to minimize the effects of reading. Whatever proficient may be, it is likely to result in substantially different definitions depending on what science measure is used. And they both are wrong!

Mean Differences as Gap Measures

Just as for cut-points, defining achievement gaps in terms of mean differences have both desirable and undesirable properties. The positive aspects of such a definition include:

  1. Given data that are approximately bell-shaped the mean is a good typical value;
  2. As opposed to a cut-point definition where not all students are affected, the mean takes into account all cases.
  3. An average is a number most persons understand.

But, as I tell my students "never a center without a spread." Figures 1E and 1F show the effects on differences between groups when the spreads differ. The difference between the figures is about 2 1/2 points, a standard deviation of 10 for the first four and between 7 and 8 for the last two. The differences in the cumulative distributions get rapidly "fatter" above the mean of 40 (incidentally, the area between cumulative distributions is equal to the difference between means for the two groups). Minimizing differences when spreads are small may mean something different than when they are large.


Because decreasing mean differences may mean different things depending on the spread of data, it creates interpretation problems across grade levels and content area. Unlike summing percents based on cut-points, there is a question of how one should sum the effects to get an overall school index.

It is possible to "game" the means, although effects may be smaller than what one gets when gaming the cut-point definitions. If one believes, for example, that there are faster and slower learners, then to focus on relatively fast learners in the lowest scoring group could provide bigger gains that focusing on each of the students.

Finally, if it were just a matter of reducing differences between means, there would not necessarily be improvement across the system. So, there should be some specification of an expected amount of improvement.

Effect Sizes and Mastery Learning

An effect size, classically defined, is the mean for a treatment group, minus the control group mean, divided by the control group standard deviation.

This standardizes mean differences making them interpretable in terms of standard deviation units. The general idea can be used in the context of gap differences. For the data I have displayed, Group 1 has a mean of 40 and Group 2 has a mean of 35. Using the larger group's standard deviation of 10, we come up with an effect size of .5, that is, Group 1 performance is on the average 1/2 of a standard deviation higher. That magnitude of effect often would be interpreted as a medium sized.

These effect sizes can be summed over content areas and grade levels in a school to produce a school index. It would take some empirical work to decide how much the index should be reduced in order to say that an achievement gap is closing.

Although effect sizes respond nicely to the question of different spreads they do not help when it comes to different shapes. When Ben Bloom in 1967 outlined the properties of his approach to Learning for Mastery, he recognized the problem of only dealing with average improvement. So his goals included not only influencing average performance but also influencing the spread and shape of performance. The goals are to raise the mean, minimize the variance, and skew the distribution! A desirable outcome, then, is a heavily positively skewed set of higher scores rather than ones that look bell-shaped.

I don't know of anyone who has argued for reducing spreads and creating positive skewness as measures related to closing the achievement gap. Perhaps someone should. It may be worth a look.

Conclusions

If I were to decide what to use as indicators for defining a gap and determining whether it has been closed, I would not use either a method based on cut-points or simple mean differences. I would start with effect sizes and then do some analyzes to determine whether indicators of reducing variation or creating positively skewed outcome data are other possible measures.

What ever measure is chosen, it should be grounded in empirical results. So, there is a major task for the assessment persons in the Kentucky Department of Education to analyze their assessment data and come up with defensible suggestions for measuring a gap, measuring how much it changes, and how much it must change before deciding that the gap has been reduced.


Caveat

I have tried to respond directly to the gap issues without divulging my reluctance to base decisions about what is a good or effective school simply on the basis of test scores. Or, for that matter, whether schools should be held accountable for "gaps" that are based only on test scores. There is what I consider a naive view that backgrounds of students should be ignored when looking at whether schools are effective. At the same time there is an almost religious belief in the efficacy of test scores as the way to determine whether a school is good. Such views defy common experience and ignore research about schools and schooling. Some schools, for example, have relatively small amounts of turnover during a school year; others turnover almost completely. Some schools have huge amount of parental participation; others have virtually none. And, it remains true that the strongest within country correlations with test scores in international studies are based on the background characteristics of students.

The effects of schooling are many, diverse, desirable and undesirable, both short term and long term. Tests get at a small number of similar, desirable, short term effects. NCLB ignores most content areas in judging schools. The Commonwealth's assessment measures fewer than half of its goals. What ever happened to self-sufficiency, effective group membership, and integration of knowledge?

Tests do not get at whether a school produces persons who are thoughtful and reflective. They do not get at whether persons are well-informed. They do not get at how well persons work together or how they well they respect other persons and other points of view. They do not get at whether a school produces good citizens. Good schools do all of the above! Those things are as worth thinking about as is the achievement gap, however defined.

[1] I produced these data. They do, however, mimic those I analyzed for a paper on the gap.

Measuring the Gap under SB 1

The new SB1 test got a first reading before the Board of Education in February and is scheduled for a second reading in April. We are in a 60-day window set aside for public comment, so let's talk about it.

One of the issues that has worried me is, How should the achievement gap be measured?

This is important because we know that however the state calculates it, teachers will plan their strategies based on whatever focus is likely to provide the best test score results. One approach might induce teachers to focus on a small subset of the kids who are said to be “in the gap.” A different approach would broaden that focus.

KSN&C spoke to KDE testing guy Ken Draut, at the AdvanceEd Conference in December, and learned of KDE’s plans to use ACT benchmarks to measure the gap. Uh oh.

Generally, KDE is looking at what they call a balanced accountability approach.

Achievement score data would come from five days of testing with the new SB1 Test for grades 3-8 built around the new Kentucky Core Achievement Standards, as they are implemented. Scores will be calculated around “proficiency,” meaning that cut scores will be set to determined performance levels. Novice = 0; Apprentice = 0.5; Proficient = 1.0; and Distinguished = 1.5. The + .5 amount given to students scoring in the Distinguished range is thought of as a Bonus, and it will be offset by a negative .5 for each Novice student in the group.

But what about measuring the achievement gap?

Draut told conference attendees that “we really feel like, in the gap world, we got a really innovative model to measure gap.” Draut explained that KDE had three problems with measuring the gap which they have tried to address.

  1. The number of different sub groups, which can number to as many as 45 different goals.
  2. A lot of the kids fall into more than one group. For example, 80 percent of Kentucky’s African American kids are receiving free and reduced lunch. 80 percent of our ELL kids are in the free and reduced lunch group. 70 percent of our special education students receive free and reduced lunch. As a result, we end up counting one student multiple times. So under the present system, a single student might be counted four times. Miss one target and you are likely to miss four.
  3. Comparing “closed gap to group.” Draut said, “You want to close African American to White; American Indian to white… Because of the way testing works, you can end up with a “wavy” pattern.” For example, in JCPS one year, the white kids at Southern Middle School dropped backwards and the African American kids stayed the same, the Courier-Journal reported that Southern Middle was closing the gap. And the opposite can happen (which was our experience at Cassidy) where white kids can go up 8 points and African American kids go up 6 points, but the gap increases.

To address this, KDE plans to present all of the gap students’ data, but will create a new single group of underperforming “gap kids” who would only be counted one time. Instead of comparing the gap kids to the group, it would be measured against the goal of 100 percent proficiency, or what is called “gap to goal.” The gap is to be divided by the number of years schools are given to reach their goal and schools would be awarded points based on the percentage of that goal they were able to close. A school that had a six goals to meet and closed three of them would earn 50 percent of their points.

Growth is to be measured by using a regression of the reading and mathematics scores, the only tests given every year from grades 3-8. It will compare a student’s progress to other students who have been performing similarly. Given a proficient 5th grade student with a scale score of 230, who then earns a score of 240 in 6th grade: the model asks if this level of growth is typical of other Kentucky students, above average, or below average, and awards points accordingly.

KSN&C caught up after his presentation.

KSN&C: Ken, as you may be aware, a number of statisticians…like Skip Kifer, say that the modeling that underlie the statistics of the ACT Benchmarks are a bunch of crap, basically. [chuckles]

Draut: Right.

KSN&C: Are you concerned about that?

Draut: Well, this is how we answered the board the other day: It’s you guys. You guys drive this. If you say the ACT is a bunch of crap, lets’ throw it out…

KSN&C: Well, not the ACT. Just the benckmarks.

Draut: Well, I’m just saying, if you all say it, and then you put something else in, we’ll line right up, because we’re trying to get them ready for you.

KSN&C: OK, but you lost me. Tell me who “you” is. Because you’re saying the board…

Draut: Universities.

KSN&C: Oh.

Draut: You see, we’re driven by the universities. We can’t get our kids into the universities unless we meet your criteria.

KSN&C: So if the universities say, this standard isn’t appropriate, or the metric’s wrong, or something, then that’s going to be a problem for you guys.

Draut: Well, we’ll put in whatever you say, but I tell you, what the issue is, and we’ve said this to several people, tell us what you’d replace it with.

KSN&C: Uh huh.

Draut: Just tell us.

KSN&C: So, the benchmarks are useful, because they are there…But you have to know what they mean or they’re meaningless. And you can’t replace it with the ACT really, because that cuts out middle school and…causes you some other problems.

Draut: Right.

KSN&C: So, then what do I replace it with. I’ve got a bad yardstick, but it’s the best one I’ve got?

Draut: And what are the universities going to accept to get the kids in the door? Because whatever the universities accept, that’s what I’ve got to get my kids ready for.

KSN&C: Are you getting that kind of pushback from the universities?

Draut: No

KSN&C: So the question’s been raised but nobody’s pushing the issue?

Draut: No. It’s kinda like just what you said, tell me what’s in its place?

KSN&C: And nothing comes to mind.

Draut: So now you open up fifty years of research saying, hey, we can tell you it works. It does predict…

KSN&C: Do we know the degree to which those benchmarks are bad, or in what direction they are bad? Or is it that we just don’t know?

Draut: I think that you’d have to do some reading, both the pro and con, when I read, and I’ve heard Skip, but when I read the ___of it, it makes a lot of sense. And when I hear Skip it makes sense, too. I can’t get a sense of which one’s right…But that whole issue is driven by CPE and the universities because if you’re sitting there in the university saying we’re only going to take the kids that make the CPE benchmark, and we’re only going to take the COMPASS, then we say, OK, and we line up with you. But if universities change…and say, you know, we’re not going to use ACT, we’re going to use some new testing, then we’ll realign everything [to that]. ..But I think it would be useful to look at both the pro and the con.

For a few months now, I've been pondering Draut's position that decisions made at CPE should drive the model ultimately adopted by the Kentucky Board of Education. Generally I agree that we can not lower standards and KDE must hit college-ready targets. But I'm much less convinced that CPE ought to dictate how the achievement gap in measured in our elementary and middle schools.

NOTE: It is my understanding that the EXPLORE can predict results on the PLAN test, but not the ACT. The PLAN test can predict performance on the ACT but not performance in college. The ACT can predict performance in college up to a point, and its arguably not the best way, but is made better by the inclusion of other measures.

Is the ACT the Sole Criteria?

It all started when Council on Postsecondary Education President Bob King argued in the Herald-Leader that CPE's High School Feedback Report allows the state "to look more deeply into actual performance measured by an external, unbiased resource — the ACT exam."

King suggested that the ACT, by itself, was superior to predictions of college-readiness derived from combinations of data. Is CPE discounting graduation rates and average GPA in favor of a single test?

That drew a response from KSN&C's Skip Kifer, demonstrating the weak relationship ACT musters, and suggesting that CPE should propose placement procedures based on a robust notion of "readiness" that includes more than just the ACT and that did not violate test score use standards. It appeared to Kifer that CPE arbitrarily uses that single test score to determine whether a student is ready for regular course work in Kentucky's public universities.

Too clarify, King stated that CPE does not rely exclusively on the ACT to make college admission or placement judgments, nor does the Council on Postsecondary Education encourage such determinations.

Perhaps King should have reviewed CPE's printed material before saying that.

In last Saturday's H-L, Kifer wrote,

I am puzzled by Bob King's response to my critique of the Council on Postsecondary Education's policy of declaring a student college ready on the basis of a single test score.

King, director of the council, says: "Please allow me to clarify that Kentucky's colleges and universities do not rely exclusively on the ACT to make college admission or placement judgments, nor does the Council on Postsecondary Education encourage such determinations."

Yet when I look on the council's Web site, it says:

"The Kentucky statewide public postsecondary placement policy in English and mathematics applies to any student entering a Kentucky public college or university. The policy is based on your ACT or SAT score and determines what type of English and math classes you will need to take when you enter college."
This is what I found in my search. Notice the date posted. That's the day Kifer's piece ran. Maybe I missed something, but I didn't catch any changes in the language from when I looked at the same material in February.

StatewidePlacementPolicy: Postsecondary Placement Policy does … Postsecondary Placement Policy, please … PLACEMENT POLICY IN http://cpe.ky.gov/nr/rdonlyres/73e9a7b3-84dc-4ec2-8f1b-6a99261b5fb4/0/statewideplacementpolicy.pdf
- 120KB - kdrummond - 3/5/2011 [View duplicates]

And I found this:
The statewide placement policy is applicable to any incoming student entering a Kentucky public postsecondary institution. ACT and SAT standards form the basis of the policy because Kentucky uses the ACT (or equivalent measures) for college admissions and placement decisions.
That language says ACT forms the basis for placement decisions but stops short of saying the ACT is the only determinant. That comes next.

Kentucky Statewide Placement Policy in English
• A student earning an ACT English sub-score of 18 or higher qualifies for placement in a credit-bearing writing course at any Kentucky public postsecondary institution.

Kentucky Statewide Placement Policy in Mathematics
Three levels of readiness are identified for placement in a credit-bearing mathematics course at any Kentucky public postsecondary institution:
• Level 1: A student earning an ACT mathematics sub-score of 19 or higher qualifies for placement in a credit-bearing mathematics course, but this course may not be a requirement for many college majors or lead to subsequent coursework in mathematics. Mathematics for liberal arts is an example of such a course.
• Level 2: A student earning an ACT mathematics sub-score of 22 or higher qualifies for placement in college algebra. College algebra (or placement in more advanced courses) is required for majors such as biology, business, economics, information systems, and technology. College algebra can lead to any major.
• Level 3: A student earning an ACT mathematics sub-score of 27 or higher qualifies for placement in calculus. Calculus is required for majors such as mathematics, physics, chemistry, computer science, engineering, biology, business, and technology.

Kentucky’s statewide public postsecondary placement policy is a guarantee of
placement in credit-bearing coursework to incoming students demonstrating
specified levels of competence.

Monday, February 14, 2011

King Clarifies Stance on ACT and College Admissions

Council on Postsecondary Education President Bob King recently offered his opinions on college admissions saying,

Historically, parents are often directed to focus on graduation rates and average GPAs as evidence of how their high school is performing. Our reports allow parents and educators to look more deeply into actual performance measured by an external, unbiased resource — the ACT exam — now required of all Kentucky students.
That drew a response from KSN&C's Skip Kifer.

Bob King, president of Kentucky's Council on Postsecondary Education applies the council's arbitrary standard of using a single test score to determine whether a student is ready for regular course work in Kentucky's public universities.

He implies a test score is a better predictor of grades in college than is a high school record. He then presents results from one high school that lump higher performing students (those with above-average high school records) with lower performing ones in a misguided approach to justify his position. A test score, however, does not make or break a student's readiness for higher education.

Today, King clairifed his stance in the Herald-Leader.

...Please allow me to clarify that Kentucky's colleges and universities do not rely exclusively on the ACT to make college admission or placement judgments, nor does the Council on Postsecondary Education encourage such determinations.

A student's entire record, including GPA, extracurricular activities, and other placement exams form a portfolio that allows campuses to make informed decisions on admission and placement.

The ACT serves as an important element in this consideration, but more importantly, it serves as an alarm bell in the student's secondary experience about preparation for life after high school.

This might have been a good place to stop. It acknowledges Kifer's concerns and clarifies King's stance. But King then makes allusions to "certain thresholds" in the ACT which serve to warn us if a student is not on track.

It warns of the need to take a deeper look at a student's college readiness if the scores fall below certain thresholds, but it is not the sole determinant when placing students in developmental courses... Far from arbitrary cutoff scores, there is a great deal of data from tens of millions of ACT score results upon which policy makers in Kentucky rely to set the scores used to indicate college readiness in key entry-level courses.

If the thresholds King has in mind are, in fact, a reference to ACT's benchmarks, which are inappropriately modeled, one wonders if King's effort to lay the issue to rest might draw yet another response from Kifer.

We'll see.

Monday, January 31, 2011

ACT: Not the Only Measure of College Readiness

Council on Postsecondary Education honcho Bob King recently argued in the Herald-Leader that CPE's High School Feedback Report allows the state "to look more deeply into actual performance measured by an external, unbiased resource — the ACT exam."

King suggested that the ACT, by itself, was superior to predictions of college-readiness derived from combinations of data. Discounting graduation rates and average GPA in favor of a single test prompted our resident testing expert to retort.

NOTE: H-L seemed to struggle editing Skip's piece, so here's the unadulterated article the paper titled:

If I were to assert that a player who cannot make 56% of his free throws is not "ready" for the NBA, a fan would point out that there is much more to basketball than shooting free throws. An astute fan with a historic prospective would point out that Wilt Chamberlain, Shaquille O'Neal and a bevy of other current players would not be "ready" using that arbitrary standard. One facet of basketball does not make or break a player's "readiness."

Bob King, president of Kentucky's Council on Postsecondary Education, in a recent op-ed piece applies the council's arbitrary standard of using a single test score to determine whether a student is "ready" for regular course work in Kentucky's public universities. He implies a test score is a better predictor of grades in college than is a high school record. He then presents results from one high school that lump higher performing students (those with above average high school records) with lower performing ones in a misguided approach to justify his position. A test score, however, does not make or break a student's "readiness" for higher education.

Decades of research indicate: performance in academic courses in high school is the single best predictor of success in higher education; a combination of the high school record and test scores predict better than the high school record alone; and, how good the prediction is and how the components are combined vary depending on the institution. Although there is a general pattern of the primacy of the high school record, there is no one-fits-all model to predict grades in different courses or different institutions.

In addition to the thoroughly suspect notion of labeling a test score readiness, the council's use of a single score for placement purposes violates standards for the proper use of tests. Those standards include the following:

In educational settings, a decision or characterization that will have major impact on a student should not be made on the basis of a single test score. Other relevant information should be taken into account if it will enhance the overall validity of the decision.

In addition:

When test scores are intended to be used as part of the process for making decisions for educational placement, promotion, or implementation of prescribed educational plans, empirical evidence documenting the relationship among particular test scores, the instructional programs, and desire student outcomes should be provided. When adequate empirical evidence is not available, users should be cautioned to weight test results accordingly in light of other relevant information about the student.

Apparently, the council determines readiness by doing statistical analyses of ACT scores and grades in first year courses without regard to institution. A certain ACT score produces a 50/50 chance of getting certain grades, say C, or better. I could find no information about this or other investigations done by the council. And, although it is possible to present information about how good a model is, I could not find any information of that kind either.

To give a sense of the power of statistical models to predict first year grades I report analyses conducted years ago on University of Kentucky student samples. The question was whether results of KIRIS, the first commonwealth assessment related to school reform, could be used for admission and placement in a university.

If a model exactly predicts grades one can say that the model accounts for 100 % of what could be known. If a model cannot at all predict grades, one can say that 0% is accounted for. One way, then, to talk about the power of a statistical model is determine what percent the model predicts.
The table below gives those percents for different courses at UK for three different statistical models: High school record only, High School record + ACT scores and High School record + KIRIS scores.



The first thing to recognize is that the models are not particularly powerful. They rarely account for 25% of what could be known leaving 75% to be explained. That 75% may be differences in students' study habits, class attendance, interests, any of a thousand other variables or simply things not explained statistically.

The pattern of results, however, is clear. Adding an ACT score to a model containing GPA makes the prediction better but not greatly so. The same is true for KIRIS scores, too. Incidentally, that was without including the KIRIS writing sample.

These are not unusual results. They point, obviously, to gathering more information about a student before making a placement decision. Here is what ACT says:

ACT offers a variety of tools to ensure postsecondary students are quickly and accurately placed in courses appropriate to their skill levels. Assessment tools from ACT offer a highly accurate and cost-effective basis for course placement. By combining students' test scores with information about their high school coursework and their needs, interests, and goals, advisors and faculty members can make placement recommendations with a high degree of validity.

To that, I would add, for obvious reasons, it is desirable for an educational agency to use tests in exemplary ways.

The op-ed piece goes on to exhort parents to ask right questions, asks an undefined "we" to fear international test results, says admissions offices should align themselves with the council's readiness standards, and the still undefined "we" to serve teachers more effectively. Such exhortations would be more convincing if, in the first instance, the council could propose placement procedures based on a robust notion of "readiness" that in addition did not violate test score use standards.