“Phonics is a fad, but it’s also good, so we should do it anyway, even though it will probably disappoint people”

Tyler Watts and Drew Bailey write:

Educational policy is often viewed as a field of fads, where one follows another for seemingly arbitrary reasons. We might like to believe that fad-based policy could be superseded by evidence-based policy, with rational actors choosing policies because of scientific evidence. But that’s been hard to achieve. In the cynical version of the story, hucksters overpromise on their pet projects, we make the changes they advertise, and they get rich while the public is left disappointed. This probably happens more often than it should, but this narrative ignores the conditions that make such disappointment ubiquitous in educational policy.

Why? They explain:

Even if the hucksters all disappeared tomorrow, we’d still have this hype cycle, because overpromising seems inevitable when effects are small, and political disappointment will always follow. . . . The educational policy research ecosystem usually selects for safe, often evidence-based, politically feasible, but ultimately oversold ideas, and we think that hurts our ability to make progress. In education, real effects are really small. . . . Why are the effects so small? The outcomes we like to measure, such as reading comprehension, depend on many personal and environmental factors, making it hard for a discrete change in schooling to make much of a difference. . . .

This is a well-known problem in education research. For example, it came up in my 2018 talk at the Society for Research on Educational Effectiveness, and the point wasn’t new then either!

But, as Watts and Bailey explain,

The extreme education-nihilist position—that educational improvements can never affect students’ skills—is actually wrong. Endorsing it is likely to lead to poor education policy. The edu-nihilist position overlooks that small effects are not always zero effects. Small effects can actually be meaningful and, in some cases, worth it from a cost-benefit perspective for society.

One way to think about this is to consider classroom teaching.

I work my butt off every semester teaching my classes. Every class takes a lot of effort, both by me and the students. But any given class can’t have much of an effect on total learning. Indeed, the entire semester doesn’t have such a large average effect, when you consider that some students already could do most of it before the semester began, some still don’t get it when the semester is over, and others learn some things which they later forget. (On the plus side, students can learn all sorts of useful things that don’t happen to be on the exam.) But the effect of these classes is not zero. Students take them and they learn some things. So of course there’s no simple switch that will improve performance by 0.2 standard deviations. We’re already doing what we can. And the flip side of this is that a small improvement is not nothing. We teach each class knowing that at best it is giving a small average improvement, but it’s still worth doing.

Watts and Bailey continue:

When education policy researchers say something “works,” most people assume that effects are much larger than they are. . . . In most cases, researchers and policymakers fall prey to effect size magnification because they develop real sympathy for a pet program. When this happens, they put forward the most promising evidence, which is likely biased. The field then develops unrealistic expectations, and researchers, policymakers, and advocates downplay or withhold information about their uncertainty.

Hey, here’s an example from celebrated economist James Heckman, who doesn’t have much of an understanding of selection bias and goes around promoting wildly exaggerated estimates of early childhood programs.

Here are Watts and Bailey on what happens next:

Educational policies are uniquely poised to fall apart when overhyped, because public scrutiny of educational outcomes is so intense. Because most of the variance in student achievement comes from factors schools can’t control, overselling sets up an inevitable letdown. When disappointing results emerge and large gaps in achievement remain, we overreact to the apparent failure. Feeling the heat, the researcher, policymaker, or advocate might retreat and say that whatever was promised was never supposed to be a “panacea.” This always feels cheap because the language rarely shows up in the press releases, grant proposals, white papers, or speeches preceding the policy. Ultimately, we are left abandoning ship, even if the policy in question was actually helping. The clearest example of this is probably the accountability reforms pushed by No Child Left Behind. Everybody hates No Child Left Behind because it was terrible and didn’t work! Right? Actually, our best policy evaluations suggest it did work; it just had small effects, which is exactly what we should have expected. Instead of taking the small effects as a sign that something is working and figuring out ways to improve it, politicians and advocates decided it was a failure and moved on to other things.

And then they turn to the latest trend:

In a country that can’t agree on much, states from Alabama to California have adopted policies designed to promote the science of reading. . . . Count us as phonics believers. Kids need to know how to sound out words when they learn to read. We taught our own kids how to read this way, and we ourselves even learned to read this way in the pre-science-of-reading dark ages of the 1990s. But the unfortunate reality is that we have no reason to believe a sudden shift to a phonics-based curriculum will significantly affect reading outcomes in this country. There are several reasons why. First, meta-analyses of phonics interventions show fadeout. This means that kids who get phonics instruction for some period of time look a lot like kids who didn’t get phonics instruction in terms of reading just a year or two after the intervention ends. Second, teaching phonics seems to help with decoding and word recognition. This is an important building block of literacy, but there isn’t much evidence that it helps with eventual reading comprehension—the ultimate goal. . . . Finally, we are fooling ourselves if we actually believe that phonics hasn’t been used in classrooms up to this point. Do we really think that it never occurred to teachers in the “woke” states to help kids break words into syllables and sound them out? The counterfactual condition here has become a strawman. . . . The best evidence on what happens when states adopt early literacy policies is probably this working paper, which finds effects of up to about .10 standard deviations . . . on high-stakes tests for 3rd graders (that’s good). They didn’t find much on low-stakes tests administered to 4th- or 8th-graders. So, if we enter a halcyon period for phonics instruction, we might help a lot of kids learn basic decoding skills when they are first learning to read. But . . . effects will be smaller than we were led to believe. Importantly, they probably won’t move the needle much on outcomes like NAEP trends, which compare U.S. cohorts of 4th and 8th graders to each other. . . . The question is: how will we interpret this when the sobering news starts to come in? Just in case we weren’t clear: doing more phonics in early-grade classrooms is, for the last time, probably a good thing.

They conclude:

Education policy is not the only field that faces conflicts between public and expert opinion. But the combination of big expectations, small effects, and constant attention to children’s educational outcomes means education policy is likely to keep experiencing these conflicts. What can we do to mitigate this educational political economy problem? . . . We should ask interested parties to make predictions on the record about how these children will differ (and by how much). And our society should reward researchers and policymakers whose mental models are closer to reality, rather than those who are best at selling potential solutions in the first place.

And my own thoughts

I agree with everything that Watts and Bailey wrote above, and also have a few things to add:

First, there’s my recent paper with Amy Krefman, Lauren Kennedy, and Jessica Hullman, Hypothesizing an effect size by considering individual variation, where we try to think systematically about how to get more realistic expectations of average treatment effects.

Second, a point from my above-linked SREE talk, that research should be less isolated from practice. From one direction, research studies should be more realistic, more connected to the real world of education—this is a point that Watts and Bailey make in their post. From the other direction, general practice should follow principles of research: students, parents, teachers, and administrators should be empowered to make choices and experiment, and, along with this, choices and outcomes should be measured so that effectiveness can be studied. We should all be trying to do better anyway, and we should design our data collection so we can all best learn from these experiences. As we discussed in a different context, local data collection, collaborative analysis, and local recommendations.

Third, ummm, this is a longer one so I’ll just quote myself:

You can conceptualize an education intervention as a vector, where the direction of the vector is the material being learned and the length of the vector is the amount that students are motivated to work. You want the material learned to be useful—you’d like the vector to have a positive “dot product” with the vector of skills, knowledge, and understanding that will be useful going forward—but, conditional on those two vectors being roughly aligned, the real gain is in the magnitude. And this magnitude will be an interaction between the teacher and student: there’s no button to push to create motivation, and if there were a button it would already have been pushed. What I’m saying is that, when thinking about acupuncture, or physical therapy, or coaching, or teaching, we have to go beyond what I’ve called the penicillin model of science, the idea that innovations come from nowhere and that the job of statistics is to design and analyze experiments to reject the null hypothesis of no effect, and in which the treatment in such experiments is considered as a black box, with the goal being to estimate an average treatment effect. I don’t think the penicillin model usually applies. Most of the time in health, education, and just about any field, improvements are incremental, and the goal is to improve the process while gaining understanding. Clinical trials and offline experiment both play a role, and you’re not going to learn much by studying a treatment as if it’s a black box. This is not to say that there cannot be new developments in any of these fields, nor is it to deny that such developments can sometimes arise serendipitously. I just think that, in any case, you have to go beyond the average treatment effect and think about the mechanism of action.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论