Opinion | When AI Tells a Story About Your Health, Is It True?
Smart-ring maker Oura is expected to go public this week, touting its promise to transform personal measurements into “trusted, actionable insights” for improving health. It’s a compelling premise, yet it has long struggled to survive contact with biology.
Two recent cardiovascular trials illustrate the point. Novo Nordisk’s ziltivekimab, an experimental drug designed to suppress inflammation, sharply lowered inflammatory biomarkers but didn’t reduce the risk of cardiovascular death, heart attack or stroke. Novartis’s pelacarsen lowered Lp(a), a lipoprotein strongly implicated in heart disease, without reducing cardiovascular events. These results highlight the complexity of biology and the challenge of predicting outcomes from measurements taken many steps upstream.
At first, this seems the sort of problem artificial intelligence could solve. In some domains AI performs spectacularly. AlphaFold predicts three-dimensional protein structure with extraordinary accuracy, drawing on decades of experimental information, strong biological constraints and an answer that can be evaluated directly. “What structure does this protein adopt?” is a more tractable question than “What will happen to this patient if we dose this compound?”
Most new technologies traverse a hype cycle. Even before the human genome was fully sequenced, the biologist Sydney Brenner said that with a complete DNA sequence and a sufficiently large computer, he could “compute the organism.” As genetic assessment became more accessible, startups promised genetically informed guidance on what to eat, what to drink, even whom to date. Most of those ambitions proved fanciful. Yet genomic technologies have transformed disease research and made contributions to clinical medicine, particularly oncology.
Stung by disappointment, seasoned drug developers approach AI with only cautious enthusiasm. But the aspiration to deliver tailored health advice to consumers is back with a vengeance, now equipped with a richer instrument panel and AI to connect the dots.
Sometimes this can be useful. But richer data, combined with AI’s ability to integrate and interpret them, also create a risk of false precision—recommendations that feel more individualized, and more certain, than the evidence warrants.
A continuous glucose monitor worn by a nondiabetic can produce a detailed account of post-meal glucose. But in this setting the measurements are less reliable, responses to the same food can vary substantially, and there is little evidence that minimizing these fluctuations improves long-term health. Wearables like Oura and Whoop (also reportedly planning an IPO) pose a similar problem: They combine heart rate, heart-rate variability, sleep and other signals into proprietary “readiness” or “recovery” scores, generally using unvalidated formulas, with little evidence that following the result improves performance, prevents injury or benefits health.
Generative AI is remarkably good at turning ambiguous, incomplete information into a coherent account. Give it your declining heart-rate variability, shortened sleep, increased mileage and stressful Tuesday, and it can fashion a plausible story connecting them, even when the explanation remains uncertain. As Daniel Kahneman observed, “No one ever made a decision because of a number. They need a story.” Personalized explanations can be seductive because they arise from our own data.
If we rely on AI to narrate our experience—interpret our numbers, synthesize our sensations, assign meaning to our setbacks—we may weaken our own capacity to make sense of it. Then again, a well-designed AI could suggest alternative interpretations, make uncertainty visible, and help us cultivate rather than surrender that capacity.
Work by Stanford psychologist Alia Crum suggests why this may be consequential. In a classic study she began as a Harvard undergraduate, hotel housekeepers were told—accurately—that their daily work already constituted meaningful exercise. Over the following weeks, investigators reported lower weight and blood pressure despite no change in their actual activity. More-recent experiments taught participants to recognize that their initial interpretation of stress was only one possible frame; they subsequently reported fewer physical symptoms.
That points toward a different opportunity for AI in personal health. Where the data support a consequential individualized insight, of course AI should report it. But where the evidence is thinner, AI could help us examine what we’re experiencing, consider alternative meanings, and decide for ourselves what to carry forward and how to act on it. Used this way, AI could strengthen rather than displace the judgment and self-belief that underlie agency.
It feels almost like poetic justice: subverting AI’s narrative fluency by enlisting it in service of our own storytelling, helping us become more intentional—and healthier—authors of our lives.
Dr. Shaywitz is chief medical scientist at Lore Health, a lecturer in the Department of Biomedical Informatics at Harvard Medical School, and an adjunct fellow at the American Enterprise Institute.