What the Study Actually Did
Researchers ran a randomized trial across 16 Kenyan primary care clinics operated by Penda Health, covering nearly 10,000 patient encounters. Half the clinical officers used AI Consult alongside their electronic notes. The other half typed notes without it.
The AI—built on OpenAI’s GPT-4o—ran in the background and flagged issues using a traffic-light system:
- Green: Notes look fine, carry on.
- Yellow: Small gap or concern, click for guidance.
- Red: Critical issue detected, act promptly.
An independent panel of six Kenyan family physicians then graded the notes blind. The verdict: clinicians using AI Consult produced measurably better diagnoses and treatment plans.
Cost per patient: four cents.
Where It Helped—and Where It Didn’t
The process improvements are real. Better notes. Better diagnoses. Better treatment plans. That’s not nothing, especially in settings where clinicians are stretched thin and specialists are rarely in the building.
There was also a 23% decrease in treatment failures—deaths or unresolved symptoms. But that number didn’t reach statistical significance. The reason is almost counterintuitive: primary care works well enough that serious failures are rare to begin with. To detect a meaningful difference in outcomes, you’d need roughly 139,000 patients in the trial. This one had 10,000.
So the honest summary is: AI Consult improved the quality of clinical work. Whether that translates to better patient outcomes remains unproven—not because the signal isn’t there, but because the study wasn’t built to find it.
“Information as an Intervention”
Dr. Bilal Mateen, chief AI officer at PATH and a co-author of the study, frames the tool’s value as “information-as-an-intervention.” That’s a useful way to think about it. The AI isn’t diagnosing patients. It’s prompting clinicians to double-check their own reasoning.
Clinical officer Njeri puts it more plainly: it’s like having a superior who says, “you could do better there” or “you’re doing okay.”
She finds the recommendations helpful about half the time. The other half, less so—but rarely wrong. That’s a reasonable batting average for a four-cent background check.
The Broader Significance
Dr. Jonathan Chen at Stanford, who wasn’t involved in the study, called it important precisely because it goes beyond simulated testing. Most AI healthcare research never leaves the lab. This one ran in real clinics, with real patients, under real conditions.
That makes it one of the first randomized trials to test this class of AI as a clinical decision support tool in primary care—particularly in a lower-resource setting where the need is acute and the safety net is thin.
The Caution Worth Keeping
Not everyone is ready to celebrate. Dr. Nicholas Okumu, an orthopedic surgeon at Kenyatta National Hospital, raises a fair point: even approved AI systems can cause harm. The tool performed well here, but “oversight has to stay active”. That’s not a reason to avoid the technology—it’s a reason to deploy it carefully.
The study’s expert panel rated most of AI Consult’s recommendations as safe and appropriate. But a tool that’s right half the time and rarely wrong still needs a human in the loop who knows the difference.
The Takeaway
AI Consult didn’t cure anything. It made clinicians more thorough, more consistent, and slightly less likely to miss something on a busy Tuesday afternoon. In a setting where one clinician might see 40 patients a day without specialist backup, that incremental improvement compounds quickly.
The gap between “better process” and “better outcomes” is real, and this study is honest about it. But if you’re evaluating clinical AI tools, that honesty is itself a useful data point. Look for tools that improve the quality of work, not just the ones that promise to fix everything. The four-cent price tag doesn’t hurt either.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!