The Honest Machine: Building Medical AI to Trust | The Uncertain Eye Ep. 8
Varun explains how rebuilding the evaluation with strict patient splits and confidence intervals produced a trustworthy 0.86 AUROC glaucoma model and why hand-computed features added nothing.
Varun presents the rebuilt version of a glaucoma-detection system in this episode of The Uncertain Eye, after catching a flaw in the earlier evaluation. The fix wasn't a smarter model. It was a stricter, more honest way of testing the one they already had.
Why the evaluation had to be rebuilt
The team had previously caught their own testing trap, so they rebuilt the whole evaluation pipeline slower and stricter. The core change was keeping training and held-out test data cleanly separated by patient, so no patient's scans could leak between the two sets. That single change closes off a common way medical AI results look better in testing than they turn out to be in practice.
Confidence intervals instead of a single lucky number
Rather than reporting one accuracy figure, the rebuilt system reports a confidence interval, a range that shows how much the result would shift if the evaluation were run again. Varun frames this as replacing a single lucky number with an honesty band attached to every result. Their clean OCT-only classifier lands around 0.86 AUROC with that band attached, a number they present as one that can actually be trusted, rather than one that happened to look good once.
Testing whether hand-computed features actually help
The team also tested a common shortcut: fusing the raw OCT image with hand-computed structural measurements, the kind of numeric features a specialist might calculate directly from a scan. The result was that these engineered features added almost nothing. The network had already learned that structural information directly from the image itself. Varun calls this a quietly important result: knowing that a technique doesn't help is real progress too, not a wasted experiment.
The open frontier: reasoning models on hard cases
The next test is still running. The team is checking whether today's most capable reasoning models, given only the honest structural signal, can deliberate over the hard cases the way a specialist would and explain every call. Varun is direct that there's no final answer yet, and frames that openness as what ongoing research actually looks like, a careful climb rather than a triumphant reveal.
The vision: a tireless first reader
The episode closes with the goal this rebuild is aimed at. A patient gets a quick, cheap OCT scan anywhere. A fast model clears the easy cases and flags the uncertain ones. The flagged cases go to a reasoning model that lays out the evidence, then to a human specialist with the case already worked up. The machine isn't positioned as the doctor. It's framed as a tireless first reader that never gets bored, never skips a scan, and knows the edge of its own knowledge, buying back time against a disease Varun calls a silent thief.
Key takeaways
- Splitting training and test data strictly by patient prevents a common source of inflated results in medical AI evaluation.
- Reporting a confidence interval instead of a single accuracy number shows how much a result would wobble on a repeat run.
- The rebuilt OCT-only glaucoma classifier reaches about 0.86 AUROC with that honesty band attached.
- Adding hand-computed structural features on top of the image did not improve performance, since the network had already learned that information.
- The next open question is whether reasoning models can deliberate over hard cases and explain their calls, and that test is still in progress.
Who this is for
Viewers following The Uncertain Eye series on AI-assisted glaucoma detection, and anyone interested in how rigorous evaluation practices, like patient-level splits and confidence intervals, separate trustworthy medical AI results from ones that only look good once.
Chapters
Full transcript(auto-generated, with timestamps)
Rebuilding the model honest to the bone
[0:00]Hi, I'm Varun. In this video is about the honest machine, the rebuild that made the result hold up. Having caught our own trap, we rebuilt the whole thing slower, stricter, and honest to the bone. This is where the research stands today and where it's heading. We rebuilt the system to be honest, strict splits, an honesty band on every number, and a
Strict patient splits and why they prevent testing bias
[0:18]Hard problem left deliberately open. First, we made the evaluation bulletproof. We split the patients cleanly, training and held out testing kept strictly apart by patient, and we stopped reporting a single lucky number. Instead, we report a confidence interval, a range that says how much the result would wobble if we ran it again.
Replacing single lucky numbers with confidence intervals
[0:35]Our clean OCT-only classifier lands around 0.86 AUROC with that honesty band attached, a number you can actually trust. Then we tested the tempting shortcut everyone reaches for, fusing the image with hand-computed structural measurements. We found they added almost nothing. The network had already learned that information from the image itself. A quietly important result. Knowing what
Testing hand-computed features (Do they actually help?)
[0:57]Doesn't help is real progress, too. And now, the frontier. We're testing whether today's most capable reasoning models, given only the honest structural signal, can deliberate over the hard cases the way a specialist does and explain every call. That experiment is live. We don't have the final answer yet. That's what ongoing research actually looks like,
The future: Tireless first readers buying back clinical time
[1:15]Not a triumphant reveal, but a careful, honest climb. Here's the vision it's all building toward. A patient gets a quick, cheap OCT scan anywhere. A fast model clears the easy cases and flags the uncertain ones. Those go to a reasoning model that lays out the evidence and then to a human specialist with the case already worked up. The machine isn't the doctor. It's the tireless first reader that never gets bored, never skips a scan, and knows the edge of its own knowledge, buying back the one thing glaucoma steals, time. Your turn. Pace this. Which medical decisions should a machine be allowed to make, and which must always stay human? That line is ours to draw. Ask it to place three specific decisions on that line and defend each placement, and to say what evidence would move one across. If it won't commit, push back. From a pale grey shimmer named by the Greeks to a mirror in 1850 to a machine that reasons and admits doubt, the hunt has never moved faster and the silent thief is running out of shadows. This has been the uncertain eye. The research continues and so does the hunt.
More from HAI
2:25The Silent Thief: Why Glaucoma is a Pattern Problem | The Uncertain Eye Ep. 1
2:06The Trap in the Data: Finding and Fixing Data Leakage | The Uncertain Eye Ep. 7
1:37Post-Closing Integration: A Plan, or a Tracker?
2:15Too Good to Be True? The Power of Peer Review | The Uncertain Eye Ep. 6
1:44Should You Respond to an Infringement Letter Right Away?
2:13