A Council of Specialists: Multi-Agent AI Debate | The Uncertain Eye Ep. 5

Varun walks through building a three-agent vision language model council that debates uncertain eye scans over three rounds before reaching a verdict.

2:13 video3 min readWatch on YouTube

Varun introduces a different way to handle the eye scans a standard classifier could not confidently call. Instead of building a bigger, more powerful classifier, his team built a conversation: three AI specialists that look at the same scan, argue about what they see, and reach a joint verdict.

Why a single classifier was not enough

A plain image classifier turns a scan into a list of numbers. It can be accurate, but it cannot explain its reasoning, and a number without a reason is hard to trust or audit in a medical setting. For the cases the classifier flagged as uncertain, Varun's team wanted something that could reason in words the way a clinician does, not just output a probability.

Vision language models as the building block

The shift starts with a different kind of model: a vision language model, trained on images and text together. Rather than producing a silent score, it can look at a scan, read the relevant measurements, and describe its reasoning in language. That difference matters most exactly where a plain classifier struggles, because a written rationale can be checked and challenged in a way a bare number cannot.

Building the three-agent council

Instead of relying on one vision language model, the team built three, each with a distinct role. A structural specialist reads the OCT nerve map. A functional specialist reads the visual field test. A director weighs both of their reads and delivers the final call. This mirrors how difficult medical decisions get made in practice: a tumor board, a panel, specialists challenging each other until a clearer picture emerges.

How the three-round critique works

The council does not just vote once. Each specialist gives its initial read, then sees what the others concluded, reconsiders its own position, and revises before the director makes the final decision. This structure forces disagreement into the open. A single model has one blind spot and no way to notice it; three models with different areas of focus have to make their disagreement explicit before they can resolve it, and that is where errors a lone model would sail past get caught.

Early results, and why they got nervous

The team ran the council specifically on the flagged uncertain cases, the exact eyes the original classifier could not confidently call. The early results were striking, genuinely and excitingly so. Varun frames that as the moment to get nervous rather than celebrate, since in research a result that looks too good is often the signal to start hunting for what was missed. The team was proud enough of the approach to submit it for peer review, and Varun notes that submission is where an expert reviewer asked a question the team had not considered, setting up the next episode.

Key takeaways

  • A vision language model can explain its reasoning in words, unlike a plain classifier that only outputs a score.
  • The council splits diagnostic work by specialty: a structural reader, a functional reader, and a director who weighs both.
  • Three rounds of critique force each specialist to reconsider its read after seeing the others' opinions.
  • Multi-agent critique catches errors a single model has no way to see on its own.
  • A striking early result is a prompt to look harder for what might be wrong, not a reason to stop checking.

Who this is for

This is for people building or evaluating AI systems for high-stakes decisions, especially in medical imaging, who want a concrete example of multi-agent design that goes beyond a single model's blind spots.

Chapters

  1. 0:00Debating the hard cases with multi-agent AI
  2. 0:18Vision Language Models (VLMs) vs. standard classifiers
  3. 0:35Building the 3-agent council: Structure, Function, and Director
  4. 0:55How a 3-round critique resolves conflicting medical signals
  5. 1:15Early striking results and why we got nervous
Full transcript(auto-generated, with timestamps)

Debating the hard cases with multi-agent AI

[0:00]Hi, I'm Varun, and this video is about a council of specialists, AI agents debating the hard cases. For the hard cases the eyes are classifier couldn't call, we tried something genuinely ambitious, not a bigger classifier, a conversation, a panel of eye specialists that would look at the same eye, argue, and reach a verdict together. For the

Vision Language Models (VLMs) vs. standard classifiers

[0:18]Hard cases we built a better conversation, three specialists, three rounds, one verdict, and results striking enough to make us nervous. To build it we used a new kind of model, a vision language model. A plain classifier turns an image into a silent list of numbers, powerful but it can't tell you why. A vision language model is

Building the 3-agent council: Structure, Function, and Director

[0:36]Trained on images and text together, so it can look at a scan, read the measurements, and reason in words the way a clinician thinks aloud. That difference matters in medicine, where a number without a reason is very hard to act on and impossible to audit. Then we did something more interesting than using one. We built three, each with a different expertise, a structural specialist that read the OCT nerve map,

How a 3-round critique resolves conflicting medical signals

[0:55]A functional specialist that read the visual field test, and a director who weighed them both. And we had them deliberate over three rounds, each specialist gives its read, then sees the others opinions, reconsiders, and revises before the director delivers the final call. Why a council instead of a single model? Because that's how hard medical decisions actually get made, a tumor

Early striking results and why we got nervous

[1:15]Board, a panel, specialists challenging each other until the truth surfaces. The research on multi-agent AI backs the intuition, models that critique each other can catch errors a lone model sails right past. A lone model has one blind spot and no way to see it. Three models with different remits have to make their disagreement explicit before they can resolve it. We ran it on the flagged uncertain cases, the exact eyes the classifier struggled with, and the early results were striking. Genuinely, excitingly striking, which in research is precisely the moment you should get nervous. Your turn, paste this. When a result looks too good, should you celebrate or start hunting for what you missed? Ask for the three most common reasons a medical AI result collapses under scrutiny and which one is hardest to detect from inside your own project. If it doesn't name that blind spot, push back. A council of specialists, we were proud enough to submit it for peer review, and that's when an expert asked the one question we hadn't. Next time.

More from HAI

Humanitarians AI Lyrical Literacy Project