Mycroft Update: What Can We Actually Prove About AI?

You can now prove what an AI system did. You still can't prove it was right. Divij breaks down the chain-of-trust security camera approach and its honest limits.

2:49 video4 min readWatch on YouTube

How do you verify an AI system when you can't prove absolute truth? That's the question Divij picks up in part two of his accountability mesh series. In part one, he established that structure can be enforced but truth cannot, at least not yet. This video asks the natural follow-up: if truth is off the table for now, what can actually be proven, and is that enough?

The security camera analogy

To make progress, Divij bolted on a new tool, one he compares to a security camera pointed at the AI. A camera doesn't know whether you lied at the checkout. It only knows that you were standing there and what the receipt says. That's the deal with this tool too: it doesn't check whether the AI's answer is true. It records exactly what the AI actually did.

That recording builds a chain of proof with four links. Link one: did the AI actually call the tool it claims to have called? That's camera-proven, verifiable directly. Link two: was the data it pulled real or made up? Also provable. Link three: does a claimed number match the real financial filing it's supposed to come from? That's checkable. Link four: did the AI's stated reasoning actually cause its final answer? That link doesn't exist yet. Nobody has built a way to prove it.

Three ways the camera can be tricked

A camera system, however thorough, can still be gamed in ways worth naming honestly. First, an AI can fetch real data right before answering without that answer actually being derived from the data it just fetched, the same way a rooster crowing before sunrise doesn't mean the rooster caused the sunrise. Second, a longer, busier recording of tool calls can look more impressive without any of those extra steps having been done correctly; doing more doesn't mean any individual step was right. Third, if the wrong person gets access to this detailed footage of the AI's actions, that itself breaks the exact confidentiality the system was built to protect in the first place.

The honest ceiling

Divij is direct about where this leaves things: we can now prove what the system did. We still can't prove it was right. He frames that as the actual boundary of what's currently possible, not a failure of the project.

Closing the gap: three approaches in progress

Three ideas are being worked on right now to push past that ceiling. The first treats a fact the way you'd treat a block in a Jenga tower: pull it out of the evidence supporting an answer and see whether the answer falls over. If removing it doesn't change anything, that fact was never really holding up the conclusion. The second looks inside the AI's actual internal wiring rather than asking it to describe its own reasoning, though right now that only works on small models researchers are allowed to open up, not the powerful models actually in use. The third idea skips the reasoning question entirely and just tracks results over time, the way a weather forecaster is judged: nobody checks whether the forecaster can explain the atmosphere, they check whether it rains on 80% of the days a forecaster said there was an 80% chance.

Why the gap hasn't closed yet

Divij offers a candid explanation for why none of this has fully closed the gap. Powerful models are locked boxes that can't be opened for inspection. Grading a track record properly takes real time, months of real predictions with no shortcuts available. And honestly, he suspects the biggest reason might not be technical at all: right now it's cheaper and faster to build something that just looks accountable, a shiny dashboard, a green checkmark, than to do the slower work of actually proving accountability. This is also the point where Divij is stepping away from this project to move to something new.

Key takeaways

  • The chain-of-trust "security camera" tool records what the AI actually did without judging whether the answer itself is true.
  • Three of four proof links are provable today: tool calls, real vs. fabricated data, and claims matching real filings; the fourth, whether stated reasoning caused the final answer, has no proof method yet.
  • Camera-style tracking can be gamed by fetched-but-unused data, padded step counts, and misuse of the detailed footage itself.
  • Jenga-style evidence removal, inspecting model internals, and weather-forecaster-style track record grading are the three approaches being explored to close the remaining gap.
  • The honest scorecard: structure is enforceable, behavior is observable, and truth remains open.

Who this is for

This update is for anyone tracking AI accountability and interpretability work, especially those following the Humanitarians AI Fellows program's ongoing accountability mesh series. Divij's closing line is meant to be the one thing anyone picking up this project next should read first: structure is enforceable, behavior is observable, truth is still open.

Chapters

  1. 0:00Introduction and the chain of trust concept
  2. 0:30The security camera analogy and its four logical links
  3. 1:15Three ways a camera system can be tricked
  4. 1:45Closing the gap: Jenga tests, inner wiring, and weather forecasting analogies
  5. 2:30The honest scorecard: Structure is enforceable, behavior is observable, truth is open
Full transcript(auto-generated, with timestamps)

Introduction and the chain of trust concept

[0:00]Hi, I'm Divage Powour. This is part two of the accountability mesh series, the chain of trust. Last time I proved structure holds. This time I'm asking what else we can actually prove and where that still falls short. Last time I ended on this structure is enforcable. Truth not yet. Today's question, if we can't prove truth yet, what can we actually prove? And is that enough? So, I bolted on a new tool. Think of it as a security camera on the AI. A camera doesn't know if you lied at the checkout. just knows you were standing there and what the receipt says. That's

The security camera analogy and its four logical links

[0:30]The deal here. This doesn't check if the AI's answer is true. It records exactly what the AI actually did. So now there's a chain of proof. Link one, did it actually call the tool camera proven? Link two, was the data real or madeup provable? Link three, does a claim number match the real financial filing checkable? Link four, did the AI stated reasoning actually cause its final answer? That link doesn't exist yet. Nobody's built it. But a camera can be tricked. When the AI can fetch real data right before answering without the answer actually coming from that data, a rooster crows before the sun rises, the rooster doesn't cause the sunrise. Two, a longer, busier recording looks more impressive. Doing more steps doesn't mean any step was done right. Three, if the wrong person gets to watch this footage, we've broken the exact rule we built the whole system to protect. So,

Three ways a camera system can be tricked

[1:16]Here's the honest ceiling. We can now prove what the system did. We still can't prove it was right. That's not a failure. That's just where the line actually is. So, what would actually close that gap? Two ideas people are working on right now. One, pull a fact out of the evidence, like pulling a block from a Jenga tower and see if the answer falls over. If it doesn't move, that fact was never really holding it up. Two, look inside the AI's actual wiring instead of asking it to describe itself. Right now, that only works on small models we're allowed to crack open, not the powerful ones we actually

Closing the gap: Jenga tests, inner wiring, and weather forecasting analogies

[1:45]Use. There's a third idea, and it skips the reasoning question completely. Stop checking the reasoning at all. just track results over time. Like a weather forecaster, nobody checks if a forecaster can explain the atmosphere. We check if it rains 80% of the time. They say 80%. So why hasn't this happened yet? The powerful models are locked boxes. We can't look inside the ones we actually use. Grading a track record takes real time, months of real predictions, no shortcuts, and honestly, the biggest reason might not be technical at all. Right now, it's cheaper and faster to build something that just looks accountable. A shiny dashboard, a green check mark, than to do the slow work of actually proving it. This is also where I'm leaving this project before moving to something new. So, here's the honest scorecard. What's real? The structure is enforced. Claims are checked against real filings where

The honest scorecard: Structure is enforceable, behavior is observable, truth is open

[2:30]Possible. And now, what actually happened is fully visible. What's still open? Whether the reasoning is genuine, that's not solved. I'm not pretending it is. If I handed this to someone else tomorrow, that's the one sentence I'd want them to read first. Structure is enforcable. Behavior is observable. Truth still open. Signing off.

More from Mycroft Financial AI

Humanitarians AI Lyrical Literacy Project