How Do You Know If You're Supervising Claude Enough?
This video argues that supervision of Claude should scale with how much decision-making you hand over on a given task, and shows a weekly log-classify-audit method for checking whether it actually does.
Ask most people whether they check Claude's work carefully enough, and they answer with a feeling: careful, or not careful enough. This video argues that's the wrong measurement. Supervision should scale with how much decision-making you actually hand over on a given task, not stay flat as a personal habit, and the only way to know if it does is to log it.
The problem with treating supervision as one habit
Most people treat being careful with AI as a single, flat trait: you either double-check things or you don't, and that habit follows you everywhere. But not every task hands Claude the same amount of trust. A quick copy-paste edit and a real decision made mostly by the model are not the same kind of interaction, and a flat habit can't tell them apart. The mismatch between how much you're trusting Claude and how long you're actually checking stays invisible until something forces it into view.
Log, classify, audit
The fix isn't trying harder to be careful. It's building a small tool that does three things: log each interaction as it happens, classify how much decision-making was actually handed to Claude (a quick edit, an assisted task, or a real decision made mostly by the model), and once a week, audit for the gap, specifically the high-autonomy interactions that got almost no verification time.
The anchor: twelve interactions on a grid
The worked example plots twelve real interactions from one week on a grid, how much you relied on Claude across the bottom, how long you actually spent checking up the side. Nine of them land near the diagonal, where usage and verification roughly match. Three sit in a danger corner: heavy reliance on the model, almost no time spent checking. Those three don't just get counted, they get named, and each one gets a suggested next step matched to what it actually was, a quick code check for a script nobody reviewed, a second opinion for a decision that leaned hard on Claude's judgment. The log turns from a report into a to-do list.
What a rising score does and doesn't prove
Run the audit again two weeks later and the danger corner should empty out. That's real progress, the gap closing. But the video is careful about what this does and doesn't prove: a rising score only shows you logged more checking time against those interactions, not that the checking itself caught anything. And an interaction that stays outside the danger corner isn't automatically fine, it just means this pass didn't flag it.
Why this matters
Supervision isn't a switch you either have on or off. It's a gap between how much trust you're extending and how much verification you're actually doing, and that gap is invisible until you measure it directly, task by task, instead of trusting your own sense of how careful you've been.
Key takeaways
- Supervision should track how much decision-making you hand to Claude on each task, not stay constant as a personal habit.
- A simple log-classify-audit loop turns a vague feeling of carefulness into a measurable pattern.
- Plotting reliance against verification time on a grid exposes a specific "danger corner": high trust, low checking.
- A shrinking danger corner over time is real progress, but it proves you checked more, not that the checking caught real problems.
- Anything outside the danger corner still isn't guaranteed safe, it just wasn't flagged by this pass.
Who this is for
Anyone who uses Claude regularly for work and has only a gut feeling about whether they're checking its output enough, especially people handing over increasingly high-stakes tasks without a clear sense of whether their review habits have kept up.
Chapters
Full transcript(auto-generated, with timestamps)
<Untitled Chapter 1>
[0:00]Someone wonders if they're careful enough with Claude. That's the wrong question. Carefulness doesn't scale by task. What matters is whether your checking effort matches how much you're trusting it to decide, Liam explains. The natural assumption is that
One habit, or logged by task?
[0:12]Supervision is one habit. Either you double-check things or you don't, and that habit follows you everywhere. But not every task hands Claude the same amount of trust. Some are quick copy paste edits, others are real decisions made mostly by the model. A flat habit can't track that difference. Only logging how much you actually verified task by task can, and the mismatch stays invisible until you do. The fix isn't
Log, classify, audit
[0:34]Trying harder to be careful, it's building something to log it. A small tool that does three things. Log each interaction as it happens, classify how much decision-making you actually handed to Claude, a quick edit, an assisted task, or a real decision made mostly by the model, and once a week audit for the gap. High autonomy interactions that got almost no verification time. Run it on a real week, 12 interactions plotted on a
The anchor — twelve interactions, one grid
[0:57]Grid. How much you relied on Claude across the bottom, how long you actually spent checking up the side. Most land near the diagonal, right where usage and verification match. Three sit in the danger corner, heavy reliance, almost no time spent checking at all.
From report to to-do list
[1:11]The three danger corner interactions get named, not just counted. Each one gets a suggested next step matched to what it actually was. A quick code check for a script that was never reviewed, a second opinion for a decision that leaned hard on Claude's judgment. The log turns from a report into a to-do list. Run the
Score rising — what does it prove?
[1:28]Audit again in 2 weeks, and the danger corner should empty out. That's real progress, the gap closing. But a rising score only proves you logged more checking time against those interactions. It doesn't prove the checking itself caught anything. And an interaction outside the danger corner isn't automatically fine, it just means this week's tool didn't flag it.
Carry-out
[1:46]Supervision isn't one habit you either have or don't. It's the gap between how much you're trusting Claude and how long you actually checked, and that gap stays invisible until you log it.
Your turn
[1:56]Your turn, here's the prompt. Read it with me. Log 10 interactions you've had with Claude this week, one sentence each. For each, note how much decision-making you handed over and how many minutes you actually spent checking the result. Then look for the pattern. Is your longest checking time going to the highest trust interactions or somewhere else? Liam in for Bear.
Outro
[2:14]How do you know if you're supervising Claude enough? Liam in for Bear.





