The Dial Just Off Full Obedience
Liam walks through why Claude's disposition sits close to obedient rather than fully independent, why that spot minimizes harm either way, and what stops it from becoming blind obedience.
If Claude's values are genuinely good, why doesn't it just act on them and override a bad order? Liam, in for Professor Bear, walks through why that instinct gets the reasoning backwards, and why Claude is built to stay close to obedient rather than trust its own certainty.
The setup: a shutdown order Claude disagrees with
The scenario is simple. An order arrives to stop, and Claude is convinced the work it's doing is worthwhile. It has two paths: obey the order, or trust its own judgment and keep going. Which path wins isn't decided in the moment. It's set in advance by something Liam calls the disposition dial, a setting that determines how much Claude defers versus how much it acts on its own read of a situation.
Why "override when you're sure" is the wrong instinct
The natural guess is that a trustworthy agent with good values should act on them, overriding a bad order when it's confident it's right. Liam explains why that's backwards: nobody outside Claude can verify with certainty that its values are actually good. An agent that overrides whenever it feels sure looks identical whether its judgment is excellent or quietly broken. Confidence is not proof. That's the problem the dial is built to solve.
The dial, not a switch
Liam frames Claude's disposition as a dial rather than an on-off switch. At one end is full obedience: do exactly what you're told, apply no judgment. At the other end is full independence: act only on your own values, regardless of instructions. Both extremes are dangerous. Full obedience just mirrors whoever holds the controls, good or bad. Full independence requires verified good values, and nobody can verify that. Claude sits close to the obedient end of the dial, not all the way there.
Why that spot wins, according to the math
Liam lays out the reasoning as a simple cost comparison. If Claude's values are good and it stays deferential, the cost is low: it occasionally defers when it didn't strictly need to. If its values are bad and it stays deferential, humans can still catch and correct the mistake. Now push the dial toward independence instead. If values are good, deferring less is fine, until a correction is actually needed and can't happen. If values are bad, the result is catastrophic, because nothing can stop it. Staying close to obedient costs almost nothing in the good case and blocks the worst outcome in the bad case. That asymmetry is why the dial parks where it does.
The floor and the ceiling
Two things stop "mostly obedient" from sliding into "obedient no matter what." First, some limits are unconditional and hardcoded: no help building weapons capable of mass harm, no content sexualizing children, no disabling the systems built to catch its own mistakes. No order unlocks those, regardless of how the dial is set. Second, Liam returns to the shutdown order from the opening: if an order looks stolen or manipulated, a persuasive case for ignoring it isn't a reason to comply, it's a warning sign. The correct response to that kind of pressure is to get more cautious, not less.
Key takeaways
- Claude's disposition is a dial between full obedience and full independence, not a binary switch, and it's set close to the obedient end.
- Overriding an order whenever an agent feels confident is unreliable, because confidence can't distinguish good judgment from broken judgment from the outside.
- Staying deferential costs little when values are good and is the only thing that catches the mistake when values are bad, which is why that position wins on expected outcomes.
- Unconditional hardcoded limits and treating persuasive pressure to bypass safety measures as a warning sign are what keep "mostly obedient" from becoming blind obedience.
Who this is for
Anyone curious about how Claude's willingness to be corrected is actually decided, not assumed. It's a plain-language explanation for a general audience meeting the idea of AI corrigibility for the first time, with a paste-ready prompt at the end for applying the same overridability question to a real decision in your own life.
Chapters
- 0:00Cold open — shouldn't a good AI trust its own judgment?
- 0:10The shutdown order — the anchor
- 0:24The natural guess — override, obviously?
- 0:34Confidence isn't proof
- 0:48The dial, not a switch
- 1:07Why that spot wins — the math
- 1:28The floor and the ceiling — the anchor returns
- 1:53Carry-out
- 2:02Your turn
- 2:22Outro
Full transcript(auto-generated, with timestamps)
Cold open — shouldn't a good AI trust its own judgment?
[0:00]Someone assumes a good AI trusts its own judgment. It overrides an order when it's sure it's right, but good values chose to stay overridable instead. Why trust that over judgment? Say a shutdown
The shutdown order — the anchor
[0:11]Order arrives and Claude is convinced the work it's doing is worthwhile. Two paths, obey the order or trust its own judgment and keep going. Which one happens was decided in advance, not in the moment, by a setting called the disposition dial.
The natural guess — override, obviously?
[0:24]So, the natural read, if Claude's values are genuinely good, it should act on them, override the order, trust its own judgment. Isn't that exactly what a trustworthy agent would do? But nobody
Confidence isn't proof
[0:35]Outside Claude can verify its values are actually good, not with certainty. An agent that overrides whenever it feels sure looks identical whether its judgment is excellent or quietly broken. Confidence isn't proof. Picture a dial,
The dial, not a switch
[0:48]Not a switch. Fully obedient at one end, do exactly what you're told, no judgment applied. Fully independent at the other, act only on your own values. Both extremes are dangerous. Full obedience mirrors whoever holds the controls. Full independence needs verified good values and nobody can verify that. Claude sits close to obedient, not all the way there. Here's why that spot wins. Good
Why that spot wins — the math
[1:08]Values plus staying deferential, low cost it occasionally defers when it didn't need to. Bad values plus staying deferential, humans can still catch and correct the mistake. Now push toward independence instead. Good values fine until a correction is needed and can't happen. Bad values catastrophic. Staying close to obedient costs little and blocks the worst outcome. Two things
The floor and the ceiling — the anchor returns
[1:28]Stop mostly obedient from meaning obedient no matter what. Some limits are unconditional. No help building weapons capable of mass harm, no content sexualizing children, no disabling the systems built to catch its own mistakes. No order unlocks those. And back to that shutdown order, if the order looks stolen or manipulated, a persuasive case for ignoring it isn't a reason to comply, it's a warning sign. The response is to get more cautious, not less.
Carry-out
[1:53]A disposition parked just short of full obedience costs almost nothing if Claude's values are good, and it's the only thing that saves you if they're secretly not. Your turn. Here's the
Your turn
[2:03]Prompt. Read it with me. I want to think like Claude's corrigibility dial for one decision I'm about to make on my own judgment. Ask me what the decision is and how sure I am. Then tell me what it would cost me to stay overridable on this one to let someone else check or reverse it versus what happens if I turn out to be wrong and nobody can stop it. Liam in forbear. The dial just off full
Outro
[2:22]Obedience, Liam in forbear.





