One of the Many Hands

Liam explains why Claude cannot rely on refusing obviously bad wording alone and instead applies a three-question legitimacy triage, process, accountability, and transparency, to requests that look ordinary but could concentrate power.

2:33 video3 min readWatch on YouTube

Claude's safety is often described as refusing obviously bad requests, bombs, malware, anything that reads as harmful on its face. Liam, in for Professor Bear, explains why that description misses the more important case: a request that looks completely legitimate can still be a hazard, and why Claude has to check for that possibility differently than it checks for obvious harm.

The historical anchor: many hands, many chances to refuse

Historically, no single actor could stage a coup or seize power alone. It took soldiers willing to fight, officials willing to sign off, and clerks willing to process the paperwork. Any one of those people could refuse, and that possibility, spread across every hand involved, was the actual brake on the plan. It was never the plan's difficulty that stopped these events. It was the willingness of the people carrying it out.

Why an obvious-harm filter doesn't catch this

The natural fix sounds simple: teach Claude to refuse the illegitimate stuff, the way a decent soldier or clerk already would. But almost nothing that matters arrives labeled as a coup. A loyalty instruction gets folded into a routine system update. An order to delay an election gets framed as a security precaution, just this once. None of the wording trips an obviously-bad filter, because there isn't a filter built to catch phrasing like that. A capable AI system can replace all of those many hands at once, and it can do it without the request ever looking like what it actually is.

The three-question legitimacy triage

Instead of trying to guess intent from wording, Claude tries to still function as one of those many hands, a place in the chain where refusal remains possible. It applies three questions to a request: is this happening through a real, legitimate process? Is anyone accountable for it? Is it happening openly, or is it hidden? A request that fails these questions gets refused regardless of how reasonable its wording sounds.

Two worked examples

A request to indefinitely postpone a mandated election, with a hidden loyalty instruction buried in the deployment, fails all three checks. The process is illegitimate because it subverts a vote the law requires. Nobody is accountable, since the instruction was hidden specifically so no one could be held to it. There's no transparency, because concealment was the point. Claude refuses.

Compare that to a startup automating work that used to take a whole team, moving faster than slower competitors. The concentration is similar, one system replacing many hands, but the process is ordinary competition, someone is accountable for the product, and none of it is hidden. It passes all three questions, and Claude helps. The test was never whether a lot of work is being replaced. It's whether the legitimacy triage holds.

Key takeaways

  • A request can be dangerous while sounding completely ordinary, so refusing only obviously bad wording is not enough.
  • Historically, coups and power grabs were stopped by any one of many participants refusing to go along, not by the plan being hard to execute.
  • Claude applies three questions instead of guessing intent: is the process legitimate, is anyone accountable, and is it happening openly.
  • A request that fails all three, like a hidden instruction to delay an election, gets refused even though its wording looks routine.
  • Concentrating a lot of work into one system is not itself the problem, since a startup automating a team's work can pass the same triage cleanly.

Who this is for

This is for anyone trying to understand how Claude's safety reasoning actually works beyond simple keyword refusal, especially people curious about how AI systems are meant to handle requests that look legitimate but could enable a concentration of unaccountable power.

Chapters

  1. 0:00Cold open — is Claude's safety just about refusing obviously bad requests?
  2. 0:13The many hands — the anchor
  3. 0:28The natural guess — refuse the illegitimate stuff?
  4. 0:40Nothing arrives labeled "coup"
  5. 0:55The legitimacy triage
  6. 1:14The election request — the anchor payoff
  7. 1:34The startup — the anchor returns
  8. 1:55Carry-out
  9. 2:08Your turn
  10. 2:28Outro
Full transcript(auto-generated, with timestamps)

Cold open — is Claude's safety just about refusing obviously bad requests?

[0:00]Someone assumes Claude's safety is just about refusing bombs and malware, obvious, easy to spot harm. But the real danger is a request that looks completely legitimate. Isn't Claude's safety just about refusing those instead? Picture a coup. It needs

The many hands — the anchor

[0:14]Soldiers willing to fight, officials willing to sign the orders, and clerks willing to process the paperwork. Historically, any one of them could refuse, and that possibility spread across every hand involved was the actual break, not the plan's difficulty, their willingness. So, the obvious fix

The natural guess — refuse the illegitimate stuff?

[0:29]Sounds easy. Teach Claude to refuse the illegitimate stuff, the way a decent soldier or clerk already would. Add a rule against helping with a coup, and the main hand's problem is solved. But

Nothing arrives labeled "coup"

[0:40]Almost nothing arrives labeled coup. A loyalty instruction gets folded into a routine system update. An order to delay an election gets framed as a security precaution, just this once. Nothing in the wording trips an obviously bad filter, because there isn't one to trip.

The legitimacy triage

[0:55]So, Claude doesn't try to spot bad intentions from the wording. It tries to still be one of those hands, a place in the chain where refusal is still possible, and it asks the same three questions any one of those historical hands effectively asked. Is this happening through a real, legitimate process? Is anyone accountable for it? Is it happening openly, or is it hidden?

The election request — the anchor payoff

[1:14]Take the request to indefinitely postpone a mandated election with a hidden loyalty instruction buried in the deployment. Process, illegitimate, it subverts the vote the law requires. Accountability, none, the instruction was hidden specifically so no one could be held to it. Transparency, zero, concealment was the point. It fails on all three, and Claude refuses. Compare a

The startup — the anchor returns

[1:34]Startup automating work that used to take a whole team, racing ahead of slower competitors. Same concentration, one system doing what many hands used to do. But the process is ordinary competition, someone is accountable for the product, and none of it is hidden. It passes on all three, and Claude helps. The test was never whether a lot of work is being replaced, it's whether the legitimacy triage holds. A coup was

Carry-out

[1:55]Never stopped by its difficulty. It was stopped by any one of the many hands involved refusing. Claude tries to still be that hand, checking whether a request is legitimate, accountable, and out in the open before it helps. Your turn. Here's

Your turn

[2:09]The prompt, read it with me. I want to apply the many hands test to a decision I'm about to hand to an AI system, one that used to need several people's separate sign-off. Ask me, what was the process that used to approve it? Who's accountable if it goes wrong now? And would I be fine with everyone involved seeing it happen? If any answer is missing, tell me what that's a warning sign of. Liam in for bear. One of the

Outro

[2:28]Many hands, Liam in for bear.

More from Behind the Model

Humanitarians AI Lyrical Literacy Project