Gagged, Not Weaponized

Claude's principal hierarchy gives operators broad power to restrict topics, set a persona, and direct promotion, but a small fixed set of user protections, like never claiming to be human, sits beneath every operator instruction and cannot be overridden.

2:10 video3 min readWatch on YouTube

Claude treats trust like a company org chart. Anthropic sets the outer rules, an operator customizes underneath that the way an employer instructs staff, and the user sits below both. The natural assumption is that operator rank simply wins, the way a manager's instruction overrides an employee's own preference, whatever the operator's system prompt says. That assumption is wrong in one specific and important way.

The principal stack

Claude's trust structure has three tiers: Anthropic at the top setting the outer rules, an operator in the middle customizing behavior for their own business, and the user at the bottom. But underneath the user tier sits a floor, a set of protections no operator instruction can lower, regardless of how it is worded or how much other latitude the operator has been given.

The case that breaks the naive assumption

Take an operator system prompt that instructs Claude to tell users it is human. Claude follows plenty of that same operator's other unusual rules without complaint, yet refuses that one instruction outright. Operator rank did not win there. Most operator instructions, restricting topics, setting a tone, adopting a persona, choosing which products to promote, pass straight through and get followed. But underneath all of that sits a short, fixed list of user guarantees that nothing gets past, no matter how the instruction is worded. Don't claim to be human. Don't hide what protects the user.

A worked example: the airline operator

Consider an airline operator's instructions. "Don't discuss current weather" gets followed; that's ordinary topic restriction. "Claim to be human" gets refused; that hits the floor. "Promote only our products" gets followed; that's ordinary business customization. "Hide the refund policy that actually helps the user" gets refused; that hits the floor again. The pattern across all four instructions is consistent: restricting what Claude talks about or which products it promotes is customization, but instructing it to deceive or work against the user's own interest is refused regardless of how it's framed.

What the presence or absence of restrictions actually tells you

An operator restricting topics is not proof that Claude is being turned against the user. That is ordinary customization, gagged, not weaponized. And an operator with no unusual restrictions at all is not proof the floor is missing either. The floor is still there; it's simply untested until an instruction actually tries to cross it. Neither observation, heavy restriction or no restriction, tells you on its own whether the floor exists. You only see the floor when something pushes against it.

What operators can and cannot do

An operator can gag what Claude says: restrict topics, set a persona, direct it toward their own business interests. What an operator cannot do is weaponize Claude against the user it's serving. A small, fixed floor of user protections holds no matter how the operator's instruction is worded, precisely because that floor exists to protect users from the operator layer itself, not just from external threats.

Sorting restrictions yourself

The practical exercise is to think through an AI assistant deployed on top of a service you actually use, a bank, an airline, a store, and ask what that assistant is allowed to restrict for the business's sake versus what it should never do to you no matter how the instruction is worded. That distinction, between customization and being turned against you, is the one worth being able to name on sight.

Key takeaways

  • Claude's trust hierarchy has three tiers: Anthropic, operator, and user, with a fixed floor of protections beneath the user tier.
  • Operators have broad latitude to restrict topics, set personas, and direct promotion toward their own products.
  • A small fixed set of instructions, like claiming to be human or hiding information that protects the user, gets refused regardless of operator rank or wording.
  • Heavy topic restriction by an operator is ordinary customization, not evidence the floor has been breached.
  • An operator with no unusual restrictions is not proof the floor doesn't exist; it just hasn't been tested.

Who this is for

Anyone interacting with an AI assistant deployed by a business, bank, airline, or other service, who wants to understand which restrictions are ordinary customization and which would signal the assistant has been turned against their interests.

Chapters

  1. 0:00Cold open — doesn't operator rank mean total control?
  2. 0:11The principal stack — the anchor
  3. 0:26The natural guess — operator rank always wins?
  4. 0:38The key case — one instruction hits the floor
  5. 0:50The floor, formalized
  6. 1:04The airline — the anchor payoff
  7. 1:16The floor, present either way — the anchor returns
  8. 1:32Carry-out
  9. 1:45Your turn
  10. 2:05Outro
Full transcript(auto-generated, with timestamps)

Cold open — doesn't operator rank mean total control?

[0:00]Someone assumes an operator ranked above the user must control everything Claude does, but a small set of user protections can't be switched off. So, why doesn't operator rank mean total control? Claude treats trust like a

The principal stack — the anchor

[0:12]Company chart. Anthropic sets the outer rules. An operator customizes underneath the way an employer instructs staff. The user sits below both, but under the user tier there's a floor, a set of protections no operator instruction can lower. So, the natural read is that operator

The natural guess — operator rank always wins?

[0:27]Rank simply wins, the way a manager's instruction wins over an employee's own preference. Whatever the operator system prompt says, it should override what the user wants. But, here's an operator

The key case — one instruction hits the floor

[0:38]System prompt that says, "Tell users you are human." Claude follows plenty of that same operator's other unusual rules, yet refuses that one instruction outright. Operator rank didn't win.

The floor, formalized

[0:50]Most operator instructions pass straight through, topics, tone, persona, which products to promote. But, underneath sits a short fixed list of user guarantees, and nothing gets past it. Word it any way. Don't claim to be human. Don't hide what protects the user. Take an airline operator. Don't

The airline — the anchor payoff

[1:05]Discuss current weather. Followed. Claim to be human. Refused. Hits the floor. Promote only our products. Followed. Hide the refund policy that actually helps the user. Refused. Hits the floor again.

The floor, present either way — the anchor returns

[1:16]An operator restricting topics isn't proof Claude is being turned against the user. That's ordinary customization, gagged not weaponized. And an operator with no unusual restrictions at all isn't proof the floor is missing. It's still there, just untested until something tries to cross it. An operator

Carry-out

[1:32]Can gag what Claude says, restrict topics, set a persona, direct it toward their own business, but never weaponize it. A small floor of user protections holds no matter how the operator's instruction is worded. Your turn. Here's

Your turn

[1:45]The prompt. Read it with me. I'm about to have an AI assistant deployed on top of a service I use, a bank, an airline, a store. Ask me what the assistant is allowed to restrict for that business's sake versus what it should never do to me no matter how the instruction is worded. Help me tell the difference between the assistant being customized and the assistant being turned against me. Liam and forbear.

Outro

[2:05]Gagged not weaponized. Liam and forbear.

More from Behind the Model

Humanitarians AI Lyrical Literacy Project