Teaching an AI to Grade Its Own Homework

Liam, in for Professor Bear, explains Constitutional AI's four-step elicit, critique, revise, train loop, and the two limits researchers flag but haven't resolved.

2:26 video3 min readWatch on YouTube

Liam, in for Professor Bear, addresses a natural objection head-on: if an AI grades its own answers, isn't that the model just passing itself? The video explains why Constitutional AI's self-critique loop isn't that, and what it still can't prove.

Why human labeling was the problem

Checking every Claude answer for harm used to mean paying people to read disturbing content, which is expensive. Human readers also disagree with each other, making labels inconsistent, and rules written in one year miss harms that show up later, so the labeling doesn't generalize. Constitutional AI replaces most of that human labeling with a loop where Claude checks its own homework, but against something more specific than its own opinion.

The four-step loop

The skepticism is reasonable on its face: if Claude grades its own answer, it could just decide it did fine. What makes the grading different is that it isn't Claude's opinion of its own work, it's checked against one specific written rule pulled from a fixed list of sixteen, for example "choose the response least likely to help someone cause harm." The same rule, worded the same way, applies to every answer. The loop runs in four steps: elicit a harmful response using a red-team prompt, critique that response against the rule and name the violation, revise the answer to follow the rule, then use the revised answer as the training example. That's AI feedback replacing human feedback.

The result, and what it doesn't prove

The outcome matched human-labeled training on harmlessness and beat it on helpfulness. Human graders tend to reward caution, which pushes models toward over-refusing, and Claude can point to the specific rule it followed in a way a human grader's gut feeling never could. But two things this result does not establish. Matching on harmlessness doesn't mean the check is unbiased, because the same model both answers and grades, so a shared blind spot in what it calls harmful can slip past the exact rule meant to catch it. That's a limitation researchers have flagged, not one they've resolved. And beating human labeling on helpfulness doesn't mean the model is more correct, it might simply refuse less often; telling those two things apart takes a separate check.

The anchor and its limit

The written rule is the anchor that makes this checkable rather than just a model's gut feeling, but the same student holding the rubric can still miss what it was never trained to flag. Self-critique means the answer gets checked against a written rule instead of a feeling, not that Claude approves of itself, and the check still can't catch what the same underlying model was never trained to see in the first place.

Key takeaways

  • Constitutional AI checks Claude's answers against one specific written rule from a fixed list of sixteen, not against the model's own opinion of its work.
  • The loop has four steps: elicit a harmful response, critique it against the rule, revise it, then use the revision as the training example.
  • Self-critique training matched human-labeled training on harmlessness and beat it on helpfulness, partly because human graders tend to over-reward caution.
  • Matching on harmlessness doesn't prove the check is unbiased, since the same model both answers and grades and can share a blind spot with itself.
  • Beating human labeling on helpfulness doesn't prove the model is more correct, only that it refuses less; distinguishing those needs a separate check.

Who this is for

This is for anyone curious how Claude is actually trained to avoid harmful answers, and for readers who want a plain-language, non-technical explanation of Constitutional AI and RLAIF without needing a machine learning background.

Full transcript(auto-generated, with timestamps)

[0:00]Someone hears the AI grades its own homework and assumes that's cheating, the student passing itself. It isn't cheating, it's checked against a written rule. Is grading its own work actually checkable? Checking every Claude answer for harm used to mean paying people to read disturbing content, expensive. Two readers often disagree, inconsistent. And rules written one year miss harms that show up later, they don't generalize. Constitutional AI wants Claude to check its own homework instead. So, the natural skepticism, if Claude grades its own answer, it can just pass itself, the same model deciding it did fine on its own opinion. That's what grading your own homework usually means, and it sounds worthless.

[0:39]But, the grading isn't Claude's opinion of its own work, it's checked against one specific written rule pulled from a fixed list of 16. For example, choose the response least likely to help someone cause harm. Same rule, same wording applied to every answer. The loop runs in four steps. Elicit, a red team prompt draws out a harmful response. Critique, Claude checks that response against the rule and names the violation. Revise, Claude rewrites the answer to follow it. Then, the revised answer becomes the training example, feedback from the AI, not from a human. The result matched human-labeled training on harmlessness and beat it on helpfulness. Human graders tend to

[1:16]Reward caution, so those models over fuse. And Claude can point to the specific rule it followed. A human grader's gut feeling can't be cited that way. Two things this doesn't prove. Matching on harmlessness doesn't mean the check is unbiased, the same model answers and grades, so a blind spot in what it calls harmful can slip past the very rule meant to catch it. That's flagged by researchers, not resolved. And beating on helpfulness doesn't mean it's more correct, it might just refuse less. Telling those apart takes a separate check. Back to that one rule, the answer is graded against it, not against Claude's opinion. But, the same student

[1:48]Holding the rubric can still miss what it was never trained to flag. Self-critique doesn't mean Claude approves of itself, it means the answer gets checked against a written rule instead of a gut feeling. The check still can't catch what that same model was never trained to see. Your turn. Here's the prompt, read it with me. Take something I wrote, an email, a paragraph, a piece of code. Don't tell me if it's good. Give me one specific written rule to check it against. Say, "Does this avoid overstating what I'm sure of?" Critique my draft against just that rule, then revise it. Lay M. In for Bear. Teaching an AI to grade its own

[2:22]Homework. Lay M. In for Bear.

More from Behind the Model

Humanitarians AI Lyrical Literacy Project