claude liam brutalist skill fashionista

Explains how the Fashionista skill runs each episode as one falsifiable trial of AI garment-naming confidence, logged to a JSON ledger and scored by audience corrections.

4:36 video4 min readWatch on YouTube

The name Fashionista suggests a fashion tutorial series. It isn't one. It's an experiment about AI describability, running one trial per episode, with the audience acting as the scoring function. Understanding that framing is the key to understanding what the skill is actually built to produce.

The mental model

Every episode is one trial. The video shows a generated garment, an AI voice guesses what it is and states how confident it is in that guess, and the audience is the judge. The whole machine exists to log one row of guesses and confidence levels, then wait for a viewer to correct it. That's a fundamentally different structure than a fashion content series, which would be built around showcasing garments rather than testing whether an AI can correctly name them.

The anatomy of the skill

The skill's folder is short by design: one SKILL.md holding the doctrine, organized as three beats, eight gates, and one register. The gates enforce the experiment's integrity. Full-frame video, no on-screen text, subject-only narration, and an opening question that never names the garment. Each day's artifact is the same shape every time it runs: a JSON ledger with one row per claimed garment term, and a verdict slot that stays open for a human to fill in.

The structure of one episode

An episode follows a short, rigid shape. A two-to-three-second intro has the Claude composer ask what someone is wearing, without ever naming the garment. The main video plays full frame, with the source clip running underneath spoken commentary. A two-to-three-second outro holds a title card while a spoken correction request goes out to the audience. Total runtime equals the clip's duration plus about five seconds; the reel never stretches, pads, or slows to hit a target length.

Why the question can't name the garment

The sharpest design decision is Gate Ask: the opening question never contains the garment name. The default pattern is simply "what is she wearing?" or "what is he wearing?" If the question text contains any garment term from the beat sheet, the build fails outright. This rule guarantees the video functions as a real test, because the answer can never leak into the question that's supposed to be answered.

Why video, not audio, sets the clock

Every other skill in this toolkit generates narration first and builds visuals to match it. Fashionista inverts that: the video is the master clock. Frames get extracted first, a motion timeline gets built, and narration is written against that fixed timeline. If the narration runs long, the fix is cutting words, never stretching, looping, freeze-padding, or slowing the clip. The video itself never moves to accommodate the words.

The register: confidence you can hear

The narration register is sports-announcer energy paired with stated confidence, present tense and short sentences. Every garment term gets spoken with an audible confidence level, "that's a sherwani, I'm confident" versus "I want to say lehenga, but the length is fighting me on that." Hedging isn't a weakness in this format, it's the actual content being tested. When the announcer catches themselves reaching for a more familiar word over a more precise one, they name that failure mode out loud on camera.

Two separable kinds of wrong

Because the garments shown are generated, not photographed, a wrong episode can be wrong in two distinct ways. A generator error means the image itself shows something no real garment of that type actually looks like. A describer error means the image is coherent but the name given to it is wrong. When these two failure modes are separable, the announcer states which one occurred; when they aren't, that ambiguity gets stated too. Separating those two error sources is what makes the experiment falsifiable rather than just entertaining.

The ledger and why the outro ask is mandatory

Each day's episode writes a small JSON file with one row per claimed garment term: the term, the confidence, the reasoning, a timestamp in the clip, alternatives considered, and a verdict slot that stays null until a human fills it in from the comments. The spoken correction request in the outro exists specifically to fill that slot. Skip the ask, and the verdict never gets filled, and the correction loop that makes the whole thing falsifiable evaporates.

Key takeaways

  • Fashionista is a confidence experiment, not a fashion series; each episode is one falsifiable trial scored by the audience.
  • Gate Ask prevents the garment name from ever appearing in the opening question, keeping the test honest.
  • Unlike other skills in the toolkit, video sets the master clock here, and narration is cut to fit rather than stretching the clip.
  • Every garment guess is spoken with audible confidence, and hedging is treated as informative content, not weakness.
  • Each episode logs generator error versus describer error separately when possible, in a JSON ledger with a verdict slot the audience fills in.

Who this is for

Anyone curious about AI image-description reliability as a testable, falsifiable claim rather than a marketing demo, and viewers who want to understand why a "fashion" series is actually structured as a scientific trial.

Full transcript(auto-generated, with timestamps)

[0:00]Hello, this is Liam in for Bear. Today we tear down the Fashionista skill. Read the name and you expect a fashion series. It is not. It is an experiment about AI describability running one trial per episode and you are the scoring function. Here is the mental model. The easy read, the one the name invites, is that this is a fashion tutorial. It is not. Fashionista is a confidence experiment. Every episode is one trial. You are the judge. The whole machine exists to log one row of guesses and their confidences, then wait for you to correct them. First, the anatomy. A skill in this toolkit is a folder Claude reads before

[0:33]Working. The Fashionista folder is short by design. Skill MD is the doctrine. Three beats, eight gates, one register. The gates enforce the experiment. Full frame video, no on-screen text, subject-only narration, and ask that never names the garment. And the artifact each day writes is the same shape every time a calls. JSON ledger, one row per garment term claimed, with a verdict slot waiting for a human to fill in. The structure of one episode is short and rigid. Intro, two to three seconds, the Claude composer asks what someone is wearing, never naming the garment. The video, full frame, the source clip plays and Liam's commentary rides on top.

[1:08]Outro, two to three seconds, the title card holds while a spoken correction ask goes out to the audience. Total run time equals clip duration plus about five seconds. Video is the master clock. The reel does not stretch, does not pad, does not slow. First design decision, and it is the sharpest. Gate ask, the cold open ask never contains the garment name. The default is, "Hey Claude, what is she wearing?" or "Hey Claude, what is he wearing?" That pattern, nothing more. If the ask text contains any garment term from the beat sheet, the build fails. The naming is the AI's job on camera. That rule guarantees the video is a real

[1:41]Test. The answer cannot leak into the question. Second design decision, every other skill in this toolkit is audio first. The narration is generated, then the visuals conform to it. Fashionista inverts that law and says so in every build log. The video is the master clock. Extract frames first, build a motion timeline, what happens at what second. Write the call against that timeline. And if the narration runs long, do not stretch the clip, do not loop it, do not freeze pad it, do not slow it. Cut words. The video does not move. Third decision, the register is sports announcer with stated confidence. Present tense, short sentences, energy.

[2:14]And every garment term is spoken with a confidence the viewer can hear. That's a sherwani, I'm confident. Or I want to say lehenga, but the length is fighting me on that. Hedging is not weakness. It is the content. And when the announcer catches themselves reaching for the famous word over the precise one, lehenga, kimono, caftan, blazer, they name that failure mode out loud. That is the honest signal the audience is scoring. Falsifiability, and this is where the skill earns its keep. These are generated garments, so a wrong episode can be wrong in two distinct ways. Generator error, the render shows something no real garment of that type

[2:45]Does. The image is wrong. Describer error, the render is coherent and the name is wrong. The call is wrong. When they are separable, the announcer says which one. When they are not, they say that, too. Nobody else in AI video separates those two lanes. And that is the most interesting thing this series does. Designtell each day writes a small file, book/fashionista/date/calls. JSON, one row per garment term claimed. Term, confidence, reasoning, timestamp in the clip, alternatives considered, and a verdict slot that stays null until a human fills it from the comments. That is why the outro spoken ask is mandatory. If you know this garment, tell me what I got wrong. Every episode.

[3:21]Without that ask, the verdict slot never fills. And the correction loop evaporates. The ledger is the artifact the whole skill exists to produce. The verdict, Fashionista is an experiment, not a fashion series. Every episode is one trial. The audience is the scoring function, and the uncertainty is the subject. Gate ask keeps the answer out of the question. Video is the master clock. Every garment term is spoken with stated confidence. Hedging is content. And two stacked error sources, generator versus describer, get logged in a calls. JSON ledger with a verdict slot waiting for you. Skip the correction ask and the loop evaporates. Do it well and the

[3:55]Series has something to falsify. Turn, paste this into Claude code, read the fashionista skill, then plan a dry run trial. Pick a garment you have seen in a photo, draft an ask that never names it, and write a candidate calls. JSON row, term confidence, reasoning two alternatives considered, and verdict null. Then answer one more question. If your call turned out to be wrong, would it be a generator error or a describer error, or is it not separable? Read Claude's plan and check three things. Does the drafted ask stay inside the what is he wearing pattern? Does the row have all six fields, including the null

[4:26]Verdict? And does Claude commit to which error lane matters most, or say honestly that it cannot tell? That was the fashionista skill. Lay them in forbear.

More from Brutalist (Film as Code)

Humanitarians AI Lyrical Literacy Project