Mycroft Update by Anjana: ECIS Episode 2
The Earnings Call Intelligence System now cross-checks two independent models across 25 companies to catch exactly where earnings call language is ambiguous.
Six months ago, this system was a proof of concept watching two stock tickers. This update shows what it looks like at a scale that actually matters: a full sector, two independent language models cross-checking each other, and a dashboard built to answer the one question static reports never quite settle, where exactly does the language of an earnings call get ambiguous.
From four readers to a two-model reader
The Earnings Call Intelligence System, ECIS, reads earnings calls using four independent readers and one triangulator that weighs their output. That was version one. In version two, the piece that used to be a single language model reading each transcript chunk is now two: Llama and Mistral, each processing every chunk independently through the same pipeline, the same prompts, and the same self-consistency checks. Critically, they do not share answers with each other before the triangulator sees them. When the two models agree, confidence in that reading goes up. When they disagree, the system knows precisely where the underlying language is genuinely ambiguous, rather than guessing.
Scaling from two tickers to a full sector
The pipeline that once covered two companies now runs across 20 to 25 companies, processing hundreds of transcripts and thousands of chunks through the same four readers, the same routing logic, and the same scorecard used in version one. That jump from a handful of tickers to a full sector is the real test of whether the design holds, since a method that works cleanly on two companies can still fall apart once volume and variety increase.
Drawing the line between maintained and nothing
At this scale, the hardest classification call is not whether guidance was raised or lowered, it is the boundary between "maintained" and "nothing." When a CEO says "we are comfortable with our current outlook," is that a reaffirmation of guidance or just a sentence with no formal signal behind it? The system now draws that boundary precisely: maintained means the company actively reaffirmed its guidance, nothing means there was no guidance given at all. Every signal behind that call, every reader's vote, every calibration curve, every model's performance, and every decision the feedback loops make, lands in one place rather than being buried across separate outputs.
A live dashboard instead of a static report
Rather than producing a report generated after the fact, the output is a live system that can be queried, drilled into, and interrogated directly. Two models, 25 companies, one scorecard that stays current rather than going stale the moment it is published. That live, queryable structure is the point: a static report answers the questions someone thought to ask before generating it, while a live dashboard can answer the questions that come up afterward.
Key takeaways
- ECIS version 2 replaces its single-model reader with two independent models, Llama and Mistral, cross-checked against each other on every transcript chunk.
- Model disagreement is treated as a signal: it marks exactly where the earnings call language is genuinely ambiguous, rather than being averaged away.
- The pipeline scaled from two tickers to a full sector of 20 to 25 companies without changing its core reader and triangulator structure.
- The system draws a precise line between "maintained" guidance, an active reaffirmation, and "nothing," where no guidance was given at all.
- Output lands in a live, queryable dashboard rather than a static report generated after the fact.
Try it yourself
Think of one process at work where you currently trust a single judgment call, one model, one reviewer, one gut check, that you have never actually questioned. Ask Claude to help you define what a genuine second, independent check would look like, what counts as agreement versus disagreement between the two, and whether disagreement would be rare enough that adding a second check is not worth the cost. This kind of systems thinking is part of the work happening on Mycroft Financial AI within the Humanitarians AI Fellows program.
Chapters
- 0:00Intro: ECIS Version 2 at scale
- 0:25The Two-Model Reader: Cross-checking Llama against Mistral
- 0:45Identifying Ambiguity: What happens when models disagree
- 1:05Scaling Up: Moving from two tickers to 25 companies
- 1:30The Guidance Boundary: Distinguishing "Maintained" from "Nothing"
- 2:00The Live Dashboard: Interrogating the data in real-time
- 2:25Your Turn: Auditing your own single-reviewer processes with Claude
Full transcript(auto-generated, with timestamps)
Intro: ECIS Version 2 at scale
[0:00]Last episode, Essis was four readers and one triangulator on two tickers. I'm Anjana. Here's the same system six months later, running at a scale that actually matters. Essis reads earnings calls with four independent readers and one triangulator. That was version one. Here is what it looks like now. The LLM reader used to be one model. Now it is two. Llama and Mistral process every chunk
The Two-Model Reader: Cross-checking Llama against Mistral
[0:26]Independently through the same pipeline, the same prompts, the same self-consistency checks. They do not share answers. The triangulator weighs both. When they agree, confidence goes up. When they disagree, the system knows exactly where the language is ambiguous. The pipeline now runs across 20 to 25
Identifying Ambiguity: What happens when models disagree
[0:45]Companies. Hundreds of transcripts, thousands of chunks, all flowing through the same four readers, the same routing, the same scorecard. What started on two tickers is now watching a full sector. At this scale, the hardest call is not raised or lowered. It is the line between maintained and nothing. A CEO
Scaling Up: Moving from two tickers to 25 companies
[1:05]Says, "We are comfortable with our current outlook." Is that guidance maintained or is that just a sentence? The system now draws that boundary precisely. Maintained means the company actively reaffirmed. Nothing means there was no guidance at all. Everything lands in one place. Every signal, every reader's vote, every calibration curve, every model's performance, and every decision the feedback loops make. Not a
The Guidance Boundary: Distinguishing "Maintained" from "Nothing"
[1:32]Report generated after the fact. A live system you can query, drill into, and interrogate. Two models, 25 companies, one scorecard that never looks away. Essis, episode two. Let's recap with Claude. The LLM reader is now two independent models cross-checked against each other. The pipeline covers 25 companies, not two. The system draws a precise line between maintained guidance and no guidance at all. And everything
The Live Dashboard: Interrogating the data in real-time
[2:01]Lands live in one queryable dashboard. Your turn. I have one process at work where I trust a single judgment call, one model, one reviewer, one gut check, and I've never questioned whether that's enough. Can you help me? One, figure out what a genuine second independent check would look like for that specific call. Two, define what counts as agreement versus disagreement between the two. And three,
Your Turn: Auditing your own single-reviewer processes with Claude
[2:25]Tell me honestly whether the disagreement case is rare enough that adding a second check isn't worth the cost. Paste that into Claude and see whether your own single model process has a Llama and Mistral problem hiding in it. Two models, 25 companies, one scorecard that never looks away. That's Inzana.
More from Mycroft Financial AI
5:36Project Mycroft: Building a Structural Enforcement Layer to Stop Silent AI Failures
2:45Anjana's update on FinBERT: Why General AI Fails to Understand the Language of Money
10:00Building AI News Monitoring Agent with n8n & FastAPI | Mycroft | Nerd Stuff with Humanitarians AI
3:01Week 2 of Mycroft's Private AI Valuation Agent: Scaling to 80 Million Rows of SEC Data
12:52What Do Mycroft Fellows Do? | Building AI-Powered Investment Intelligence | Professor Bear Explains
2:01