What is LLM Visibility? Auditing AI Recommendation Bias | Karishma
Karishma introduces LLM visibility as distinct from search ranking, breaking it into four measurable pillars and explaining why persona-based auditing needs repeated, documented, human-reviewed tests to avoid noise and stereotyping.
When Google gives you ten links, visibility is easy to understand: where did you rank? When an AI assistant gives you one synthesized answer instead, what does visibility even mean? Karishma walks through a framework for measuring how brands show up in AI recommendations, and why it takes more than one number to describe it.
Why LLM visibility is not just a new search ranking
LLM visibility observes which brands enter AI recommendations, how prominently they appear, and how the answer changes depending on the question and the person asking it. Traditional search visibility asks where a brand lands in a results list. An AI assistant instead synthesizes a single answer and may mention only a handful of options. Strong search engine performance can coexist with weak visibility inside generated recommendations. Search rankings and LLM recommendations are related discovery surfaces, but they are not the same measurement, and treating them as interchangeable misses real gaps.
The four pillars of a visibility audit
A visibility audit needs more than one percentage. Four things matter: mention frequency, prominence within the answer, competitor share, and coverage across different prompts. Because LLM responses can vary from run to run, any single answer should be treated as one observation, not a final result. Consider a laptop brand appearing in 40 percent of recommendations. That aggregate average mixes together different questions and different user needs, which can hide where the brand is genuinely relevant, where it disappears entirely, or whether one particular prompt family is driving the whole result. Aggregate visibility is a starting point for investigation, not a diagnosis on its own.
Why the framework has to change with the user
Watch how the criteria that matter shift depending on who is asking. A student may prioritize price and battery life. A developer may need memory and sustained performance. A creative professional may care about display quality and media workflows. An enterprise buyer may emphasize security, management, support, and total cost. The product category stays constant across all of these, but the criteria that determine a good recommendation do not.
Persona-aware visibility: asking to whom, not just how often
This shift leads to persona-aware LLM visibility. Instead of asking only how often a brand gets recommended, the better question is: to whom, for which task, and against which alternatives? The result of this kind of audit is not a single universal score, but a map connecting user needs, prompts, recommendation position, and competitors.
The limits of persona testing
Persona-aware testing has real limitations. Results can shift with wording, model choice, context, and even repeated runs of the identical prompt. Persona descriptions themselves can also introduce stereotypes rather than useful personalization if they are not built carefully. This is exactly why a good audit needs multiple prompts, repeated tests, documented settings, and human review. One polished AI answer proves very little on its own.
The Madison project
This visibility framework is what Karishma is exploring through the Madison project at Humanitarians AI, building the audit methodology into something repeatable rather than anecdotal.
Key takeaways
- LLM visibility measures which brands enter AI-synthesized recommendations, how prominently, and how the answer shifts by question and asker, distinct from traditional search ranking.
- A visibility audit rests on four pillars: mention frequency, prominence, competitor share, and coverage across prompts.
- A single aggregate percentage, like "appears in 40 percent of recommendations," can hide where a brand is actually relevant or irrelevant.
- Persona-aware auditing asks to whom and for which task a brand is recommended, not just how often, since criteria change by user type.
- Because LLM answers vary with wording, model, and repetition, and personas risk introducing stereotypes, audits need multiple prompts, repeated tests, documented settings, and human review.
Who this is for
Marketers, researchers, and brand teams trying to understand how their products show up inside AI assistant recommendations, and anyone building a methodology to measure AI-generated visibility rather than relying on traditional search rank alone.
Chapters
- 0:00The Death of the 10-Blue-Links: What is LLM Visibility?
- 0:25The 4 Pillars of a Visibility Audit: Mentions, Prominence, Competitors, and Prompts
- 0:50Persona-Based Auditing: Customizing Criteria Across Students, Developers, and Enterprises
- 1:15The Limitations of Persona Testing: Noise, Wording, and Stereotype Risks
- 1:40The Madison Project: Designing Repeated, Documented, and Human-Reviewed Tests
Full transcript(auto-generated, with timestamps)
The Death of the 10-Blue-Links: What is LLM Visibility?
[0:00]When Google gives you 10 links, visibility is easy to understand. Where did you rank? But when an AI assistant gives you one answer, what does visibility mean? I'm Karishma, and in this Humanitarians AI Explainer, we're looking at not just whether a brand appears, but who it appears for. LLM visibility is not simply a new search ranking. It observes which brands enter AI recommendations,
The 4 Pillars of a Visibility Audit: Mentions, Prominence, Competitors, and Prompts
[0:25]How prominently they appear, and how the answer changes with the question and the person asking it. Traditional search visibility asks where a brand appears in a results list. An AI assistant instead synthesizes an answer and may mention only a few options. Strong search performance can coexist with weak visibility in generated recommendations. These are related discovery surfaces,
Persona-Based Auditing: Customizing Criteria Across Students, Developers, and Enterprises
[0:50]But not the same measurement. A visibility audit needs more than one percentage. We can look at four things: mention frequency, prominence in the answer, competitor share, and coverage across different prompts. And because LLM responses can vary, one answer should be treated as one observation, not the final result. Suppose a laptop brand appears in 40% of recommendations. That average mixes different questions
The Limitations of Persona Testing: Noise, Wording, and Stereotype Risks
[1:16]And needs. It can hide where the brand is relevant, where it disappears, or whether one prompt family drives the result. Aggregate visibility is a starting point, not the diagnosis. Watch the framework change with the user. A student may prioritize price and battery. A developer may need memory and sustained performance. A creative professional may care about display and
The Madison Project: Designing Repeated, Documented, and Human-Reviewed Tests
[1:40]Media workflows. An enterprise buyer may emphasize security, management, support, and total cost. The category stays the same, the criteria do not. That leads to persona-aware LLM visibility. Instead of asking only, "How often is this brand recommended? Ask to whom, for which task, and against which alternatives. The result is not one universal score, but a map of needs, prompts, recommendation position, and competitors. But persona-aware testing has limitations. Results can change with wording, model, context, and even repeated runs. Persona descriptions can also introduce stereotypes instead of useful personalization. That's why a good audit needs multiple prompts, repeated tests, documented settings, and human review. One polished AI answer proves very little. The verdict: LLM visibility observes mentions, prominence, competitors, and prompt coverage. Persona-aware visibility adds user needs. Because generated answers vary, the measurement must be repeated, documented, and interpreted, not treated as a permanent ranking. This is what I'm exploring through the Madison project at Humanitarians AI. Now, try it yourself. Pick a product category, define a few different user personas, and ask the same underlying question for each one. Track which brands appear, where they appear, and which competitors show up. Repeat the experiment and see what changes. The goal isn't one perfect answer, it's finding the pattern. What is LLM visibility? Karishma for Humanitarians AI.
More from Humanitarians AI Fellows
2:39Bending the Line: How AI and Passports Scale the Circular Luxury Economy
2:49The AI Saturation Crisis on LinkedIn | Yatra Rawat
2:05The Cast That Hid the Bug: Bridging the Gap to Production-Ready | Sai Nikhil
2:24Testing Gordy: Running a 5-Stage AI Tool Review | Yatra
3:13Preparing Data for YOLO: Training and Evaluating Object Detection Models | Swara Joshi
4:00