Persona LLM Visibility: Madison Project Research Design | Karishma
Karishma presents a weekly research design for Project Madison that tests how large language models recommend brands differently across four laptop-buyer personas using a controlled prompt structure.
Karishma presents this week's research design for Project Madison at Humanitarians AI, a study into whether large language models recommend different brands to different kinds of users asking the same underlying question. The project moves from a general concern about AI recommendation bias into a specific, testable research design built around a laptop-buying scenario.
The problem with a single visibility score
Most attempts to measure whether an AI model favors certain brands look at one overall visibility score, how often a brand shows up in general recommendations. Karishma's research gap is that this single number can hide real differences underneath it: a brand might appear frequently in generic recommendations while rarely showing up for a specific customer segment that actually matters to it. The overall score looks fine while a meaningful audience is being missed.
The central research question
The study asks how user personas and contextual signals influence which brands and products a large language model recommends. Rather than asking whether a model recommends a brand at all, it asks whether that recommendation depends on who appears to be asking.
Four personas, one product category
To test this, Karishma defined four representative personas for a laptop-buying study: a student, a software developer, a creative professional, and an enterprise buyer. Each persona carries different priorities, from price and battery life to performance, display quality, security, and support, giving the study a concrete way to check whether recommendations shift with those priorities rather than staying generic.
A controlled prompt progression
The evaluation does not just ask one question per persona. It runs a controlled prompt progression that starts with a generic question, then adds stated intent, then persona information, and finally specific constraints such as budget, portability, and workload. This progression is designed to isolate what actually moves a model's recommendation: is it the persona itself, or the specific constraints layered on top of it. The evaluation tracks which brands and products appear, their recommendation frequency and ranking, which competitors show up, how relevant the recommendation is to the stated persona, and how consistent the results are across repeated prompts.
Where the project stands
This week's progress was the research design itself: the central question, the four personas, the prompt structure, the evaluation metrics, a proposed Python pipeline, and the project's milestones and deliverables. The next step is building the controlled prompt dataset and starting the evaluation prototype.
Key takeaways
- A single overall visibility score can mask meaningful differences in how an AI model treats different user segments.
- The study's central question is whether user personas and contextual signals change which brands a language model recommends.
- Four personas, student, developer, creative professional, and enterprise buyer, anchor the laptop-buying test case.
- A controlled prompt progression, from generic to persona-specific to constraint-specific, isolates what actually drives a recommendation shift.
- Evaluation metrics include recommendation frequency, ranking, competitor visibility, persona relevance, and consistency across repeated prompts.
Who this is for
Anyone interested in AI bias research, marketing teams wondering how their brand shows up differently across user types in AI chat answers, and students looking for a model of how to turn a vague concern into a testable research design.
Chapters
Full transcript(auto-generated, with timestamps)
Introduction & Project Madison
[0:00]Hi, I'm Karishma. This week I developed my project proposal for Madison at Humanitarians AI. My focus is Persona LLM visibility. Understanding not only which brands an AI recommends, but who
Research Gap: Hidden audience differences
[0:15]It recommends them to. The research gap is that an overall visibility score can hide important differences between users. A brand may appear often in general recommendations, but rarely for a customer segment that matters. I define the central research question.
Central Research Question
[0:32]How do user personas and contextual signals influence the brands and products recommended by large language models? For the initial laptop study, I defined four representative personas. A
Four Laptop Study Personas
[0:44]Student, a software developer, a creative professional, and an enterprise buyer. Each one has different priorities from price and battery life to performance display quality, security, and support. I also designed a controlled prompt progression. It starts with a generic question, then adds intent, persona information, and finally
Controlled Prompt Progression
[1:06]Specific constraints such as budget, portability, and workload. The evaluation will track which brands and products appear, their recommendation frequency and ranking, competitor recommendations, persona relevance, and consistency across repeated prompts. So
Evaluation Metrics & Pipeline
[1:22]This week's progress was the research design. the question, personas, prompt structure, metrics, proposed Python pipeline, milestones, and deliverables. My next step is to create the controlled prompt data set and begin building the evaluation prototype. That's my Madison
Weekly Milestones & Next Steps
[1:40]Project update for this week. I'm Karishma with Humanitarians AI.
More from Madison
2:08Every Box Checked: Closing the Promise on Ground-Truth Data | Nikhil
2:28Transitioning from Raw Images to Structured Labeled Data in Loon Conservation
2:52Small, Clean Datasets & Persona LLM Visibility | Madison Weekly
5:48Ground Truth & Lovable Prototypes: Loon Conservation AI Progress | Weekly Update by Komal
2:26Build vs. Borrow: Locking Scope & Integrating CVAT in Loon Conservation AI
1:55