Case studies

Three brands. Three very different AI visibility problems.

A SaaS that ChatGPT kept confusing with a namesake. An industrial leader invisible in the answers its buyers read. A D2C brand recommended everywhere — with a warning attached. Here's what the scans found, what was fixed, and what moved.

Client identities withheld — industries, mechanics and outcomes are representative of real engagements; numbers rounded.

B2B SaaS · ~40 people · France

The SaaS ChatGPT kept mistaking for someone else.

The situation

An HR-tech SaaS with solid traction and a name problem: a larger US company with a near-identical name. Sales calls kept opening with corrections — prospects arrived quoting a pricing model the company had retired two years earlier, or features that belonged to the American namesake.

What the scans found

Visibility was actually decent — mentioned on 6 of 10 buyer prompts. The problem was what the engines said: of 24 claims extracted by the Brand Accuracy dimension, 6 were wrong or outdated, and on 3 prompts ChatGPT blended the two companies into one.

"…is a US-based payroll provider founded in 2012, with plans starting at $8 per employee…" — ChatGPT, describing a French company that does neither

The fix — a 6-week sprint

One canonical brand description deployed across the site, social profiles and every listing. A Wikidata entry that made the two entities unambiguous. Corrections pushed to the four sources the engines cited most. A BLUF pricing page the engines could quote instead of inventing.

The outcome — 3 scans later

Fact risks: 6 → 1 Namesake confusion: 3 prompts → 0 Lead-pick rate: 20% → 40%

Brand Accuracy — fact risks per scan

Scan 16 risks
Scan 23 risks
Scan 31 risk

24 claims checked per scan, every verdict linked to the AI response

Illustration — anonymized engagement data

Industrial equipment · family-owned · DACH

The market leader that didn't exist in AI answers.

The situation

Seventy years of engineering reputation, distributors in twelve countries — and a mention rate of 1 out of 10 when buyers asked the engines for equipment recommendations. Two younger competitors, both smaller offline, were being recommended on nearly every prompt.

What the scans found

The Sources dimension explained the mystery: the competitors were carried by industry directories and a trade publication the engines cited constantly — while this company's only citation source was its own website. On Perplexity, competitor comparison pages sat in the top-3 cited sources for 7 of 10 prompts.

"The engines weren't ignoring them out of preference. They had literally nothing to read except the company's own brochure."

The fix — a 4-month managed program

Listings built on the five directories the engines actually cite in this vertical. Two high-citability pages engineered for the exact "best X for Y" prompts buyers use. Digital PR in the two trade publications the engines source from. Every action chosen because a scan showed the gap — not from a generic checklist.

The outcome — over 8 scans

Mention rate: 1/10 → 5/10 (Perplexity) AI share of voice: 4% → 19% In cited sources on 6 prompts

AI share of voice — Perplexity

Before — scan 1

Competitor A
33%
Competitor B
28%
The client
4%

After — scan 8

Competitor A
27%
The client
19%
Competitor B
18%

Bi-weekly scans, 4 months apart

Illustration — anonymized engagement data

D2C e-commerce · home goods · UK

Recommended everywhere — with a warning attached.

The situation

On paper, great AI visibility: mentioned on 8 of 10 prompts. But conversion from AI-referred traffic was oddly weak. The brand suspected pricing; the scans found something else entirely.

What the scans found

The Sentiment dimension flagged a cautious tone on 45% of mentions — nearly every recommendation came with a shipping-delay caveat. The Sources dimension found the culprit: a 2024 forum thread and a review page with twelve stale reviews, cited on 5 of 10 prompts. The logistics problem had been fixed 18 months earlier. The engines didn't know.

"…a solid choice, though some customers have reported long delivery times…" — cited from a thread older than the fix

The fix — sprint, then monitoring

A systematic review-collection push on the platforms the engines actually cite. Listings updated with current delivery facts. A factual BLUF shipping & returns page giving the engines something newer to quote. Fresh editorial mentions to dilute the stale thread.

The outcome — over 5 scans

Cautious tone: 45% → 12% Positive tone: 48% → 79% Stale thread out of cited sources by scan 4

Tone of mentions — all engines

Before — scan 1

Positive 48% Cautious 45%

After — scan 5

Positive 79% Cautious 12%

Every tone verdict traces to the exact wording the engine used

Illustration — anonymized engagement data

Three industries, one method.

None of these started with a hunch. The scans found the gap — with the AI responses as evidence — the actions targeted exactly that gap, and the next scans measured whether it moved. That's the whole method, and it's the same one running in your trial.

Your story starts with a scan

Five engines, evidence included — and your own chapter one, in minutes.

FAQ

Why don't you name the companies?

Because AI visibility gaps are competitively sensitive. A case study that says 'this brand was invisible on the prompts that matter' or 'ChatGPT repeated wrong facts about them for months' is not something clients want their competitors reading with their name attached. So we publish the industry, the market, the findings and the numbers — and withhold the identity. The mechanics are what you can learn from anyway.

Are these results measured with the same platform I would use?

Yes — the same scans. Every number in these stories comes from the platform's dimensions: visibility and AI citation rate, competitors and AI share of voice, brand accuracy claims, cited sources, and sentiment. When you start a trial, your first scan produces exactly this kind of evidence about your own brand.

How quickly can we expect similar movement?

Timeline depends on starting position and competitive density. Quick wins — llms.txt, schema markup, listing corrections, canonical brand descriptions — can show up within a few scans. Brand-query visibility typically improves within 4 to 8 weeks. Category queries — 'best X for Y' — require deeper structural work and show significant movement within 3 to 6 months. The platform re-scans every two weeks, so you see the trend, not just the destination.

Are these results typical?

They are representative of what happens when a specific, diagnosed gap gets a specific fix — not a guarantee. What generalizes is the method: the scans find the gap with evidence, the actions target that gap, and the next scans measure whether it moved. Brands with low competitive density often move faster; saturated categories move slower.

What happens after the initial work is done?

AI answers drift — models update, sources change, competitors publish. That's why Storyzee is a monitoring subscription: your Monitors keep scanning every two weeks, trend lines show any regression early, and each scan ends with a fresh prioritized action plan. Execute it with your team, or hand it to the consulting arm.