← All experience

A measurement product that couldn't measure its own effect

SummonLMCo-founder Oct 2025 — Jan 2026GEO analytics SaaS

Tracking how brands show up in AI answers, with two co-founders and five pilots.

The short version

SummonLM showed brands how AI answers treat them: whether Google AI Overview, ChatGPT, Gemini, Perplexity and Grok mention them, where they rank, how they're described and which sources get cited. Three of us co-founded it. I owned the roadmap, the design and the frontend, and later sales.

Twenty leads became five pilots, and two of them used it regularly. One reached a verbal agreement at ₹60,000 a month. It was never signed and no money changed hands.

That customer came in through the checks that told them what to fix, not the dashboards we pitched. And because we never tracked whether any pilot's rankings moved, we couldn't prove the product caused anything. We wound it down in January 2026.

A narrated walkthrough of the shipped product: the brand setup, the analytics screens, and the technical checks. Because the recording has narration, your browser may ask for one press of play inside the player. Open it in a new tab if the player doesn’t load.

Where it started

My co-founder was a product analyst at Inito and watched their impressions halve in a single core-update week. That was the starting point: not a hunch about where search was heading, but the problem showing up inside a real business.

SummonLM tracked how brands appear in AI answers across Google AI Overview, ChatGPT, Gemini, Perplexity and Grok. We spent about two weeks validating the problem and about two weeks building the first version, then roughly two months building and selling at the same time.

The page I'd show first

If I could show one screen, it wouldn't be a dashboard. It would be the technical roadmap.

The product audited a client's real site and gave every issue an impact score. The roadmap turned those scores into phases, so a team knew what to fix first.

Phase 195–100impact score
Phase 285–94impact score
Phase 370–84impact score
Phase 450–69impact score
Every itemNot startedIn progressCompletedSkipped
Skipping needs a reasonNot relevantBlockedNeed infoAssigned elsewhereOther

Every item had a status: not started, in progress, completed or skipped. Skipping required one of five reasons, and finishing a phase set off a small celebration.

Those two details are the design. A skip reason exists because someone will decide not to act on a recommendation, and we should know why. The celebration exists because the page was built for a person working through a list of fixes, not just reading a report.

The technical roadmap, issues ordered by impact and grouped into phases, each with its fix and status. Screens on this page are from a static rebuild of the shipped product, using a representative data snapshot for one pilot brand.

The checks behind it

The audit ran about 25 checks in four phases, each with a named tool and a way to verify it. I built them into the product end to end, and they ran on a scheduler against real client sites.

The first two phases are conventional SEO and Core Web Vitals. The last two are specific to AI search:

  • Explicit Allow rules for GPTBot, Google-Extended, ClaudeBot, CCBot and PerplexityBot, because security plugins block them by default
  • An llms.txt file at the site root
  • HTML-to-text ratio, treated as token efficiency against a model's context window
  • Entity salience, verified through Google's Cloud Natural Language API
  • The citation test: search for a unique stat from your own site without naming the brand, and see whether Perplexity cites you

The citation test works by taking the brand out of the question, so the model has to find you on substance alone.

Four questions, four screens

The pitch asked four things about a brand: are we mentioned, where do we rank, how are we described, and which sources get cited. The product had exactly four analytics screens to match them: Mention, Ranking, Perception and Citation.

The overview: mention rate, ranking position, perception and citations, then share of voice by brand and by answer engine.

Every screen used the same cuts: model, persona, geography, intent and competition. Each page could have grown whatever charts were easiest to build. Instead they all answered the same questions in the same way.

Mention rate cut by model, persona, intent and geography, the same cuts every analytics screen used.

Most of the operational complexity sat in prompts. Prompt templates were crossed with persona and geography to produce tracked variants, each on its own daily or weekly schedule, written by hand, generated, or imported from CSV.

Tracked prompts, each with its persona, geography, intent, refresh frequency and next scheduled run.

I built the frontend in about two months: roughly 25,500 lines of TypeScript across 200 files on Next.js, with a full custom theme system and multi-user roles from the start.

What the pilots told us

Leads
~20
Pilots
5
Used regularly
2one daily, one weekly
Agreed to pay
1verbal, ₹60,000/month, never signed

Pilots came through a friend's employer, my co-founder's previous company, a growth community, and friends in growth and product roles. Two used it regularly, one daily and one weekly.

The daily user, a bootstrapped e-sourcing business, reached a verbal agreement at ₹60,000 a month, roughly ₹7.2 lakh a year. It was never contracted and never invoiced.

What matters is how they got there. Our pitch, our site copy and our case for a moat were all built around the analytics screens. This customer came in through the other door: the technical checks ran on their real site and surfaced issues they wanted to fix, and that is the customer who agreed to pay.

DescriptiveWhat we pitched

Where do we stand?

MentionRankingPerceptionCitation
PrescriptiveWhere the paying intent came from

What do we fix?

Technical roadmapTopical targetTopical gap analysis

One customer is a signal, not a proof. But it pointed at which half of the product carried the value.

We planned to charge in credits, one per brand-data query, because our costs rose with usage and the margins were thin.

Why we stopped

There was an external reason and an internal one.

Externally, the larger leads were already on competing products or had handed the category to agencies. The market had its winners, and what was left had become a commodity with little value to extract.

Internally, we had no proof the product caused anything. We never tracked a pilot's outcome, so we couldn't tell whether anyone acted on a recommendation or whether a ranking moved because of it.

Our pitch said the moat was a growing dataset of how AI models perceive and rank brands. That claim couldn't be tested. Neither could our bet on long-tail target keywords around frequent, high-traffic queries, because no content was ever published using them.

A measurement product that can't measure its own effect can't prove its data is worth anything.

Both reasons are true. The second is the one that makes the first believable.