Methodology

Where every number on this product comes from.

Every figure AnswerLift shows counts something that actually happened: an answer a paid API returned, or an answer a person read with their own eyes. There is no estimate, no model output, and no computed score anywhere in the product. This page states the method in enough detail that reading the source code afterwards holds no surprise.

An earlier version of this product showed a score out of 100 derived by arithmetic from the brief the customer typed. It measured nothing, and a number in that position is indistinguishable from a measurement to the person reading it. It has been removed entirely rather than relabelled.

Part one

What is queried, and how.

A tracked prompt is a question a team decided is worth asking, with a cadence attached — anything from every 6 hours upward, up to 50 tracked prompts per audit. An hourly scheduler picks up the prompts whose own interval has elapsed and sends each one to the providers chosen for it. Two providers are called, both official and paid:

Perplexity Sonar API

Official paid API at api.perplexity.ai. The prompt is sent as a chat completion; the answer text and the search sources Perplexity reports are stored verbatim.

OpenAI Responses API, web_search tool

Official paid API at api.openai.com/v1/responses with the built-in web-search tool. The answer text and every url_citation annotation are stored verbatim.

An API answer is not the consumer product’s answer. The Perplexity Sonar API and the OpenAI Responses API with web search are separate products from what a person sees in the consumer ChatGPT or Perplexity apps: different models, different retrieval, different sources. Every record collected this way is stored and displayed under the API’s name, never the consumer product’s, and carries the model that answered it.

A run that fails, or that finds no API key configured, is stored as a failed or skipped run and shown as one. Nothing is substituted for an answer that was not obtained. If no provider key is configured on the deployment, collection is simply off and every surface says so.

Deliberately not done

  • The consumer ChatGPT, Perplexity and Google interfaces are never scraped, automated, or driven by a headless browser. Where a consumer surface matters, a person opens it and records what they saw.
  • Google AI Overviews is not collected by any means. There is no route to it we are willing to use, so the product simply does not claim to cover it.
  • Gemini grounding is not used. Its terms forbid this kind of analytics use.
  • No estimated, modelled or computed score exists anywhere in the product. There is no fallback number when collection is off.
  • Nothing is inferred about a prompt that was not asked, or a date on which nothing ran.

Part two

What is recorded, and how it gets verified.

Collected answers and hand-recorded observations land in the same ledger, in the same shape, and pass the same gate. A person can also bulk-import evidence they already gathered as CSV — up to 500 rows at a time, validated row by row, with a preview before anything is written and no partial imports. That is an entry path for human-gathered evidence, not a second collection path.

Recorded on every observationWhat it means
PromptThe exact question that was asked, stored verbatim.
Prompt versionA normalised form of the prompt keys a version counter, so re-asking the same question later appends v2, v3 … instead of overwriting v1.
Engine or providerFor collected records: the API name and the model that answered. For hand-recorded records: which consumer interface a person looked at. The two are never conflated.
Answer textFor collected records, the provider's answer stored verbatim, so a reviewer checks the actual response and not a summary of it.
Cited sourcesEvery URL the provider reported as a source for that answer.
Source label and URLWhat the snapshot is, and where it can be re-checked.
Captured at / observed atWhen the source material was captured, and when the answer was seen.
Collection methodManual observation · imported snapshot · API collection.
Brand classificationNot mentioned · mentioned but not cited · cited. Mention and citation are recorded separately and never conflated.
Classification basisWhether a person made the call, or it is a provisional machine match awaiting review.
ConfidenceHigh, medium or low. Machine-provisional records always start at low.
VolatilityStable, emerging or volatile — derived only from repeated observations of the same prompt on the same provider, never from a single answer.
Evidence noteWhat the record establishes, in plain words, including any qualification.
Reviewer and review timeSet only by the human QA step.

The provisional call, and why it is not a finding.

When an answer is collected, the software makes a first-pass call: if any cited source is on the client’s own domain the record is marked cited; if the brand name appears in the answer text with no such citation, mentioned but not cited; otherwise not mentioned. That is a string match, and string matches are wrong in both directions — a brand named “Apex” matches text that has nothing to do with it, and a brand described without being named is missed.

So the provisional call is stored as exactly that: flagged machine provisional, at low confidence, with QA status pending, and with the full answer text saved so a reviewer checks the response rather than the guess.

The QA gate.

A person reviews the record, enters their own name, and sets it verified or needs review. The reviewer name and time are stored on the record. The gate is enforced where the data is read, not in the page markup: the public share-report query returns only verified observations, so a pending record cannot reach a client screen, and a QA decision never rewrites the snapshot it was made about.

If nothing has been verified yet, the client report says so in those words. It does not fall back to a projection, a placeholder or a score.

Volatility, and why it is usually blank.

Volatility is derived only from the classifications already recorded for the same prompt on the same provider. Fewer than two prior records and it stays “emerging”; three or more that agree makes it “stable”; any disagreement makes it “volatile”. A single answer never produces a volatility claim.

AI-crawler logs.

A customer can import a summary of their own server or CDN logs — up to 5,000 rows, by date, bot, path and status — covering crawlers such as GPTBot, PerplexityBot, ClaudeBot and Google-Extended. AnswerLift does not fetch or obtain those logs itself. A crawler hit records that a bot fetched a page: it is not a citation and not a mention, and it is labelled that way everywhere it appears.

Part three

What is software and what is a person.

The software does

  • Ask tracked prompts on schedule through the two paid provider APIs.
  • Store each answer verbatim with its cited sources, provider, model and time.
  • Make a provisional classification for a reviewer to check.
  • Derive volatility from repeated observations only.
  • Enforce the human-QA gate on what can reach a client report.
  • Render, export and expire the client report.

A person does

  • Choose the prompts and the cadence.
  • Ask questions in the consumer interfaces, where no API exists.
  • Read every collected answer and confirm or correct the classification.
  • Sign off every record before a client can see it.
  • Write the interpretation and the prioritised remediation work.
  • Export and supply the crawler logs from their own servers.

Concretely, a Visibility Sprint covers 1 brand in 1 market, briefed with 35 priority topics. Prompts tracked and reviewed: At least 10, up to the 50 per audit the product enforces — the working set is agreed at kickoff. Delivery: 20 business days from kickoff. Work performed by: AnswerLift — the same reviewer runs the observations and signs off every record before it can reach your client. Provider API usage costs: Included in the fee — collection runs on AnswerLift's own Perplexity and OpenAI API keys, and provider usage is never billed on to you. Delivery and the reviewer are commitments we make; the 50-prompt ceiling and the provider-cost position are properties of the software described above.

What this method cannot tell you.

  • It cannot tell you what the consumer ChatGPT or Perplexity apps say. Records from the APIs describe the APIs; records from the consumer apps describe one person’s session at one moment.
  • It cannot cover Google AI Overviews at all.
  • It cannot give you a share of voice or a rank. Those need a defined denominator across a sampled prompt universe, and this product samples only the prompts you choose.
  • It cannot tell you whether a change you made caused a change in an answer. There is no controlled before-and-after here, and answers move on their own.
  • It cannot promise citations, placement, traffic or revenue, and nothing in the product should be read as doing so.

Still to build. Collection today is the honest core: two providers, a scheduler, verbatim storage, provisional classification, and the QA gate. Not yet built, and not claimed: competitor share-of-voice across a sampled prompt universe, automated alerting on classification changes, per-run cost reporting, and any coverage of Google AI Overviews. This page will be rewritten the day any of that ships.

Still deciding?

See what the Sprint delivers.

Review Sprint scope

Collection code lives in convex/collectionProviders.ts, convex/collection.ts and convex/crons.ts. The QA gate is in convex/audits.ts, in the share-report query.

    Methodology | AnswerLift