Methodology
What is measured, how, and what is not claimed. The whole document.
This is the whole methodology, and the only copy of it: what a run measures, how every number is produced, and what it refuses to claim. For the mechanics of how one run moves through the nine pipeline stages, see How it works; for the code that implements each piece, see ARCHITECTURE.md.
What does Saylent measure?
Saylent asks four AI assistants the questions a real buyer would ask, records every answer, and fetches the pages those answers cited. It measures whether the brand is named and recommended, and whether the brand actually appears on the pages those answers were built from. It also tests, live, whether each assistant's crawler can read the brand's own site. Every number traces back to a stored answer or a fetched page you can open. Nothing is modeled, extrapolated, or bought from a third-party panel.
Where do the questions come from?
questions.ts holds a fixed template library with a quota per type: category
8, comparison 5, problem 4, branded 3, and one each for integration, migration
and trust, 23 in all. Placeholders ({brand}, {category}, {icp},
{competitor}, {competitor2}, {problem}, {year}) come from a brand model
derived from the site; competitors rotate by iteration index so the first rival
does not absorb every comparison, and duplicates are dropped.
The set is then frozen onto the brand, {year} included. A verify run reuses
it verbatim and asks the same engines, so before and after is compared set
against set; changing the inputs that shape questions creates a new set and
visibly rebaselines the trend.
How are the engines asked?
Four answer engines (ChatGPT, Claude, Gemini, Perplexity), each through its official API with the provider's own web search on. Nothing is scraped from a consumer app. Each engine is its own step, so a provider outage costs one engine, not the run, and the report labels the gap.
Sampling is per profile: full takes two draws per engine on scored
questions, and a third only when the first two disagree or one failed -
under a majority of three, two agreeing draws are already the majority, so
the extra draw is asked only when it can change the outcome. smoke takes
one draw per engine and never tiebreaks (a fast, cheap read is the whole
point of the profile). --samples <n> (1-5, also AUDIT_SAMPLES or
saylent.config sampling.samples) overrides either profile's default -
above 2, every draw is taken up front and majority-voted with no separate
tiebreak step. The verdict is the majority vote over mention_type; when
every draw differs, the tied labels are sorted by strength and the middle
taken, rounding toward the weaker label. Because one draw is close to a coin
flip, the recommended count is a {min, max} band across sample sets: a
single number would imply a precision the measurement does not have.
How is an answer judged?
- Presence first, deterministic. A whole-word, case-insensitive alias regex, so "Acme" never matches "Acmeology".
- Cross-family judge. ChatGPT and Gemini answers are judged by the Anthropic family, Claude and Perplexity by the OpenAI family. Models prefer their own output, so the split is a deliberate control.
- Reasoning first. The judge quotes the deciding sentence before emitting
mention_type,prominenceandsentiment, because constrained JSON degrades reasoning when the verdict decodes first. - The rubric. A segment pick is a recommendation even when another brand wins overall; the loser of a head to head, and a split verdict, are "compared"; a branded question has no list, so an affirmative verdict is "recommended" and a skeptical one "dismissed"; prominence is rank among the brands named, not character offset or answer length; a one line endorsement inside a rival's comparison counts; an answer mostly about a different company of the same name is flagged as entity confusion.
- Code wins. Regex absence forces "absent" and "none" after the judge returns. The answer is fenced as untrusted data, and a failure yields a counted neutral verdict after one retry.
- Single-family mode. With one provider key that family serves every
judgment role and the run is stamped
judge_mode: "single-family". The judge model still differs from the answering model, but the control is weaker and the report says so. - The golden set. v2, public and self-contained, in
fixtures/golden-judge.json: 24 entries drawn from one real smoke run against the fictional brand Kestrel Uptime, with the raw answer text of every one inlined in the fixture itself. 18 of the 24 are judgeable (6 engine calls failed in the source run and are kept, markedjudgeable: false, so the file still accounts for the whole bundle); of those 18, the brand is deterministically absent in 12, so only 6 actually exercise the qualitative rubric. Rerunscripts/judge-golden.tson any judge change; it needs onlyANTHROPIC_API_KEY+OPENAI_API_KEY, no database and no run ids.--offlinereplays the verdict the production judge already produced for each answer inside the source bundle, at $0, instead of calling the judge again. Limitation: 18 labels is a floor and a regression guard, not a production agreement rate, and every reference sentiment is neutral, so the set cannot validate positive or negative calibration.
What is checked on the cited pages?
corpus.ts aggregates citations by normalized URL, counting per engine and per
question, then fetches the most-cited pages and checks their text for
whole-word presence of the brand aliases and each competitor. When a fetch
fails, or the page is a video shell, brand_present is null and the row
reads "unverified"; it is never guessed. Pages that block unknown agents fall
back once to a public Internet Archive snapshot, and final_url then points at
the archive copy actually read while the page's identity stays the original
URL. Rows merge on the normalized final URL, so Gemini's redirect links land on
the same row as direct links.
Limitation: word presence is not entailment. A page can name the brand and still contradict the answer, so the report says "brand present on the cited page", never "the page supports the claim".
What do the site access checks know?
| Class | Agents | If blocked |
|---|---|---|
| Training crawler | GPTBot, ClaudeBot | warn: costs model weights, not citations |
| Search-index crawler | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot | fail: removes citation eligibility |
| User-fetch agent | ChatGPT-User, Claude-User, Perplexity-User | fail: kills live page reads |
Google-Extended is information, not a gate: a training opt-out token, not a
crawler. Rules aimed at retired names (anthropic-ai, Claude-Web) are
flagged as instructions no current bot reads.
Robots is a polite request, so six agents are also probed live. A 401, 403 or 429 while robots.txt allows that agent is a failure with a named cause: a CDN or WAF rule overriding robots.txt. Every run carries the caveat that IP-verifying CDNs may treat the probe differently from a real bot, and that your server logs are ground truth.
Page checks: noindex fails, nosnippet warns, a missing Organization JSON-LD
fails, missing Product or FAQPage warns; schema and meta are read from the full
document, so a block at the end of a huge page is still seen. When the crawl
read zero pages, page-level checks are reported as skipped, not failed,
with the likely cause named.
How are fixes produced?
Diagnosis is deterministic: fixes.ts reads the stored checks, answers and
corpus rows and emits fixes in a fixed order. No model decides what is wrong;
it only drafts the artifact afterwards.
Weight families: access blocked 9.5, coverage gap, citation source gap,
negative or wrong claims 7.0, unclear entity 7.0, missing schema 5.5, stale
freshness 4.0. The two middle families follow the brand's own citation mix:
with t the third-party share of engine-cited pages, the coverage hub weighs
7 + 2(1 - t) and outreach 7.5 + 1.5t, so a brand the engines reach only
through third parties gets outreach ranked above publishing; under five cited
pages the static defaults (9.0 and 8.5) apply. These weights, and the 0.45
coverage threshold, are calibrated defaults tuned on internal runs;
saylent.config.{json,js,mjs,ts} thresholds.fixWeights (read by
fixes.ts) and thresholds.coverage (read by coverage.ts) override any
subset of them without editing code; naming a fix family there pins it to
your number and switches its citation-mix calibration off.
A page a competitor owns is never pitched; those roll into one counter-fix. Risk claims are grouped by text across engines, so a fix leads with the claim the most engines agree on, says "N of M engines" only when at least two do, and gains half a point when three or more assert it.
The top fixes get a drafted artifact. Evidence is fenced as untrusted quoted data; bracket placeholders are stripped; two pitches over 60 per cent word-trigram overlap trigger one bounded redraft; a drafter failure yields placeholder text rather than failing the run. A chain-of-verification pass that deletes claims with no evidence line exists behind a flag, off by default.
What is scored, and what is not claimed?
Scores cover the scored types only, category and problem, 12 of the 23: per
engine and overall, answered, recommended, mentioned, and two rates that are
null rather than zero when nothing was answered. Share of voice counts how
often each other brand is named, canonicalized so one company does not split
across rows. The summary (brief.ts) composes all of it into typed blocks with
no model call; cards hide themselves when data is absent, and every number
carries a receipt handle to the stored row. Not claimed: no traffic estimate,
no revenue effect, no invented prompt-volume figures, no market ranking, no
causal claim that a fix caused a movement.
Which models are used?
models.ts is the only place a model name exists, and every value is an
environment override with a default: gpt-5.4, claude-sonnet-4-6,
gemini-3.6-flash and sonar for answers; claude-haiku-4-5 and gpt-5-mini
for the judges; claude-haiku-4-5 for the brand model, claude-sonnet-4-6 for
the drafter. Every run records its resolved role-to-model map, judge mode,
frozen engine set and template version, so two runs months apart can be read
knowing what moved underneath them.
What are the limits?
- One run is a snapshot; signal comes from frozen sets compared as sets.
- We measure the official API surface. Consumer apps can answer differently.
- Google AI Overviews has no API, is not scraped, and is not measured.
- Cited-page checks are word presence, not entailment.
- Access checks report what our fetches saw. Server logs are ground truth.
- The coverage threshold and fix weights are calibrated defaults; a
saylent.configthresholdsoverride changes the number, not the underlying calibration. - The golden set is small (18 judgeable labels) and its brand is deterministically absent in most of them, so it is a regression floor, not a production agreement rate.
- Share of voice counts mentions in the answers we drew, not market share.
- Fix weights rank likely leverage. Movement after a fix is correlation.
What sample data is published?
Layer one, publishable for real brands, is what the assistants answered: who was recommended per question, how often, and which sources they leaned on, attributed to the engine and dated, quoted only where neutral or positive. Layer two, the full report with sentiment, risk claims and generated fixes, is published only for a brand we own or one that gave written consent; a fictional sample brand stands in until then. Layer three is aggregate study data with no opinion about any named company. Every public sample carries the date, the model names, a correction and takedown link, and the statement that it records what the assistants answered on that date rather than our assessment of the company. The line: no verbatim negative claim and no generated fix about a named company we do not own, anywhere public.
Last verified against the code on 2026-09-09.