Saylent

What it costs

Real numbers, by profile and engine count, before you run anything.

Saylent runs on your own provider keys - every dollar goes to OpenAI, Anthropic, Google, or Perplexity directly, never through us. The CLI always prints its estimate and asks you to confirm before it sends anything.

By profile

ProfileQuestionsWhat it doesCost with four engines
smoke (default)6A quick read: is the brand present at all, on the two required engines or on all fourabout $0.60 to $1.20
full23The complete question set across up to 4 engines, for a real gate or a client reportabout $3.70 to $5.50

Both ranges are computed by the same estimator the app shows before a run, for 23 questions across four engines: the full profile at two draws per scored question plus the adaptive tiebreak, the smoke profile at one. They are rounded outward to the nearest $0.10 and widened to cover the runs we actually recorded, so no real number falls outside a printed one. Fewer engines cost less, and the estimate falls with them before you confirm.

Two real recorded runs: smoke audits of a fictional test brand across all four engines (cross-family judge, one draw per scored question, since smoke samples once) cost $0.93 in 392 seconds and $1.11 in 232 seconds, both inside the printed estimate range. That's the honest shape of the number: the estimate is a range because sampling itself is variable, and the run's actual cost is always what gets recorded and shown to you, never the estimate.

What drives the cost

  • Engine count. Each additional answer engine (Gemini, Perplexity) adds its own answer calls on top of the two required ones.
  • Draws per question. Only category and problem questions are scored. smoke takes 1 draw per engine, no tiebreak. full takes 2 draws per engine, with a 3rd only when the first two disagree - most scored questions on full cost 2 draws, not 3.
  • Samples. --samples <n> (1-5, also AUDIT_SAMPLES or saylent.config sampling) directly scales the printed estimate before you spend anything
    • raising it beyond 2 asks every scored question more times per engine and always majority-votes over every draw it takes, with no separate tiebreak step.
  • The judge. One judge call per answer, on the cheaper "judge" tier model, not the answering model.
  • Fixes. A handful of drafted artifacts (the top-ranked fixes only), one call each.
  • The brand model. One call, roughly $0.02, built once per audit (or reused verbatim on a verify).

The daily cap

A local ledger at ~/.saylent/spend.json tracks what you've spent today and refuses to start a run that would push you over DAILY_SPEND_CAP_USD (default $20) - raise it, lower it, or override a single run with --max-usd <n>. The self-hosted app has the same idea at the deployment level: an operator-set daily cap that trips the kill switch when crossed, so every further run is refused until the operator releases it.

Infrastructure, not just providers

Everything above is provider spend, and for the CLI that is the entire bill: it runs on your machine and writes to a folder. The full app adds a database and a place to run it, so a total-cost answer has to include both.

On your laptop or your own server, there is no third bill. Docker Compose brings up Postgres and the job runner next to the app, and what you pay is the machine you already have.

In the cloud, the two pieces are a Supabase project and a host for the Next.js app, usually Vercel. Both have free tiers that a single small team fits inside, with two limits worth knowing before you rely on them.

Supabase

The free plan gives 500 MB of database storage per project, 5 GB of egress, and two free projects per organization. Runs are rows, not files, so the database grows slowly - the cited pages and answers of an audit are text.

The limit that surprises people is not size. Supabase pauses free projects with no activity over a 7-day period. A paused project is restored from the dashboard, but a deployment nobody has opened in a week is asleep when you come back to it, and anything scheduled against it did not run. If the deployment matters, either keep using it or move it off the free plan. Check Supabase's current plan page for the numbers that apply when you read this.

Vercel

The app is a normal Next.js deployment, with one thing to check: the Inngest webhook that drives the audit pipeline declares a 300-second maximum duration, because a full audit's stages are much longer than a typical web request. A plan that caps function duration below that will cut long runs off. Vercel's function-duration limits are per plan, so read the current limits page against the plan you are on.

You do not need Vercel Cron. Saylent's scheduled runs, the weekly verify and the monthly audit, are Inngest functions rather than platform cron jobs. They are off until you set FLAG_SCHEDULED_RUNS=1, and turning them on means the free-tier Inngest account you already need for the pipeline is what schedules them. Locally, npx inngest-cli dev needs no account and no keys at all.

Neither of these is a Saylent bill. There is no hosted Saylent service to pay for.

See the cost before you spend anything

npx saylent audit example.com --dry-run

Prints the full preflight block - keys, engines, judge mode, the cost estimate - and exits. $0 sent to any provider.

npx saylent questions example.com --no-brand-model --print

Prints the question set itself for $0 and no keys at all (template defaults instead of a brand-model call), so you can see exactly what would be asked before spending anything on answers.

How to reduce it

  • Stay on smoke for iterating; save full for a gate or a client-facing report.
  • Use one provider (OPENAI_API_KEY or ANTHROPIC_API_KEY alone) while testing your setup - the report says judge_mode: single-family and the band is honestly wider, but nothing else changes, and it's roughly half the cost of cross-family.
  • saylent verify reuses the frozen question set and answering pattern from the baseline, so it costs about the same as the original smoke run - there's no separate "cheap re-check" mode, because the whole point is asking the identical questions again.
  • Skip engines you don't need: --engines chatgpt,claude runs exactly the two required ones.
  • --skip drafts drops the single most expensive stage per run (the drafter writes copy-ready artifacts for the top fixes) - diagnosis still runs, so the fix plan is unchanged except for the drafted copy. --skip corpus,gates cost $0 in LLM spend either way (they're fetches and deterministic checks, not billed roles), but skip them anyway to save the time. Preflight and --dry-run show the lower estimate before you commit.

See What gets sent to providers for exactly what those dollars buy - no cost is hidden behind a call the docs don't mention.

On this page