What it costs
Real numbers, by profile and engine count, before you run anything.
Saylent runs on your own provider keys - every dollar goes to OpenAI, Anthropic, Google, or Perplexity directly, never through us. The CLI always prints its estimate and asks you to confirm before it sends anything.
By profile
| Profile | Questions | What it does | Cost with four engines |
|---|---|---|---|
smoke (default) | 6 | A quick read: is the brand present at all, on the two required engines or on all four | about $0.60 to $1.20 |
full | 23 | The complete question set across up to 4 engines, for a real gate or a client report | about $3.70 to $5.50 |
Both ranges are computed by the same estimator the app shows before a run, for 23 questions across four engines: the full profile at two draws per scored question plus the adaptive tiebreak, the smoke profile at one. They are rounded outward to the nearest $0.10 and widened to cover the runs we actually recorded, so no real number falls outside a printed one. Fewer engines cost less, and the estimate falls with them before you confirm.
Two real recorded runs: smoke audits of a fictional test brand across all
four engines (cross-family judge, one draw per scored question, since
smoke samples once) cost $0.93 in 392 seconds and $1.11 in 232
seconds, both inside the printed estimate range. That's the honest shape of
the number: the estimate is a range because sampling itself is variable, and
the run's actual cost is always what gets recorded and shown to you, never
the estimate.
What drives the cost
- Engine count. Each additional answer engine (Gemini, Perplexity) adds its own answer calls on top of the two required ones.
- Draws per question. Only
categoryandproblemquestions are scored.smoketakes 1 draw per engine, no tiebreak.fulltakes 2 draws per engine, with a 3rd only when the first two disagree - most scored questions onfullcost 2 draws, not 3. - Samples.
--samples <n>(1-5, alsoAUDIT_SAMPLESorsaylent.configsampling) directly scales the printed estimate before you spend anything- raising it beyond 2 asks every scored question more times per engine and always majority-votes over every draw it takes, with no separate tiebreak step.
- The judge. One judge call per answer, on the cheaper "judge" tier model, not the answering model.
- Fixes. A handful of drafted artifacts (the top-ranked fixes only), one call each.
- The brand model. One call, roughly $0.02, built once per audit (or
reused verbatim on a
verify).
The daily cap
A local ledger at ~/.saylent/spend.json tracks what you've spent today and
refuses to start a run that would push you over DAILY_SPEND_CAP_USD
(default $20) - raise it, lower it, or override a single run with
--max-usd <n>. The self-hosted app has the same idea at the deployment
level: an operator-set daily cap that trips the kill switch when crossed, so
every further run is refused until the operator releases it.
Infrastructure, not just providers
Everything above is provider spend, and for the CLI that is the entire bill: it runs on your machine and writes to a folder. The full app adds a database and a place to run it, so a total-cost answer has to include both.
On your laptop or your own server, there is no third bill. Docker Compose brings up Postgres and the job runner next to the app, and what you pay is the machine you already have.
In the cloud, the two pieces are a Supabase project and a host for the Next.js app, usually Vercel. Both have free tiers that a single small team fits inside, with two limits worth knowing before you rely on them.
Supabase
The free plan gives 500 MB of database storage per project, 5 GB of egress, and two free projects per organization. Runs are rows, not files, so the database grows slowly - the cited pages and answers of an audit are text.
The limit that surprises people is not size. Supabase pauses free projects with no activity over a 7-day period. A paused project is restored from the dashboard, but a deployment nobody has opened in a week is asleep when you come back to it, and anything scheduled against it did not run. If the deployment matters, either keep using it or move it off the free plan. Check Supabase's current plan page for the numbers that apply when you read this.
Vercel
The app is a normal Next.js deployment, with one thing to check: the Inngest webhook that drives the audit pipeline declares a 300-second maximum duration, because a full audit's stages are much longer than a typical web request. A plan that caps function duration below that will cut long runs off. Vercel's function-duration limits are per plan, so read the current limits page against the plan you are on.
You do not need Vercel Cron. Saylent's scheduled runs, the weekly verify and
the monthly audit, are Inngest functions rather than platform cron jobs. They
are off until you set FLAG_SCHEDULED_RUNS=1, and turning them on means the
free-tier Inngest account you already need for the pipeline is what schedules
them. Locally, npx inngest-cli dev needs no account and no keys at all.
Neither of these is a Saylent bill. There is no hosted Saylent service to pay for.
See the cost before you spend anything
npx saylent audit example.com --dry-runPrints the full preflight block - keys, engines, judge mode, the cost estimate - and exits. $0 sent to any provider.
npx saylent questions example.com --no-brand-model --printPrints the question set itself for $0 and no keys at all (template defaults instead of a brand-model call), so you can see exactly what would be asked before spending anything on answers.
How to reduce it
- Stay on
smokefor iterating; savefullfor a gate or a client-facing report. - Use one provider (
OPENAI_API_KEYorANTHROPIC_API_KEYalone) while testing your setup - the report saysjudge_mode: single-familyand the band is honestly wider, but nothing else changes, and it's roughly half the cost of cross-family. saylent verifyreuses the frozen question set and answering pattern from the baseline, so it costs about the same as the originalsmokerun - there's no separate "cheap re-check" mode, because the whole point is asking the identical questions again.- Skip engines you don't need:
--engines chatgpt,clauderuns exactly the two required ones. --skip draftsdrops the single most expensive stage per run (the drafter writes copy-ready artifacts for the top fixes) - diagnosis still runs, so the fix plan is unchanged except for the drafted copy.--skip corpus,gatescost $0 in LLM spend either way (they're fetches and deterministic checks, not billed roles), but skip them anyway to save the time. Preflight and--dry-runshow the lower estimate before you commit.
See What gets sent to providers for exactly what those dollars buy - no cost is hidden behind a call the docs don't mention.