Python API¶
Generated from the docstrings. For a guided tour see the API overview.
Comparing variants¶
flowprompt.testing.compare.compare ¶
compare(prompts: dict[str, Any], inputs: list[dict[str, Any]], model: str = 'gpt-4o', *, expected: list[Any] | None = None, eval_metric: str | Callable[..., Any] = 'contains', success_fn: Callable[..., Any] | None = None, metric_fn: Callable[[Any], float] | None = None, confidence_level: float = 0.95, runs_per_input: int = 1, temperature: float = 0.0, test_type: str = 'auto', dry_run: bool = False, control: str | None = None, comparisons: str = 'control') -> ComparisonResult
Compare variants on the same inputs with paired significance tests.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prompts
|
dict[str, Any]
|
Dict mapping variant names to variants: Prompt subclasses,
|
required |
inputs
|
list[dict[str, Any]]
|
Input dicts; every variant runs on every input. |
required |
model
|
str
|
Default model for Prompt variants. |
'gpt-4o'
|
expected
|
list[Any] | None
|
Expected outputs (same length as |
None
|
eval_metric
|
str | Callable[..., Any]
|
Scorer for |
'contains'
|
success_fn
|
Callable[..., Any] | None
|
Custom pass/fail function |
None
|
metric_fn
|
Callable[[Any], float] | None
|
Numeric score |
None
|
confidence_level
|
float
|
Confidence level (default 0.95, i.e. alpha 0.05). |
0.95
|
runs_per_input
|
int
|
Runs per input and variant. Repeats are averaged per input; they measure run-to-run variation but do not increase the sample size, which is the number of inputs. |
1
|
temperature
|
float
|
Default temperature for Prompt variants. |
0.0
|
test_type
|
str
|
|
'auto'
|
dry_run
|
bool
|
Only estimate the cost; make no calls. |
False
|
control
|
str | None
|
Name of the baseline variant (default: the first one). |
None
|
comparisons
|
str
|
|
'control'
|
Returns:
| Name | Type | Description |
|---|---|---|
A |
ComparisonResult
|
class: |
ComparisonResult
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
On fewer than 2 variants, no inputs, mismatched
|
Example
result = compare( ... {"short": Short, "detailed": Detailed}, ... inputs=[{"text": t} for t in texts], ... expected=labels, ... eval_metric="exact", ... model="gpt-4o-mini", ... ) print(result.verdict)
flowprompt.testing.compare.acompare
async
¶
acompare(prompts: dict[str, Any], inputs: list[dict[str, Any]], model: str = 'gpt-4o', *, expected: list[Any] | None = None, eval_metric: str | Callable[..., Any] = 'contains', success_fn: Callable[..., Any] | None = None, metric_fn: Callable[[Any], float] | None = None, confidence_level: float = 0.95, runs_per_input: int = 1, temperature: float = 0.0, test_type: str = 'auto', dry_run: bool = False, control: str | None = None, comparisons: str = 'control') -> ComparisonResult
Async version of :func:compare; variants run concurrently.
Same arguments and result. Prompt variants use Prompt.arun();
async callables are awaited; sync callables run in a worker thread.
flowprompt.testing.compare.ComparisonResult
dataclass
¶
Result of comparing variants.
Attributes:
| Name | Type | Description |
|---|---|---|
winner |
str | None
|
Name of the variant that is significantly better than the others (after multiple-comparison adjustment), or None. |
variants |
dict[str, VariantResult]
|
Per-variant results keyed by name. |
statistical_result |
StatisticalResult | None
|
The comparison that decides the outcome (the winner's comparison, or the closest call when there is none). |
confidence_level |
float
|
Confidence level used. |
total_runs |
int
|
Total number of runs across all variants. |
estimated_cost |
dict[str, Any] | None
|
Pre-run cost estimate (also filled in dry-run mode). |
has_expected |
bool
|
True when |
comparisons |
list[StatisticalResult]
|
All pairwise comparisons that were tested. |
control |
str | None
|
Name of the control (baseline) variant. |
outcome |
str
|
What was measured: |
has_ground_truth |
bool
|
False when outputs could not be graded (no expected outputs, success_fn or metric_fn). |
n_inputs |
int
|
Number of inputs. |
runs_per_input |
int
|
Runs per input and variant. |
correction |
str | None
|
Multiple-comparison correction applied ( |
comparison_mode |
str
|
|
model |
str | None
|
Default model used for Prompt variants. |
notes |
list[str]
|
Caveats worth showing next to the result. |
total_cost_usd
property
¶
Metered cost of all runs (0.0 when no price was known).
is_proportion
property
¶
True when scores are pass/fail rates (shown in points).
enough_data
property
¶
False when no difference could reach significance at this size.
With n inputs the smallest attainable two-sided p-value of the
paired tests is 2 / 2**n (every input favours the same variant);
with m comparisons Holm multiplies it by up to m.
sample_size_plan ¶
sample_size_plan(min_detectable_difference: float = 0.1, *, power: float = 0.8) -> SampleSizePlan | None
Inputs needed to detect a given difference, using this run as a pilot.
Uses the discordance (pass/fail) or the SD of per-input differences (numeric scores) observed in the deciding comparison, and the observed cost per run when it is known.
save_report ¶
Write the report; the format follows the extension (.html/.md/.txt/.json).
flowprompt.testing.compare.VariantResult
dataclass
¶
Results for a single variant.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Variant name. |
samples |
int
|
Number of runs completed (inputs x runs_per_input). |
successes |
int
|
Runs that passed (pass/fail outcomes) or completed without error (numeric outcomes). |
success_rate |
float
|
|
mean_latency_ms |
float
|
Average latency per run in milliseconds. |
total_cost_usd |
float
|
Total metered cost in USD (0.0 when unknown; see
|
outputs |
list[Any]
|
Outputs of the runs that did not raise. |
errors |
list[str]
|
Error messages of runs that raised (counted as failures). |
label |
str
|
Human-readable description (prompt class, model, function). |
mean_score |
float | None
|
Input-level mean score (accuracy for pass/fail). |
ci_low |
float | None
|
Lower confidence bound for |
ci_high |
float | None
|
Upper confidence bound for |
p95_latency_ms |
float
|
95th percentile latency per run. |
cost_known |
bool
|
True when every LLM call had a known price. |
cost_per_correct |
float | None
|
|
total_tokens |
int
|
Tokens used across all LLM calls. |
llm_calls |
int
|
Number of LLM calls (including retries and cache hits). |
scores |
list[list[float]]
|
Per-input lists of run scores (errors scored 0). |
records |
list[RunRecord]
|
Every individual run. |
flowprompt.testing.compare.estimate_compare_cost ¶
estimate_compare_cost(prompts: dict[str, Any], inputs: list[dict[str, Any]], model: str = 'gpt-4o', *, runs_per_input: int = 1, estimated_output_tokens: int = 100) -> dict[str, Any]
Estimate the cost of running compare() without making API calls.
Prompt variants are rendered for each input and their input tokens are counted with litellm; prices come from litellm's model cost map. Custom callables cannot be inspected, so they count calls but no tokens.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict with keys: model, total_calls, estimated_input_tokens, |
dict[str, Any]
|
estimated_output_tokens, estimated_cost_usd, per_variant. Cost fields |
dict[str, Any]
|
are None when a price is unknown. |
Paired statistics¶
flowprompt.testing.paired.paired_test ¶
paired_test(control: Sequence[float | Sequence[float]], treatment: Sequence[float | Sequence[float]], *, confidence_level: float = 0.95, control_name: str = 'control', treatment_name: str = 'treatment') -> StatisticalResult
Compare two variants evaluated on the same inputs.
Each element of control / treatment is the score for one input,
or a list of scores from repeated runs on that input. Repeated runs are
averaged per input first, so the input (not the run) is the unit of
analysis.
- Pass/fail scores with one run per input: exact McNemar test with an Agresti-Min confidence interval for the difference in accuracy.
- Anything else: paired sign-flip permutation test on the per-input mean differences with a BCa bootstrap confidence interval (inputs are resampled as whole clusters, so repeated runs are not pseudo-replicated).
Returns:
| Type | Description |
|---|---|
StatisticalResult
|
A StatisticalResult whose |
StatisticalResult
|
(in accuracy points for pass/fail data), with |
StatisticalResult
|
|
StatisticalResult
|
equals |
StatisticalResult
|
control_mean`` or None when the control mean is 0. |
flowprompt.testing.paired.mcnemar_exact ¶
Two-sided exact McNemar p-value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
b
|
int
|
Inputs where the control was right and the treatment wrong. |
required |
c
|
int
|
Inputs where the control was wrong and the treatment right. |
required |
mid_p
|
bool
|
Return the mid-p variant (less conservative; Fagerland, Lydersen and Laake 2013 recommend it for small samples). |
False
|
Returns:
| Type | Description |
|---|---|
float
|
The two-sided p-value |
float
|
|
Example
round(mcnemar_exact(4, 0), 4) 0.125
flowprompt.testing.paired.paired_sign_flip_test ¶
paired_sign_flip_test(differences: Sequence[float], *, seed: int = 0, draws: int = _MONTE_CARLO_DRAWS) -> tuple[float, str, float]
Two-sided paired sign-flip permutation test for a zero mean difference.
Under the null hypothesis that the two variants are interchangeable on every input, the sign of each per-input difference is a fair coin flip. The p-value is the probability, over all sign assignments, of a total at least as extreme as the observed one. Inputs where both variants tie (difference 0) do not change the test statistic.
The p-value is exact when the differences lie on a common grid (pass/fail scores averaged over a fixed number of runs) or when at most 16 inputs differ; otherwise it is estimated by Monte Carlo with a fixed seed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
differences
|
Sequence[float]
|
Per-input differences (treatment - control). |
required |
seed
|
int
|
Seed for the Monte Carlo branch (results are reproducible). |
0
|
draws
|
int
|
Number of Monte Carlo sign assignments. |
_MONTE_CARLO_DRAWS
|
Returns:
| Type | Description |
|---|---|
float
|
|
str
|
|
float
|
the Monte Carlo estimate (0.0 when exact). |
flowprompt.testing.paired.paired_bootstrap_interval ¶
paired_bootstrap_interval(differences: Sequence[float], confidence_level: float = 0.95, *, resamples: int = 4000, seed: int = 0) -> tuple[float, float]
BCa bootstrap confidence interval for the mean per-input difference.
Resamples whole inputs (a cluster bootstrap when an input has several runs), then applies the bias-corrected and accelerated adjustment of Efron (1987), JASA 82:171-185. Inputs where the variants tie contribute a difference of 0 and are kept. Deterministic for a given seed.
flowprompt.testing.paired.agresti_min_interval ¶
Confidence interval for p_treatment - p_control with paired data.
Adds 1/2 to each of the four cells of the paired 2x2 table, then uses the Wald interval (Agresti & Min 2005). Good coverage even for small n.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
b
|
int
|
Control right, treatment wrong. |
required |
c
|
int
|
Control wrong, treatment right. |
required |
n
|
int
|
Total number of paired inputs. |
required |
confidence_level
|
float
|
Confidence level of the interval. |
0.95
|
flowprompt.testing.paired.holm_adjust ¶
Holm-Bonferroni step-down adjusted p-values (same order as input).
Controls the family-wise error rate at alpha under arbitrary dependence between the tests.
Example
holm_adjust([0.01, 0.04, 0.03]) [0.03, 0.06, 0.06]
flowprompt.testing.paired.plan_sample_size ¶
plan_sample_size(min_detectable_difference: float = 0.1, *, discordance: float | None = None, baseline_accuracy: float | None = None, power: float = 0.8, alpha: float = 0.05, sd_difference: float | None = None, n_variants: int = 2, runs_per_input: int = 1, cost_per_call: float | None = None, model: str | None = None, input_tokens_per_call: int | None = None, output_tokens_per_call: int | None = None) -> SampleSizePlan
Plan how many inputs a paired prompt comparison needs.
For pass/fail outcomes the paired-proportions formula of Connor (1987) is used::
n = (z_{1-a/2} * sqrt(psi) + z_{power} * sqrt(psi - d^2))^2 / d^2
where d is the difference to detect and psi the discordance rate
(fraction of inputs where the two variants disagree). The discordance
matters as much as the accuracies: two prompts that fail on the same hard
inputs disagree rarely and need far fewer inputs. Estimate it from a
small pilot (ComparisonResult.sample_size_plan uses the discordance it
observed). If it is not given, it is computed from baseline_accuracy
(default 0.5) assuming the variants err independently, which is the
conservative choice for prompts that are positively correlated.
For numeric scores pass sd_difference (standard deviation of the
per-input differences) instead; then n = ((z_{1-a/2} + z_{power}) *
sd / d)^2.
With more than two variants, alpha is split across the
n_variants - 1 comparisons against the control (Bonferroni), a
conservative stand-in for the Holm procedure used in the analysis.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
min_detectable_difference
|
float
|
Absolute difference to detect, e.g. 0.1 for 10 accuracy points. |
0.1
|
discordance
|
float | None
|
Expected fraction of inputs where the variants disagree. |
None
|
baseline_accuracy
|
float | None
|
Expected accuracy of the control, used only to derive a default discordance (default 0.5). |
None
|
power
|
float
|
Target power (default 0.8). |
0.8
|
alpha
|
float
|
Two-sided significance level (default 0.05). |
0.05
|
sd_difference
|
float | None
|
For numeric scores: SD of per-input differences. |
None
|
n_variants
|
int
|
Variants evaluated on each input (default 2). |
2
|
runs_per_input
|
int
|
Runs per input and variant (default 1). Repeats do
not reduce |
1
|
cost_per_call
|
float | None
|
Known average cost of one call in USD. |
None
|
model
|
str | None
|
Model name to look up a price (litellm cost map) when
|
None
|
input_tokens_per_call
|
int | None
|
Average prompt tokens per call (for pricing). |
None
|
output_tokens_per_call
|
int | None
|
Average completion tokens per call. |
None
|
Returns:
| Type | Description |
|---|---|
SampleSizePlan
|
A SampleSizePlan. |
SampleSizePlan
|
known. |
Example
plan_sample_size(0.1, discordance=0.2).n_inputs 155
flowprompt.testing.paired.SampleSizePlan
dataclass
¶
How many inputs a paired comparison needs.
Attributes:
| Name | Type | Description |
|---|---|---|
n_inputs |
int
|
Inputs needed (each evaluated by every variant). |
min_detectable_difference |
float
|
Absolute difference the plan targets. |
power |
float
|
Target probability of detecting that difference. |
alpha |
float
|
Significance level per comparison (after Bonferroni for several treatments). |
n_variants |
int
|
Number of variants that will run on each input. |
runs_per_input |
int
|
Runs per input and variant. |
total_calls |
int
|
|
estimated_cost_usd |
float | None
|
Estimated cost of the calls, or None if no price is known. |
method |
str
|
Formula used. |
assumptions |
str
|
Plain-language statement of the assumptions. |
flowprompt.testing.paired.SequentialMcNemar ¶
Always-valid McNemar test you may check after every input.
Classical p-values are only valid if you look once, at a sample size fixed in advance. "Peeking" after every batch and stopping at the first p < 0.05 inflates the false-positive rate far beyond 5%. This class tracks a test martingale instead:
M_n = Integral prod_i [theta^x_i (1 - theta)^(1 - x_i) / (1/2)] dBeta(theta; a, a)
= 2^n_d * B(a + c, a + b) / B(a, a)
where b/c count discordant inputs favouring the control/treatment
and n_d = b + c. Under the null hypothesis (no difference between
variants) M_n is a non-negative martingale with M_0 = 1, so by
Ville's inequality P(sup_n M_n >= 1/alpha) <= alpha. Stopping the
first time M_n >= 1/alpha therefore keeps the false-positive rate at
most alpha no matter how often you look. The always-valid p-value is
min(1, 1 / max_k M_k).
The price of continuous monitoring is power: for a fixed sample size the sequential test needs more data than the one-look exact McNemar test.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
alpha
|
float
|
Significance level (default 0.05). |
0.05
|
prior_strength
|
float
|
|
1.0
|
Example
test = SequentialMcNemar(alpha=0.05) for control_ok, treatment_ok in [(False, True)] * 9: ... test.update(control_ok, treatment_ok) test.rejected True
update ¶
Add one paired observation and return rejected.
flowprompt.testing.statistics.StatisticalResult
dataclass
¶
Result of statistical significance test.
Attributes:
| Name | Type | Description |
|---|---|---|
significant |
bool
|
Whether the result is statistically significant (using
|
p_value |
float
|
Raw (unadjusted) p-value of the test. |
confidence_level |
float
|
Confidence level used. |
effect_size |
float
|
Estimated effect size. For the paired tests used by
|
confidence_interval |
tuple[float, float] | None
|
Confidence interval for the difference. |
power |
float | None
|
Statistical power of the test. |
sample_size_recommendation |
int | None
|
Recommended sample size if not significant. |
test_name |
str
|
Name of the statistical test used. |
details |
dict[str, Any]
|
Additional test details. |
method |
str
|
Name of the method (e.g. |
difference |
float | None
|
Treatment minus control (absolute). |
ci_low |
float | None
|
Lower confidence bound for |
ci_high |
float | None
|
Upper confidence bound for |
n_inputs |
int | None
|
Number of paired inputs the test is based on. |
adjusted_p |
float | None
|
p-value after the multiple-comparison correction (Holm), or None when there was a single comparison. |
relative_lift |
float | None
|
|
control |
str | None
|
Name of the control variant. |
treatment |
str | None
|
Name of the treatment variant. |
Variants, scorers and offline testing¶
flowprompt.testing.variants ¶
Variants: anything that turns an input into an output.
The comparison engine only needs a callable variant(input: dict) -> output
(sync or async). Prompt classes and "same prompt, different model" setups
are thin adapters onto that protocol:
compare(
{
"baseline": ExtractUser, # Prompt class, default model
"mini": (ExtractUser, "gpt-4o-mini"), # Prompt class on another model
"regex": lambda inp: my_rule_based(inp["text"]), # any callable
"pipeline": PromptVariant(ExtractUser, model="gpt-4o", temperature=0.2),
},
inputs=..., expected=...,
)
LLM calls made through FlowPrompt inside any variant (including your own callables) are metered automatically, so cost per variant works for custom pipelines too.
PromptVariant
dataclass
¶
Adapter running a Prompt class on one model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prompt
|
type
|
The Prompt subclass. Each input dict is passed as keyword arguments to its constructor. |
required |
model
|
str | None
|
Model identifier. None means "use compare()'s |
None
|
temperature
|
float | None
|
Sampling temperature. None means "use compare()'s". |
None
|
run_kwargs
|
dict[str, Any]
|
Extra keyword arguments for |
dict()
|
FunctionVariant
dataclass
¶
Adapter for a plain callable fn(input: dict) -> output (sync or async).
as_variant ¶
Adapt a user-supplied variant definition to the engine protocol.
Accepts a Prompt subclass, a (PromptClass, "model") tuple, a
:class:PromptVariant, or any callable taking the input dict.
model_variants ¶
model_variants(prompt: type, models: list[str], *, temperature: float | None = None) -> dict[str, PromptVariant]
Variants that run the same Prompt class on several models.
Example
compare(model_variants(ExtractUser, ["gpt-4o-mini", "gpt-4o"]), inputs, ...)
flowprompt.testing.scorers ¶
Scorers: how an output is graded against the expected answer.
A scorer is any callable (output, expected) -> bool | float. True /
False (or 1.0 / 0.0) give pass/fail outcomes, which compare() tests
with the exact McNemar test; other floats are treated as numeric scores and
tested with a paired permutation test.
>>> from flowprompt.testing import scorers
>>> compare(variants, inputs, expected=answers, eval_metric=scorers.numeric(abs_tol=0.01))
String shortcuts accepted by compare(eval_metric=...): "exact",
"contains", "regex", "numeric", "similarity".
exact ¶
Output equals the expected value.
Strings are compared after stripping whitespace and, by default,
case-insensitively. A structured (pydantic) output matches a dict
expectation when output.model_dump() == expected.
contains ¶
The expected value appears somewhere in the output.
regex ¶
The expected value is a regular expression the output must match.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flags
|
int
|
|
IGNORECASE
|
full_match
|
bool
|
Require the whole (stripped) output to match instead of searching for the pattern anywhere. |
False
|
numeric ¶
Output is numerically close to the expected number.
Numbers are extracted from text outputs (the last number in the text is used, so "The answer is 42." scores against 42).
similarity ¶
Character-level similarity ratio (difflib) of at least threshold.
resolve_scorer ¶
Turn a scorer name or callable into a scorer callable.
flowprompt.testing.fake_llm.FakeLLM ¶
Context manager that answers LLM calls locally.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
responder
|
Responder | str | None
|
Produces the reply. Either a string (always returned), or
a callable receiving the list of chat messages (and, if it accepts
a second argument, the full request as a dict, e.g. to check
|
None
|
latency_s
|
float
|
Optional artificial delay per call, in seconds. |
0.0
|
Attributes:
| Name | Type | Description |
|---|---|---|
calls |
list[dict[str, Any]]
|
The keyword arguments of every call made while active. |
Usage, caching and tracing¶
flowprompt.core.usage ¶
Token usage and cost accounting for LLM calls.
Every call made through Prompt.run() / Prompt.arun() records a
:class:CallUsage (tokens, cost, latency, cache hit) to any active
:func:track_usage collector. compare() uses this to report real cost
per variant; you can use it directly:
>>> from flowprompt import track_usage
>>> with track_usage() as calls:
... MyPrompt(text="hi").run(model="gpt-4o-mini")
>>> sum(c.cost_usd or 0 for c in calls)
Collectors are stored in a :class:contextvars.ContextVar, so they are
isolated per thread and per asyncio task and can be nested.
CallUsage
dataclass
¶
Usage of one LLM call.
Attributes:
| Name | Type | Description |
|---|---|---|
model |
str
|
Model identifier passed to the provider. |
prompt_tokens |
int
|
Input tokens reported by the provider. |
completion_tokens |
int
|
Output tokens reported by the provider. |
cost_usd |
float | None
|
Cost in USD, or None when no price is known for the model. Cache hits cost 0. |
latency_ms |
float
|
Wall-clock latency of the call. |
cached |
bool
|
True when the response came from the FlowPrompt cache. |
track_usage ¶
Collect the usage of every LLM call made inside the block.
estimate_cost ¶
Cost of a call from litellm's model price map, or None if unknown.
usage_from_response ¶
Build a CallUsage from a litellm (OpenAI-style) response object.
flowprompt.core.cache.configure_cache ¶
configure_cache(backend: CacheBackend | None = None, default_ttl: float | None = 3600, enabled: bool = True) -> PromptCache
Configure the global cache used by Prompt.run() and arun().
After this call, identical requests (same messages, model, temperature, output schema and generation parameters) are answered from the cache without an LLM call. Streaming calls are not cached.
Note
A cache returns the same answer for repeated identical requests. If
you use compare(..., runs_per_input>1) to measure run-to-run
variation at temperature > 0, disable the cache for that run.
flowprompt.tracing.otel.configure_tracer ¶
Configure the global tracer.