Skip to content

Python API

Generated from the docstrings. For a guided tour see the API overview.

Comparing variants

flowprompt.testing.compare.compare

compare(prompts: dict[str, Any], inputs: list[dict[str, Any]], model: str = 'gpt-4o', *, expected: list[Any] | None = None, eval_metric: str | Callable[..., Any] = 'contains', success_fn: Callable[..., Any] | None = None, metric_fn: Callable[[Any], float] | None = None, confidence_level: float = 0.95, runs_per_input: int = 1, temperature: float = 0.0, test_type: str = 'auto', dry_run: bool = False, control: str | None = None, comparisons: str = 'control') -> ComparisonResult

Compare variants on the same inputs with paired significance tests.

Parameters:

Name Type Description Default
prompts dict[str, Any]

Dict mapping variant names to variants: Prompt subclasses, (PromptClass, "model") tuples, :class:PromptVariant objects or any callable fn(input: dict) -> output (sync or async).

required
inputs list[dict[str, Any]]

Input dicts; every variant runs on every input.

required
model str

Default model for Prompt variants.

'gpt-4o'
expected list[Any] | None

Expected outputs (same length as inputs). Outputs are graded with eval_metric.

None
eval_metric str | Callable[..., Any]

Scorer for (output, expected): "exact", "contains" (default), "regex", "numeric", "similarity", or a callable returning bool (pass/fail) or a float score. See :mod:flowprompt.testing.scorers.

'contains'
success_fn Callable[..., Any] | None

Custom pass/fail function (output) or, with expected, (output, expected). Takes precedence.

None
metric_fn Callable[[Any], float] | None

Numeric score (output) -> float. Used as the outcome when neither expected nor success_fn is given.

None
confidence_level float

Confidence level (default 0.95, i.e. alpha 0.05).

0.95
runs_per_input int

Runs per input and variant. Repeats are averaged per input; they measure run-to-run variation but do not increase the sample size, which is the number of inputs.

1
temperature float

Default temperature for Prompt variants.

0.0
test_type str

"auto" (default): exact McNemar for pass/fail with one run per input, otherwise a paired sign-flip permutation test. "z_test", "chi_squared", "t_test" and "bayesian" select the deprecated unpaired tests.

'auto'
dry_run bool

Only estimate the cost; make no calls.

False
control str | None

Name of the baseline variant (default: the first one).

None
comparisons str

"control" (default) compares each variant with the control; "all" compares every pair. Holm's correction is applied whenever there is more than one comparison.

'control'

Returns:

Name Type Description
A ComparisonResult

class:ComparisonResult. print(result) shows a report;

ComparisonResult

result.save_report("report.html") writes Markdown or HTML.

Raises:

Type Description
ValueError

On fewer than 2 variants, no inputs, mismatched expected or an unknown option.

Example

result = compare( ... {"short": Short, "detailed": Detailed}, ... inputs=[{"text": t} for t in texts], ... expected=labels, ... eval_metric="exact", ... model="gpt-4o-mini", ... ) print(result.verdict)

flowprompt.testing.compare.acompare async

acompare(prompts: dict[str, Any], inputs: list[dict[str, Any]], model: str = 'gpt-4o', *, expected: list[Any] | None = None, eval_metric: str | Callable[..., Any] = 'contains', success_fn: Callable[..., Any] | None = None, metric_fn: Callable[[Any], float] | None = None, confidence_level: float = 0.95, runs_per_input: int = 1, temperature: float = 0.0, test_type: str = 'auto', dry_run: bool = False, control: str | None = None, comparisons: str = 'control') -> ComparisonResult

Async version of :func:compare; variants run concurrently.

Same arguments and result. Prompt variants use Prompt.arun(); async callables are awaited; sync callables run in a worker thread.

flowprompt.testing.compare.ComparisonResult dataclass

Result of comparing variants.

Attributes:

Name Type Description
winner str | None

Name of the variant that is significantly better than the others (after multiple-comparison adjustment), or None.

variants dict[str, VariantResult]

Per-variant results keyed by name.

statistical_result StatisticalResult | None

The comparison that decides the outcome (the winner's comparison, or the closest call when there is none).

confidence_level float

Confidence level used.

total_runs int

Total number of runs across all variants.

estimated_cost dict[str, Any] | None

Pre-run cost estimate (also filled in dry-run mode).

has_expected bool

True when expected outputs were provided.

comparisons list[StatisticalResult]

All pairwise comparisons that were tested.

control str | None

Name of the control (baseline) variant.

outcome str

What was measured: "accuracy" (expected outputs), "success" (success_fn), "score" (numeric metric) or "no_error" (nothing to grade against).

has_ground_truth bool

False when outputs could not be graded (no expected outputs, success_fn or metric_fn).

n_inputs int

Number of inputs.

runs_per_input int

Runs per input and variant.

correction str | None

Multiple-comparison correction applied ("holm") or None for a single comparison.

comparison_mode str

"control" (each variant vs the control) or "all" (all pairs).

model str | None

Default model used for Prompt variants.

notes list[str]

Caveats worth showing next to the result.

total_cost_usd property

total_cost_usd: float

Metered cost of all runs (0.0 when no price was known).

is_proportion property

is_proportion: bool

True when scores are pass/fail rates (shown in points).

enough_data property

enough_data: bool

False when no difference could reach significance at this size.

With n inputs the smallest attainable two-sided p-value of the paired tests is 2 / 2**n (every input favours the same variant); with m comparisons Holm multiplies it by up to m.

verdict property

verdict: str

One plain-English sentence summarising the outcome.

sample_size_plan

sample_size_plan(min_detectable_difference: float = 0.1, *, power: float = 0.8) -> SampleSizePlan | None

Inputs needed to detect a given difference, using this run as a pilot.

Uses the discordance (pass/fail) or the SD of per-input differences (numeric scores) observed in the deciding comparison, and the observed cost per run when it is known.

__str__

__str__() -> str

Plain-text report.

to_markdown

to_markdown() -> str

Markdown report (for PR comments, job summaries, notebooks).

to_html

to_html() -> str

Self-contained HTML report (no external assets).

save_report

save_report(path: str | Path) -> Path

Write the report; the format follows the extension (.html/.md/.txt/.json).

to_dict

to_dict() -> dict[str, Any]

Serialize to a dictionary.

flowprompt.testing.compare.VariantResult dataclass

Results for a single variant.

Attributes:

Name Type Description
name str

Variant name.

samples int

Number of runs completed (inputs x runs_per_input).

successes int

Runs that passed (pass/fail outcomes) or completed without error (numeric outcomes).

success_rate float

successes / samples.

mean_latency_ms float

Average latency per run in milliseconds.

total_cost_usd float

Total metered cost in USD (0.0 when unknown; see cost_known).

outputs list[Any]

Outputs of the runs that did not raise.

errors list[str]

Error messages of runs that raised (counted as failures).

label str

Human-readable description (prompt class, model, function).

mean_score float | None

Input-level mean score (accuracy for pass/fail).

ci_low float | None

Lower confidence bound for mean_score.

ci_high float | None

Upper confidence bound for mean_score.

p95_latency_ms float

95th percentile latency per run.

cost_known bool

True when every LLM call had a known price.

cost_per_correct float | None

total_cost_usd / number of correct answers.

total_tokens int

Tokens used across all LLM calls.

llm_calls int

Number of LLM calls (including retries and cache hits).

scores list[list[float]]

Per-input lists of run scores (errors scored 0).

records list[RunRecord]

Every individual run.

flowprompt.testing.compare.estimate_compare_cost

estimate_compare_cost(prompts: dict[str, Any], inputs: list[dict[str, Any]], model: str = 'gpt-4o', *, runs_per_input: int = 1, estimated_output_tokens: int = 100) -> dict[str, Any]

Estimate the cost of running compare() without making API calls.

Prompt variants are rendered for each input and their input tokens are counted with litellm; prices come from litellm's model cost map. Custom callables cannot be inspected, so they count calls but no tokens.

Returns:

Type Description
dict[str, Any]

Dict with keys: model, total_calls, estimated_input_tokens,

dict[str, Any]

estimated_output_tokens, estimated_cost_usd, per_variant. Cost fields

dict[str, Any]

are None when a price is unknown.

Paired statistics

flowprompt.testing.paired.paired_test

paired_test(control: Sequence[float | Sequence[float]], treatment: Sequence[float | Sequence[float]], *, confidence_level: float = 0.95, control_name: str = 'control', treatment_name: str = 'treatment') -> StatisticalResult

Compare two variants evaluated on the same inputs.

Each element of control / treatment is the score for one input, or a list of scores from repeated runs on that input. Repeated runs are averaged per input first, so the input (not the run) is the unit of analysis.

  • Pass/fail scores with one run per input: exact McNemar test with an Agresti-Min confidence interval for the difference in accuracy.
  • Anything else: paired sign-flip permutation test on the per-input mean differences with a BCa bootstrap confidence interval (inputs are resampled as whole clusters, so repeated runs are not pseudo-replicated).

Returns:

Type Description
StatisticalResult

A StatisticalResult whose difference is treatment minus control

StatisticalResult

(in accuracy points for pass/fail data), with ci_low/ci_high,

StatisticalResult

p_value, method and n_inputs filled in. effect_size

StatisticalResult

equals difference; relative_lift is ``difference /

StatisticalResult

control_mean`` or None when the control mean is 0.

flowprompt.testing.paired.mcnemar_exact

mcnemar_exact(b: int, c: int, *, mid_p: bool = False) -> float

Two-sided exact McNemar p-value.

Parameters:

Name Type Description Default
b int

Inputs where the control was right and the treatment wrong.

required
c int

Inputs where the control was wrong and the treatment right.

required
mid_p bool

Return the mid-p variant (less conservative; Fagerland, Lydersen and Laake 2013 recommend it for small samples).

False

Returns:

Type Description
float

The two-sided p-value 2 * P(X <= min(b, c)) (capped at 1) with

float

X ~ Binomial(b + c, 1/2).

Example

round(mcnemar_exact(4, 0), 4) 0.125

flowprompt.testing.paired.paired_sign_flip_test

paired_sign_flip_test(differences: Sequence[float], *, seed: int = 0, draws: int = _MONTE_CARLO_DRAWS) -> tuple[float, str, float]

Two-sided paired sign-flip permutation test for a zero mean difference.

Under the null hypothesis that the two variants are interchangeable on every input, the sign of each per-input difference is a fair coin flip. The p-value is the probability, over all sign assignments, of a total at least as extreme as the observed one. Inputs where both variants tie (difference 0) do not change the test statistic.

The p-value is exact when the differences lie on a common grid (pass/fail scores averaged over a fixed number of runs) or when at most 16 inputs differ; otherwise it is estimated by Monte Carlo with a fixed seed.

Parameters:

Name Type Description Default
differences Sequence[float]

Per-input differences (treatment - control).

required
seed int

Seed for the Monte Carlo branch (results are reproducible).

0
draws int

Number of Monte Carlo sign assignments.

_MONTE_CARLO_DRAWS

Returns:

Type Description
float

(p_value, how, monte_carlo_error) where how is "exact" or

str

"monte_carlo" and monte_carlo_error is the standard error of

float

the Monte Carlo estimate (0.0 when exact).

flowprompt.testing.paired.paired_bootstrap_interval

paired_bootstrap_interval(differences: Sequence[float], confidence_level: float = 0.95, *, resamples: int = 4000, seed: int = 0) -> tuple[float, float]

BCa bootstrap confidence interval for the mean per-input difference.

Resamples whole inputs (a cluster bootstrap when an input has several runs), then applies the bias-corrected and accelerated adjustment of Efron (1987), JASA 82:171-185. Inputs where the variants tie contribute a difference of 0 and are kept. Deterministic for a given seed.

flowprompt.testing.paired.agresti_min_interval

agresti_min_interval(b: int, c: int, n: int, confidence_level: float = 0.95) -> tuple[float, float]

Confidence interval for p_treatment - p_control with paired data.

Adds 1/2 to each of the four cells of the paired 2x2 table, then uses the Wald interval (Agresti & Min 2005). Good coverage even for small n.

Parameters:

Name Type Description Default
b int

Control right, treatment wrong.

required
c int

Control wrong, treatment right.

required
n int

Total number of paired inputs.

required
confidence_level float

Confidence level of the interval.

0.95

flowprompt.testing.paired.holm_adjust

holm_adjust(p_values: Sequence[float]) -> list[float]

Holm-Bonferroni step-down adjusted p-values (same order as input).

Controls the family-wise error rate at alpha under arbitrary dependence between the tests.

Example

holm_adjust([0.01, 0.04, 0.03]) [0.03, 0.06, 0.06]

flowprompt.testing.paired.plan_sample_size

plan_sample_size(min_detectable_difference: float = 0.1, *, discordance: float | None = None, baseline_accuracy: float | None = None, power: float = 0.8, alpha: float = 0.05, sd_difference: float | None = None, n_variants: int = 2, runs_per_input: int = 1, cost_per_call: float | None = None, model: str | None = None, input_tokens_per_call: int | None = None, output_tokens_per_call: int | None = None) -> SampleSizePlan

Plan how many inputs a paired prompt comparison needs.

For pass/fail outcomes the paired-proportions formula of Connor (1987) is used::

n = (z_{1-a/2} * sqrt(psi) + z_{power} * sqrt(psi - d^2))^2 / d^2

where d is the difference to detect and psi the discordance rate (fraction of inputs where the two variants disagree). The discordance matters as much as the accuracies: two prompts that fail on the same hard inputs disagree rarely and need far fewer inputs. Estimate it from a small pilot (ComparisonResult.sample_size_plan uses the discordance it observed). If it is not given, it is computed from baseline_accuracy (default 0.5) assuming the variants err independently, which is the conservative choice for prompts that are positively correlated.

For numeric scores pass sd_difference (standard deviation of the per-input differences) instead; then n = ((z_{1-a/2} + z_{power}) * sd / d)^2.

With more than two variants, alpha is split across the n_variants - 1 comparisons against the control (Bonferroni), a conservative stand-in for the Holm procedure used in the analysis.

Parameters:

Name Type Description Default
min_detectable_difference float

Absolute difference to detect, e.g. 0.1 for 10 accuracy points.

0.1
discordance float | None

Expected fraction of inputs where the variants disagree.

None
baseline_accuracy float | None

Expected accuracy of the control, used only to derive a default discordance (default 0.5).

None
power float

Target power (default 0.8).

0.8
alpha float

Two-sided significance level (default 0.05).

0.05
sd_difference float | None

For numeric scores: SD of per-input differences.

None
n_variants int

Variants evaluated on each input (default 2).

2
runs_per_input int

Runs per input and variant (default 1). Repeats do not reduce n_inputs; they only add calls.

1
cost_per_call float | None

Known average cost of one call in USD.

None
model str | None

Model name to look up a price (litellm cost map) when cost_per_call is not given.

None
input_tokens_per_call int | None

Average prompt tokens per call (for pricing).

None
output_tokens_per_call int | None

Average completion tokens per call.

None

Returns:

Type Description
SampleSizePlan

A SampleSizePlan. estimated_cost_usd is None when no price is

SampleSizePlan

known.

Example

plan_sample_size(0.1, discordance=0.2).n_inputs 155

flowprompt.testing.paired.SampleSizePlan dataclass

How many inputs a paired comparison needs.

Attributes:

Name Type Description
n_inputs int

Inputs needed (each evaluated by every variant).

min_detectable_difference float

Absolute difference the plan targets.

power float

Target probability of detecting that difference.

alpha float

Significance level per comparison (after Bonferroni for several treatments).

n_variants int

Number of variants that will run on each input.

runs_per_input int

Runs per input and variant.

total_calls int

n_inputs * n_variants * runs_per_input.

estimated_cost_usd float | None

Estimated cost of the calls, or None if no price is known.

method str

Formula used.

assumptions str

Plain-language statement of the assumptions.

summary

summary() -> str

One-line plain-English summary.

flowprompt.testing.paired.SequentialMcNemar

Always-valid McNemar test you may check after every input.

Classical p-values are only valid if you look once, at a sample size fixed in advance. "Peeking" after every batch and stopping at the first p < 0.05 inflates the false-positive rate far beyond 5%. This class tracks a test martingale instead:

M_n = Integral prod_i [theta^x_i (1 - theta)^(1 - x_i) / (1/2)] dBeta(theta; a, a)
    = 2^n_d * B(a + c, a + b) / B(a, a)

where b/c count discordant inputs favouring the control/treatment and n_d = b + c. Under the null hypothesis (no difference between variants) M_n is a non-negative martingale with M_0 = 1, so by Ville's inequality P(sup_n M_n >= 1/alpha) <= alpha. Stopping the first time M_n >= 1/alpha therefore keeps the false-positive rate at most alpha no matter how often you look. The always-valid p-value is min(1, 1 / max_k M_k).

The price of continuous monitoring is power: for a fixed sample size the sequential test needs more data than the one-look exact McNemar test.

Parameters:

Name Type Description Default
alpha float

Significance level (default 0.05).

0.05
prior_strength float

a in the symmetric Beta(a, a) mixing prior (default 1, uniform). Larger values favour detecting smaller effects later; smaller values favour large effects early.

1.0
Example

test = SequentialMcNemar(alpha=0.05) for control_ok, treatment_ok in [(False, True)] * 9: ... test.update(control_ok, treatment_ok) test.rejected True

log_evidence property

log_evidence: float

log M_n for the current data.

evidence property

evidence: float

The test martingale M_n (likelihood-ratio evidence against H0).

p_value property

p_value: float

Always-valid p-value: min(1, 1 / max_k M_k).

rejected property

rejected: bool

True once the evidence has crossed 1/alpha at any look.

difference property

difference: float

Observed accuracy difference (treatment - control) so far.

update

update(control_correct: bool, treatment_correct: bool) -> bool

Add one paired observation and return rejected.

flowprompt.testing.statistics.StatisticalResult dataclass

Result of statistical significance test.

Attributes:

Name Type Description
significant bool

Whether the result is statistically significant (using adjusted_p when a multiple-comparison correction was applied).

p_value float

Raw (unadjusted) p-value of the test.

confidence_level float

Confidence level used.

effect_size float

Estimated effect size. For the paired tests used by compare() this is the absolute difference (treatment minus control, e.g. 0.12 = 12 accuracy points). The legacy unpaired tests report the relative improvement here.

confidence_interval tuple[float, float] | None

Confidence interval for the difference.

power float | None

Statistical power of the test.

sample_size_recommendation int | None

Recommended sample size if not significant.

test_name str

Name of the statistical test used.

details dict[str, Any]

Additional test details.

method str

Name of the method (e.g. "mcnemar_exact").

difference float | None

Treatment minus control (absolute).

ci_low float | None

Lower confidence bound for difference.

ci_high float | None

Upper confidence bound for difference.

n_inputs int | None

Number of paired inputs the test is based on.

adjusted_p float | None

p-value after the multiple-comparison correction (Holm), or None when there was a single comparison.

relative_lift float | None

difference / control_rate; None when undefined.

control str | None

Name of the control variant.

treatment str | None

Name of the treatment variant.

summary

summary() -> str

Generate a human-readable summary.

Variants, scorers and offline testing

flowprompt.testing.variants

Variants: anything that turns an input into an output.

The comparison engine only needs a callable variant(input: dict) -> output (sync or async). Prompt classes and "same prompt, different model" setups are thin adapters onto that protocol:

compare(
    {
        "baseline": ExtractUser,                     # Prompt class, default model
        "mini": (ExtractUser, "gpt-4o-mini"),         # Prompt class on another model
        "regex": lambda inp: my_rule_based(inp["text"]),  # any callable
        "pipeline": PromptVariant(ExtractUser, model="gpt-4o", temperature=0.2),
    },
    inputs=..., expected=...,
)

LLM calls made through FlowPrompt inside any variant (including your own callables) are metered automatically, so cost per variant works for custom pipelines too.

PromptVariant dataclass

Adapter running a Prompt class on one model.

Parameters:

Name Type Description Default
prompt type

The Prompt subclass. Each input dict is passed as keyword arguments to its constructor.

required
model str | None

Model identifier. None means "use compare()'s model".

None
temperature float | None

Sampling temperature. None means "use compare()'s".

None
run_kwargs dict[str, Any]

Extra keyword arguments for Prompt.run().

dict()
bind
bind(model: str, temperature: float) -> PromptVariant

Fill in defaults from compare().

FunctionVariant dataclass

Adapter for a plain callable fn(input: dict) -> output (sync or async).

as_variant

as_variant(obj: Any, *, model: str, temperature: float) -> Variant

Adapt a user-supplied variant definition to the engine protocol.

Accepts a Prompt subclass, a (PromptClass, "model") tuple, a :class:PromptVariant, or any callable taking the input dict.

model_variants

model_variants(prompt: type, models: list[str], *, temperature: float | None = None) -> dict[str, PromptVariant]

Variants that run the same Prompt class on several models.

Example

compare(model_variants(ExtractUser, ["gpt-4o-mini", "gpt-4o"]), inputs, ...)

flowprompt.testing.scorers

Scorers: how an output is graded against the expected answer.

A scorer is any callable (output, expected) -> bool | float. True / False (or 1.0 / 0.0) give pass/fail outcomes, which compare() tests with the exact McNemar test; other floats are treated as numeric scores and tested with a paired permutation test.

>>> from flowprompt.testing import scorers
>>> compare(variants, inputs, expected=answers, eval_metric=scorers.numeric(abs_tol=0.01))

String shortcuts accepted by compare(eval_metric=...): "exact", "contains", "regex", "numeric", "similarity".

exact

exact(*, case_sensitive: bool = False, strip: bool = True) -> Scorer

Output equals the expected value.

Strings are compared after stripping whitespace and, by default, case-insensitively. A structured (pydantic) output matches a dict expectation when output.model_dump() == expected.

contains

contains(*, case_sensitive: bool = False) -> Scorer

The expected value appears somewhere in the output.

regex

regex(*, flags: int = re.IGNORECASE, full_match: bool = False) -> Scorer

The expected value is a regular expression the output must match.

Parameters:

Name Type Description Default
flags int

re flags (default case-insensitive).

IGNORECASE
full_match bool

Require the whole (stripped) output to match instead of searching for the pattern anywhere.

False

numeric

numeric(*, abs_tol: float = 1e-06, rel_tol: float = 0.0) -> Scorer

Output is numerically close to the expected number.

Numbers are extracted from text outputs (the last number in the text is used, so "The answer is 42." scores against 42).

similarity

similarity(threshold: float = 0.7) -> Scorer

Character-level similarity ratio (difflib) of at least threshold.

resolve_scorer

resolve_scorer(metric: str | Callable[..., Any]) -> Callable[..., Any]

Turn a scorer name or callable into a scorer callable.

flowprompt.testing.fake_llm.FakeLLM

Context manager that answers LLM calls locally.

Parameters:

Name Type Description Default
responder Responder | str | None

Produces the reply. Either a string (always returned), or a callable receiving the list of chat messages (and, if it accepts a second argument, the full request as a dict, e.g. to check request["model"] or request["response_format"]) and returning a string, a dict / pydantic model (serialised to JSON), or None to use the default reply. Defaults to echoing the last user message, or schema-shaped JSON for structured prompts.

None
latency_s float

Optional artificial delay per call, in seconds.

0.0

Attributes:

Name Type Description
calls list[dict[str, Any]]

The keyword arguments of every call made while active.

Usage, caching and tracing

flowprompt.core.usage

Token usage and cost accounting for LLM calls.

Every call made through Prompt.run() / Prompt.arun() records a :class:CallUsage (tokens, cost, latency, cache hit) to any active :func:track_usage collector. compare() uses this to report real cost per variant; you can use it directly:

>>> from flowprompt import track_usage
>>> with track_usage() as calls:
...     MyPrompt(text="hi").run(model="gpt-4o-mini")
>>> sum(c.cost_usd or 0 for c in calls)

Collectors are stored in a :class:contextvars.ContextVar, so they are isolated per thread and per asyncio task and can be nested.

CallUsage dataclass

Usage of one LLM call.

Attributes:

Name Type Description
model str

Model identifier passed to the provider.

prompt_tokens int

Input tokens reported by the provider.

completion_tokens int

Output tokens reported by the provider.

cost_usd float | None

Cost in USD, or None when no price is known for the model. Cache hits cost 0.

latency_ms float

Wall-clock latency of the call.

cached bool

True when the response came from the FlowPrompt cache.

track_usage

track_usage() -> Iterator[list[CallUsage]]

Collect the usage of every LLM call made inside the block.

record_usage

record_usage(usage: CallUsage) -> None

Append usage to every active collector.

estimate_cost

estimate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float | None

Cost of a call from litellm's model price map, or None if unknown.

usage_from_response

usage_from_response(response: Any, model: str, latency_ms: float = 0.0) -> CallUsage

Build a CallUsage from a litellm (OpenAI-style) response object.

flowprompt.core.cache.configure_cache

configure_cache(backend: CacheBackend | None = None, default_ttl: float | None = 3600, enabled: bool = True) -> PromptCache

Configure the global cache used by Prompt.run() and arun().

After this call, identical requests (same messages, model, temperature, output schema and generation parameters) are answered from the cache without an LLM call. Streaming calls are not cached.

Note

A cache returns the same answer for repeated identical requests. If you use compare(..., runs_per_input>1) to measure run-to-run variation at temperature > 0, disable the cache for that run.

flowprompt.tracing.otel.configure_tracer

configure_tracer(service_name: str = 'flowprompt', enabled: bool = True) -> Tracer

Configure the global tracer.