Changelog¶
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Unreleased¶
Planned¶
- Always-valid (sequential) tests for live-traffic experiments
- Redis cache backend
- Langfuse integration
[0.5.0] - Unreleased¶
This release makes compare() statistically sound and fixes the crash that
broke the README quickstart and every example. If you
used compare() before, re-check conclusions drawn from it: the old
analysis could report significant differences that were not there (see
Fixed), and results now change accordingly.
Fixed¶
- Prompts defined as in the README crashed.
class MyPrompt(Prompt): system = "..."(no type annotation) raisedPydanticUserError: Field 'system' defined on a base class was overridden by a non-annotated attributeon every Pydantic 2 release we tested (2.0.3 through 2.13.5), so the quickstart and all examples failed; the test suite only used annotated fields and did not catch it. Unannotated overrides of prompt fields now work on Pydantic 2.0 through 2.13 and Python 3.10 through 3.14; annotated definitions are unchanged. compare()p-values ignored that variants share inputs. Every variant runs on the same inputs, but results were tested with an unpaired pooled z-test. Pass/fail results now use the exact McNemar test. Example: 8/10 versus 4/10 correct, with the four disagreements all favouring the first prompt, gives p = 0.125 (previously reported as 0.068).- Repeated runs were counted as extra samples. With
runs_per_input=5and deterministic outputs, the p-value for the example above dropped from 0.068 to below 0.0001 without any new information. Runs are now averaged per input; the number of inputs is the sample size. - Wrong winner when the control scored 0. The winner was chosen by the
relative lift, which is undefined (and was set to 0) when the control
scores 0, so a significantly better treatment lost. Winners now follow
the sign of the absolute difference. The same bug in
ABTestRunner(summary and auto-completion) is fixed too. - No correction for multiple comparisons. With three or more variants,
p-values are now Holm-adjusted before a winner is named, in
compare()and inABTestRunner. - Comparisons without expected outputs reported a failed test ("Zero
standard error"). They now say plainly that outputs were not graded and
compare error rates only (
has_ground_truth=False). - Caching, tracing and cost tracking were not connected to
run(). Two identical calls made two LLM calls, the tracer recorded nothing andcompare()reported $0.00.configure_cache()now serves identical requests from the cache, a configured tracer records every call, and token usage and cost are taken from the provider response. - YAML prompts with an
output_schemafailed withPydanticSchemaGenerationError. Output models are now built withpydantic.create_model(enums, nested objects, typed arrays, required fields).
Changed¶
compare()uses paired tests by default (test_type="auto"): the exact McNemar test for pass/fail with one run per input, otherwise a paired sign-flip permutation test with a bootstrap confidence interval. Expect larger, honest p-values on small evaluation sets; with fewer than six inputs no difference can be significant, and the result says so.StatisticalResult.effect_sizefromcompare()is the absolute difference (0.12 = 12 accuracy points); the relative lift moved torelative_lift(None when undefined). The unpaired tests keep their old meaning.test_type="z_test","chi_squared","t_test"and"bayesian"still work withcompare()but emit aDeprecationWarning.print(result)shows a report: accuracy with confidence intervals, latency (mean and p95), cost and cost per correct answer, each paired comparison with its interval and adjusted p-value, and a one-line verdict.PromptTestResult.assert_significant()and.p_valueuse the Holm-adjusted p-value when several variants were compared.- The cache is only used after
configure_cache();get_cache()alone does not turn it on.
Added¶
- Variants can be any callable
fn(input) -> output(sync or async),(PromptClass, "model")tuples,PromptVariantormodel_variants(Prompt, [...]), so models and pipelines can be compared, not only prompts. - Scorers:
exact,contains,regex,numeric(with tolerances),similarity, or any function returning a bool or a score (flowprompt.testing.scorers). ComparisonResult:comparisons,verdict,enough_data,notes,to_markdown(),to_html()(self-contained, light and dark),save_report(),sample_size_plan();StatisticalResultgainsmethod,difference,ci_low,ci_high,n_inputs,adjusted_p,relative_lift,control,treatment.plan_sample_size(): inputs needed for a given difference and power (Connor 1987 for paired pass/fail outcomes), with a cost estimate when a price is known.SequentialMcNemar: an always-valid paired test for stopping early without inflating false positives.- Statistics helpers:
paired_test,mcnemar_exact(with mid-p),paired_sign_flip_test,paired_bootstrap_interval,agresti_min_interval,wilson_interval,holm_adjust. track_usage()andCallUsagefor per-call tokens, cost, latency and cache hits.FakeLLM: answers LLM calls offline and deterministically for tests, CI and demos.flowprompt compare VARIANTS.py DATASET.jsonlwith Markdown, HTML and JSON reports and--fail-on-regressionfor CI.compare(control=..., comparisons="all")to choose the baseline or test every pair.- Documentation site (MkDocs Material), pages on CI, statistical methods
and observability, and a simulation study:
Your prompt A/B test is probably lying to you
(
benchmarks/false_winners.py). - CI runs every example and the README offline, and tests prompt definitions on Pydantic 2.0, 2.9 and the latest 2.x.
Removed¶
- README claims that could not be verified: "<100ms import" (importing
flowprompttakes about 0.15 s; litellm, about 2 s, loads on the first call), "the only Python LLM framework with built-in A/B testing", and the unverified comparison table.
0.4.0 - 2026-03-20¶
Not published to PyPI. This version was tagged on GitHub, but the package version in the source was still 0.3.0, so the PyPI upload could not succeed. Everything listed here ships in 0.5.0.
Added¶
expectedparameter forcompare()andacompare(): Ground-truth evaluation for prompt outputs- Pass
expected=[...]alongsideinputsto measure accuracy instead of just success - Output displays "accuracy" instead of "success" when expected values are provided
- Built-in eval metrics:
exact_match,contains_match,similarity_match(stdlib only, no new deps) eval_metricparameter accepts string names or custom(output, expected) -> boolcallablesresolve_eval_metric()for programmatic metric resolution- Pytest integration (auto-discovered via
pytest11entry point, zero config): fpfixture (session-scoped):FlowPromptHelperwith.compare,.acompare,.estimate_cost,.Promptfp_comparefixture (function-scoped): wrapscompare(), returnsPromptTestResult@pytest.mark.prompt_testand@pytest.mark.slow_promptmarkers--no-slow-promptsCLI option to skip expensive testsPromptTestResultassertion wrapper:.assert_significant(threshold=0.05)-- fail unless statistically significant.assert_winner(expected)-- fail unless the named variant won.assert_no_errors()-- fail if any variant recorded errors.is_significant,.winner,.p_valueconvenience properties- New
pytestoptional dependency:pip install flowprompt-ai[pytest] - New example:
examples/11_pytest_testing.py
0.3.0 - 2026-02-06¶
Added¶
compare()convenience function: One-call prompt A/B testing with statistical significancecompare(prompts, inputs, model)-- compare prompt variants in a single callacompare()async variant with parallel execution viaasyncio.gatherComparisonResultandVariantResultdataclasses with__str__()andto_dict()- Top-level exports:
from flowprompt import compare, acompare, ComparisonResult - Prompt Lab example (
examples/10_prompt_lab.py): End-to-end sentiment analysis comparison demo
Fixed¶
- Native JSON schema for structured outputs: Use
litellm.supports_response_schema()to detect models with nativejson_schemaresponse format instead of always appending schema text to the system message. Falls back gracefully for older models. - Dynamic model pricing:
UsageInfo.calculate_cost()now looks uplitellm.model_costfor current per-token pricing instead of using a hardcoded 9-model table. Built-in fallback prices remain for when litellm is unavailable.
Changed¶
- README rewritten to lead with
compare()as the hero feature, honest comparison table - Quickstart docs updated with "Compare Prompts" section as second thing users learn
MODEL_PRICINGrenamed to_FALLBACK_PRICING(private, backward-compatible)
0.2.1 - 2026-01-16¶
Added¶
- CLI
optimizecommand: One-command prompt optimization from terminal flowprompt optimize my_prompt.py examples.json --strategy fewshot- Shows before/after accuracy comparison
- Supports fewshot, instruction, and bootstrap strategies
- Added
a-b-testing,prompt-testing,langchain-alternativekeywords for discoverability
Fixed¶
- Fixed 404 documentation URL (now points to GitHub docs)
- Fixed package name in installation docs (
flowprompt→flowprompt-ai) - Fixed silent exception swallowing in optimizer (now logs warnings)
- Added upper bounds to dependencies for security (litellm, jinja2, pyyaml)
- Removed
devfrom[all]extra (don't ship linters to users) - Made basic example actually runnable with API key detection
Changed¶
- README repositioned to lead with unique value proposition: "Stop guessing which prompt works. Measure it."
- A/B Testing moved to top of feature comparison tables
0.2.0 - 2026-01-10¶
Added¶
Prompt Optimization (DSPy-style)¶
- FewShotOptimizer: Automatic few-shot example selection and optimization
- InstructionOptimizer: LLM-powered instruction refinement
- OptunaOptimizer: Hyperparameter search with Optuna integration
- BootstrapOptimizer: Self-improving bootstrapping optimization
- ExampleDataset: Structured dataset management for optimization
- Metrics: ExactMatch, ContainsMatch, F1Score, StructuredAccuracy, RegexMatch, CustomMetric
- Convenience function: Simple
optimize()function for quick optimization
A/B Testing Framework¶
- ExperimentConfig: Define A/B experiments with multiple variants
- Allocation strategies: Random, RoundRobin, Weighted, EpsilonGreedy, UCB1, ThompsonSampling
- Statistical tests: Z-test, Chi-squared, T-test, Bayesian A/B testing
- ExperimentStore: Persist and analyze experiment results
- Multi-armed bandit support: Adaptive allocation based on performance
Multimodal Support¶
- ImageContent: Support for base64, URL, and file-based images
- VideoContent: Video frame extraction and processing
- DocumentContent: PDF, DOCX, HTML, and plain text processing
- MultimodalPrompt: Unified prompt class for multimodal inputs
Fixed¶
- Resolved ruff linting errors across the codebase
- Added proper exception chaining with
from err - Fixed unused variable warnings
- Added
strict=Trueto zip() calls for safety - Improved code quality with set comprehensions
0.1.0 - 2026-01-10¶
Added¶
Core Features¶
- Prompt class with Pydantic v2 integration: Type-safe prompt definitions with full IDE autocomplete and validation
- Structured output validation: Automatic parsing and validation of LLM responses via Pydantic models
- Multi-provider support via LiteLLM: Works with OpenAI, Anthropic, Google, Ollama, and 100+ providers
- Async support: Full async/await support with
arun()method
Streaming¶
- Sync streaming: Real-time streaming responses with
stream()method - Async streaming: Async streaming support with
astream()method - StreamChunk model: Structured streaming chunks with delta content and metadata
Caching¶
- Prompt caching system: Built-in caching for 50-90% cost reduction
- Memory cache backend: Fast in-memory caching with TTL support
- File cache backend: Persistent file-based caching for cross-session reuse
- Cache statistics: Track hits, misses, and hit rates
- Configurable TTL: Time-to-live settings for cache entries
Observability¶
- OpenTelemetry tracing: Full distributed tracing support
- Cost tracking: Automatic cost calculation per request
- Token tracking: Input/output token counting
- Latency metrics: Request timing and performance monitoring
- Usage summaries: Aggregate statistics across all requests
Prompt Loading¶
- YAML prompt files: Load prompts from YAML files for team collaboration
- JSON prompt files: Alternative JSON format support
- Prompt registry: Load and manage multiple prompts from directories
- Schema validation: JSON Schema support for output validation in files
CLI Tools¶
flowprompt init: Initialize new FlowPrompt projects with best-practice structureflowprompt test: Test all prompts in a directoryflowprompt run: Run individual prompts from command lineflowprompt stats: View usage and cost statisticsflowprompt cache-stats: View cache performance metricsflowprompt cache-clear: Clear the prompt cacheflowprompt list-prompts: List all available prompts
Template System¶
- Python format strings: Simple
{variable}interpolation - Jinja2 templates: Full Jinja2 support for complex logic (conditionals, loops)
- Mixed template support: Use both formats as needed
Versioning¶
- Prompt versioning: Explicit version tracking with
__version__attribute - Content hashing: Automatic content hash generation for change detection
- Version comparison: Track prompt evolution over time
Technical¶
- Python 3.10+ support
- Pydantic v2 integration
- MIT License
- Modern tooling: ruff, mypy, pytest, pre-commit
- GitHub Actions CI/CD ready
- Comprehensive test suite
- Full type annotations