Skip to content

FlowPrompt

Stop guessing which prompt works. Measure it.

FlowPrompt is a Python library for prompts as typed classes and for comparing prompts, models and pipelines with statistics you can trust.

  • Typed prompts. Prompts are Pydantic classes with templates and a structured Output model, run against 100+ providers through LiteLLM.
  • Paired A/B tests. compare() runs every variant on the same inputs and tests the difference with the exact McNemar test (pass/fail) or a paired permutation test (scores), corrected for multiple comparisons.
  • Answers, not just p-values. Each result reports accuracy with confidence intervals, latency, cost and cost per correct answer, a plain-English verdict, and how many inputs you would need to detect a smaller difference.
  • Built for CI. Markdown and HTML reports, a flowprompt compare command that fails the build on a significant regression, a pytest plugin, and FakeLLM for offline tests.

A first comparison

from flowprompt import Prompt, compare


class Concise(Prompt):
    system = (
        "Classify the sentiment. Reply with one word: positive, negative or neutral."
    )
    user = "Review: {text}"


class Chatty(Prompt):
    system = "You are a helpful assistant. Analyze the sentiment of the review."
    user = "Review: {text}"


result = compare(
    {"concise": Concise, "chatty": Chatty},
    inputs=[{"text": text} for text, _ in reviews],
    expected=[label for _, label in reviews],
    eval_metric="exact",
    model="gpt-4o-mini",
)
print(result.verdict)
result.save_report("report.html")

Comparison report

The complete script, which runs offline with a simulated model, is examples/12_ab_test_report.py.

Where to go next

Getting help

FlowPrompt is released under the MIT License.