Quick Start¶
This guide will get you up and running with FlowPrompt in under 5 minutes.
Your First Prompt¶
Create a simple prompt that extracts information from text:
from flowprompt import Prompt
from pydantic import BaseModel
class ExtractUser(Prompt):
"""Extract user information from text."""
system = "You are a precise data extractor."
user = "Extract the name and age from: {text}"
class Output(BaseModel):
name: str
age: int
# Run the prompt
result = ExtractUser(text="John Smith is 25 years old").run(model="gpt-4o")
print(f"Name: {result.name}") # Name: John Smith
print(f"Age: {result.age}") # Age: 25
Compare Prompts¶
The fastest way to find which prompt works better -- with statistical significance:
from flowprompt import Prompt, compare
class PromptV1(Prompt):
system = "You are a precise data extractor."
user = "Extract the name and age from: {text}"
class PromptV2(Prompt):
system = "Extract structured data. Be accurate and concise."
user = "From the following text, extract name and age: {text}"
result = compare(
{"v1": PromptV1, "v2": PromptV2},
inputs=[
{"text": "John Smith is 25 years old"},
{"text": "Alice (age 30) joined today"},
{"text": "Bob, 42, from NYC"},
],
expected=["John Smith, 25", "Alice, 30", "Bob, 42"], # ground-truth answers
model="gpt-4o-mini",
)
print(result)
print(result) shows each variant's accuracy with a confidence interval,
latency, cost and cost per correct answer, the paired test that compared
them, and a one-line verdict. With three inputs the verdict will be "Not
enough data": no difference can be significant with so few examples. Use an
evaluation set of dozens to hundreds of inputs; plan_sample_size() tells
you how many (see the A/B testing guide).
Use dry_run=True to preview estimated cost before spending API credits:
result = compare(
{"v1": PromptV1, "v2": PromptV2},
inputs=test_data,
model="gpt-4o-mini",
dry_run=True, # No API calls -- just estimates cost
)
print(result) # Shows estimated cost and number of API calls
For async execution with parallel variant runs:
from flowprompt import acompare
result = await acompare(
{"v1": PromptV1, "v2": PromptV2},
inputs=test_data,
model="gpt-4o-mini",
)
Tip: FlowPrompt includes a pytest plugin for running prompt tests in CI. Install
flowprompt-ai[pytest]and use thefp_comparefixture. See the A/B Testing Guide for details.
Streaming Responses¶
For real-time output, use streaming:
for chunk in ExtractUser(text="John is 25").stream(model="gpt-4o"):
print(chunk.delta, end="", flush=True)
Async streaming:
async for chunk in ExtractUser(text="John is 25").astream(model="gpt-4o"):
print(chunk.delta, end="", flush=True)
Using Different Providers¶
Switch providers by changing the model string:
# OpenAI
result = prompt.run(model="gpt-4o")
# Anthropic Claude
result = prompt.run(model="anthropic/claude-3-5-sonnet-20241022")
# Google Gemini
result = prompt.run(model="gemini/gemini-2.0-flash-exp")
# Local Ollama
result = prompt.run(model="ollama/llama3")
Enable Caching¶
Reduce costs by caching identical requests:
from flowprompt import configure_cache
# Enable caching with 1-hour TTL
configure_cache(enabled=True, default_ttl=3600)
# First call hits the API
result1 = ExtractUser(text="John is 25").run(model="gpt-4o")
# Second call returns cached result
result2 = ExtractUser(text="John is 25").run(model="gpt-4o") # Instant!
Track Usage and Costs¶
Measure the tokens and cost of a block of calls:
from flowprompt import track_usage
with track_usage() as calls:
result = ExtractUser(text="John is 25").run(model="gpt-4o")
print(sum(c.cost_usd or 0 for c in calls), sum(c.total_tokens for c in calls))
Or trace every call in your application:
from flowprompt import configure_tracer
tracer = configure_tracer(service_name="my-app") # before running prompts
result = ExtractUser(text="John is 25").run(model="gpt-4o")
summary = tracer.get_summary()
print(f"Total cost: ${summary['total_cost_usd']:.4f}")
print(f"Total tokens: {summary['total_tokens']}")
See Caching, tracing and cost.
Load Prompts from Files¶
Create prompts/extract_user.yaml:
name: ExtractUser
version: "1.0.0"
system: You are a precise data extractor.
user: "Extract from: {{ text }}"
output_schema:
type: object
properties:
name:
type: string
age:
type: integer
required: [name, age]
Load and use it:
from flowprompt import load_prompt
ExtractUser = load_prompt("prompts/extract_user.yaml")
result = ExtractUser(text="John is 25").run(model="gpt-4o")
Using the CLI¶
Initialize a new project:
Run a prompt:
Compare prompt variants on a JSONL dataset (and fail CI on a regression):
flowprompt compare variants.py dataset.jsonl --model gpt-4o-mini --metric exact --fail-on-regression
See Prompt tests in CI.
Next Steps¶
- Read the A/B testing guide and Statistical methods
- Read the API Reference for detailed documentation
- Check out the examples directory
- Join our GitHub Discussions