Skip to content

Quick Start

This guide will get you up and running with FlowPrompt in under 5 minutes.

Your First Prompt

Create a simple prompt that extracts information from text:

from flowprompt import Prompt
from pydantic import BaseModel


class ExtractUser(Prompt):
    """Extract user information from text."""

    system = "You are a precise data extractor."
    user = "Extract the name and age from: {text}"

    class Output(BaseModel):
        name: str
        age: int


# Run the prompt
result = ExtractUser(text="John Smith is 25 years old").run(model="gpt-4o")
print(f"Name: {result.name}")  # Name: John Smith
print(f"Age: {result.age}")  # Age: 25

Compare Prompts

The fastest way to find which prompt works better -- with statistical significance:

from flowprompt import Prompt, compare


class PromptV1(Prompt):
    system = "You are a precise data extractor."
    user = "Extract the name and age from: {text}"


class PromptV2(Prompt):
    system = "Extract structured data. Be accurate and concise."
    user = "From the following text, extract name and age: {text}"


result = compare(
    {"v1": PromptV1, "v2": PromptV2},
    inputs=[
        {"text": "John Smith is 25 years old"},
        {"text": "Alice (age 30) joined today"},
        {"text": "Bob, 42, from NYC"},
    ],
    expected=["John Smith, 25", "Alice, 30", "Bob, 42"],  # ground-truth answers
    model="gpt-4o-mini",
)
print(result)

print(result) shows each variant's accuracy with a confidence interval, latency, cost and cost per correct answer, the paired test that compared them, and a one-line verdict. With three inputs the verdict will be "Not enough data": no difference can be significant with so few examples. Use an evaluation set of dozens to hundreds of inputs; plan_sample_size() tells you how many (see the A/B testing guide).

Use dry_run=True to preview estimated cost before spending API credits:

result = compare(
    {"v1": PromptV1, "v2": PromptV2},
    inputs=test_data,
    model="gpt-4o-mini",
    dry_run=True,  # No API calls -- just estimates cost
)
print(result)  # Shows estimated cost and number of API calls

For async execution with parallel variant runs:

from flowprompt import acompare

result = await acompare(
    {"v1": PromptV1, "v2": PromptV2},
    inputs=test_data,
    model="gpt-4o-mini",
)

Tip: FlowPrompt includes a pytest plugin for running prompt tests in CI. Install flowprompt-ai[pytest] and use the fp_compare fixture. See the A/B Testing Guide for details.

Streaming Responses

For real-time output, use streaming:

for chunk in ExtractUser(text="John is 25").stream(model="gpt-4o"):
    print(chunk.delta, end="", flush=True)

Async streaming:

async for chunk in ExtractUser(text="John is 25").astream(model="gpt-4o"):
    print(chunk.delta, end="", flush=True)

Using Different Providers

Switch providers by changing the model string:

# OpenAI
result = prompt.run(model="gpt-4o")

# Anthropic Claude
result = prompt.run(model="anthropic/claude-3-5-sonnet-20241022")

# Google Gemini
result = prompt.run(model="gemini/gemini-2.0-flash-exp")

# Local Ollama
result = prompt.run(model="ollama/llama3")

Enable Caching

Reduce costs by caching identical requests:

from flowprompt import configure_cache

# Enable caching with 1-hour TTL
configure_cache(enabled=True, default_ttl=3600)

# First call hits the API
result1 = ExtractUser(text="John is 25").run(model="gpt-4o")

# Second call returns cached result
result2 = ExtractUser(text="John is 25").run(model="gpt-4o")  # Instant!

Track Usage and Costs

Measure the tokens and cost of a block of calls:

from flowprompt import track_usage

with track_usage() as calls:
    result = ExtractUser(text="John is 25").run(model="gpt-4o")

print(sum(c.cost_usd or 0 for c in calls), sum(c.total_tokens for c in calls))

Or trace every call in your application:

from flowprompt import configure_tracer

tracer = configure_tracer(service_name="my-app")  # before running prompts
result = ExtractUser(text="John is 25").run(model="gpt-4o")

summary = tracer.get_summary()
print(f"Total cost: ${summary['total_cost_usd']:.4f}")
print(f"Total tokens: {summary['total_tokens']}")

See Caching, tracing and cost.

Load Prompts from Files

Create prompts/extract_user.yaml:

name: ExtractUser
version: "1.0.0"
system: You are a precise data extractor.
user: "Extract from: {{ text }}"
output_schema:
  type: object
  properties:
    name:
      type: string
    age:
      type: integer
  required: [name, age]

Load and use it:

from flowprompt import load_prompt

ExtractUser = load_prompt("prompts/extract_user.yaml")
result = ExtractUser(text="John is 25").run(model="gpt-4o")

Using the CLI

Initialize a new project:

flowprompt init my-project
cd my-project

Run a prompt:

flowprompt run prompts/extract_user.yaml --var text="John is 25"

Compare prompt variants on a JSONL dataset (and fail CI on a regression):

flowprompt compare variants.py dataset.jsonl --model gpt-4o-mini --metric exact --fail-on-regression

See Prompt tests in CI.

Next Steps