Quick Start
This guide gets you to a green tool-call test in about 60 seconds, then covers the rest of the API.
Install Toolscore
pip install tool-scorer
The fastest path: toolscore init + snapshots
toolscore init detects your agent framework and scaffolds a working pytest
suite plus an optional CI workflow — so your first run records a snapshot and
passes immediately.
1. Scaffold a test suite
toolscore init # auto-detects your framework, prompts to confirm
toolscore init --yes # accept the detected framework non-interactively
toolscore init --framework langgraph # or pick one explicitly
This writes tests/test_agent_tools.py (a passing suite wired to a snapshot),
.toolscore/snapshots/ (your reviewed baselines live here), and
.github/workflows/toolscore.yml (unless --no-ci).
2. Run pytest — the first run records a snapshot
Edit the # TODO: import your agent block to call your real agent, then:
pytest
The toolscore_snapshot fixture records the agent’s tool calls into a
pending snapshot and warns. Nothing is asserted yet — there is no baseline to
compare against.
3. Review and approve the snapshot
toolscore snapshots list # see what was recorded
toolscore snapshots show <name> # inspect the tool calls
toolscore approve --all # lock them in as the baseline
From now on every run replays against the approved baseline and fails on
drift. When the behavior changes on purpose, re-record with
pytest --toolscore-update (or toolscore record --update -- pytest) and
commit the updated snapshot file. See Snapshot Testing for the full story.
The one-line assertion
If you would rather assert against an explicit spec than a recorded snapshot:
import toolscore
toolscore.assert_tools(
expected=[{"tool": "get_weather", "args": {"city": "NYC"}}],
actual=my_agent("weather in NYC"), # raw provider response — auto-detected
min_score=0.9,
)
Evaluating captured traces
Toolscore can also score pre-captured trace files from the command line:
tool-scorer eval examples/gold_calls.json examples/trace_openai.json --html report.html
Console output shows the metrics; open report.html for a detailed breakdown.
Basic Usage
Command Line Interface
Evaluate a trace:
tool-scorer eval gold_calls.json trace.json
Generate both JSON and HTML reports:
tool-scorer eval gold_calls.json trace.json --html report.html
Specify trace format explicitly:
tool-scorer eval gold_calls.json trace.json --format openai
Validate trace file format:
tool-scorer validate trace.json
Python API (In-Memory)
from toolscore import evaluate
result = evaluate(
expected=[{"tool": "get_weather", "args": {"city": "NYC"}}],
actual=[{"tool": "get_weather", "args": {"city": "NYC"}}],
)
print(result.score) # 1.0
Pass raw LLM provider responses directly — auto-detected:
from openai import OpenAI
from toolscore import evaluate
client = OpenAI()
response = client.chat.completions.create(model="gpt-4o", messages=[...], tools=[...])
# No from_openai() needed — auto-detected!
result = evaluate(expected=[...], actual=response)
Python API (File-Based)
from toolscore import evaluate_trace
result = evaluate_trace(
gold_file="gold_calls.json",
trace_file="trace.json",
format="auto"
)
print(f"Selection Accuracy: {result.metrics['selection_accuracy']:.2%}")
Creating Gold Standards
A gold standard defines the expected tool calls for a task. Create a gold_calls.json file:
Note
Omitting ``args`` means “do not check arguments” — the tool must be
called, but any arguments are accepted. An explicit "args": {} means
“expect the tool to be called with no arguments.” See
Argument Matchers for the full contract and for matchers like ANY and
Regex that assert on argument shape instead of exact values.
[
{
"tool": "make_file",
"args": {
"filename": "poem.txt",
"lines_of_text": ["Roses are red,", "Violets are blue."]
},
"side_effects": {
"file_exists": "poem.txt"
},
"description": "Create a file with a poem"
}
]
Supported Trace Formats
OpenAI Format
[
{
"role": "assistant",
"function_call": {
"name": "get_weather",
"arguments": "{\"location\": \"Boston\"}"
}
}
]
Anthropic Format
[
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_123",
"name": "search",
"input": {"query": "Python"}
}
]
}
]
Custom Format
{
"calls": [
{
"tool": "read_file",
"args": {"path": "data.txt"},
"result": "file contents"
}
]
}
Other Trace Sources
Sessions recorded with toolscore mcp record and OpenTelemetry exports
({"resourceSpans": [...]}) are detected automatically, so the same
toolscore eval gold.json trace.json command scores them. See
Framework Integration Guide.
Check behavior and safety
Every evaluation also reports failed calls, blind retries, the required calls completed and credentials passed into tool arguments. Add calls the agent must never make, and fail the build when one happens:
toolscore eval gold.json trace.json --forbidden forbidden.json --fail-on-violations
Next Steps
Read the User Guide for detailed usage
Lock in behavior with Snapshot Testing
Assert on argument shape with Argument Matchers
Catch failed calls, leaked credentials and forbidden calls with Behavior and Safety Checks
Pass raw framework responses with the Framework Integration Guide extractors
Test an MCP server with Testing MCP Servers – or run
toolscore demoto grade a bundled sample server in seconds (no setup, no API key)Add semantic scoring with the LLM-as-a-Judge
Explore example scripts in the examples/ directory
Check out the complete API Reference
Learn how to Contributing to Toolscore