Toolscore Documentation
Toolscore is a Python package for evaluating LLM tool usage against gold standard specifications. It helps developers benchmark different models, validate agent behavior, and track improvements in function calling accuracy over time.
What is Toolscore?
Toolscore evaluates LLM tool usage - it doesn’t call LLM APIs directly. Think of it as a testing framework for function-calling agents:
✅ Evaluates tool usage traces from OpenAI, Anthropic, Gemini, agent frameworks, recorded MCP sessions, OpenTelemetry spans, or custom sources
✅ Compares actual behavior against expected gold standards
✅ Reports detailed metrics on accuracy, efficiency, and correctness, plus failed calls, leaked credentials and forbidden calls
✅ Grades and lints MCP servers, including references to missing tools and tool poisoning
❌ Does NOT call LLM APIs or execute tools (you capture traces separately)
Quick Start
pip install tool-scorer
# Run evaluation
toolscore eval examples/gold_calls.json examples/trace_openai.json --html report.html
# Grade an MCP server
toolscore mcp test "python my_server.py"
Key Features
Comprehensive Metrics Suite: Tool invocation accuracy, selection accuracy, sequence edit distance, argument matching, redundant and repeated calls, required calls completed, and side-effect validation
Behavior and Safety Checks: Failed calls, blind retries, credentials passed into tool arguments, and forbidden-call policies, with a CI gate (Behavior and Safety Checks)
Snapshot Testing: Record, approve and replay your agent’s tool calls in pytest (Snapshot Testing)
MCP Scorecard, Lint and Recorder: Grade any MCP server, lint it for missing-tool references and tool poisoning, and record real sessions (Testing MCP Servers)
Native Everywhere: OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, CrewAI, MCP and OpenTelemetry traces (Framework Integration Guide)
CLI, Python API and GitHub Action: Command-line interface, programmatic usage and CI gates
Rich Reports: Console, HTML, Markdown, CSV and machine-readable JSON reports, plus a versioned record for evaluation harnesses
Extensible: Easy to add custom metrics and validators
Contents
Getting Started
Guides
- User Guide
- Understanding Metrics
- Scoring Semantics
- Working with Trace Formats
- Capturing Traces
- Creating Effective Gold Standards
- Side-Effect Validation
- Generating Reports
- Batch Evaluation
- End-to-End Agent Testing
- Data-Driven Testing with
@toolscore.cases() - Pytest Integration
- Interactive Tutorials
- Tips and Tricks
- Troubleshooting
- Snapshot Testing
- Argument Matchers
- Framework Integration Guide
- Fluent API —
expect() - Behavior and Safety Checks
- LLM-as-a-Judge
- Testing MCP Servers
- Extending Toolscore
- Toolscore vs. Other Tools
API Reference
Development