Toolscore Documentation

PyPI version License Downloads

Toolscore is a Python package for evaluating LLM tool usage against gold standard specifications. It helps developers benchmark different models, validate agent behavior, and track improvements in function calling accuracy over time.

What is Toolscore?

Toolscore evaluates LLM tool usage - it doesn’t call LLM APIs directly. Think of it as a testing framework for function-calling agents:

✅ Evaluates tool usage traces from OpenAI, Anthropic, Gemini, agent frameworks, recorded MCP sessions, OpenTelemetry spans, or custom sources

✅ Compares actual behavior against expected gold standards

✅ Reports detailed metrics on accuracy, efficiency, and correctness, plus failed calls, leaked credentials and forbidden calls

✅ Grades and lints MCP servers, including references to missing tools and tool poisoning

❌ Does NOT call LLM APIs or execute tools (you capture traces separately)

Quick Start

pip install tool-scorer

# Run evaluation
toolscore eval examples/gold_calls.json examples/trace_openai.json --html report.html

# Grade an MCP server
toolscore mcp test "python my_server.py"

Key Features

  • Comprehensive Metrics Suite: Tool invocation accuracy, selection accuracy, sequence edit distance, argument matching, redundant and repeated calls, required calls completed, and side-effect validation

  • Behavior and Safety Checks: Failed calls, blind retries, credentials passed into tool arguments, and forbidden-call policies, with a CI gate (Behavior and Safety Checks)

  • Snapshot Testing: Record, approve and replay your agent’s tool calls in pytest (Snapshot Testing)

  • MCP Scorecard, Lint and Recorder: Grade any MCP server, lint it for missing-tool references and tool poisoning, and record real sessions (Testing MCP Servers)

  • Native Everywhere: OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, CrewAI, MCP and OpenTelemetry traces (Framework Integration Guide)

  • CLI, Python API and GitHub Action: Command-line interface, programmatic usage and CI gates

  • Rich Reports: Console, HTML, Markdown, CSV and machine-readable JSON reports, plus a versioned record for evaluation harnesses

  • Extensible: Easy to add custom metrics and validators

Contents

Guides

Development

Indices and tables