API Reference
This section contains the complete API documentation for Toolscore.
Overview
Toolscore provides a simple, Pythonic API for evaluating LLM tool usage. The main entry point is the evaluate_trace() function, which handles loading traces, computing metrics, and returning results.
Basic Usage
from toolscore import evaluate_trace
result = evaluate_trace(
gold_file="gold_calls.json",
trace_file="trace.json",
format="auto"
)
print(f"Accuracy: {result.metrics['selection_accuracy']:.2%}")
Main Components
Core Module - Core evaluation logic
Adapters Module - Trace format adapters (OpenAI, Anthropic, Gemini, MCP, OpenTelemetry, LangChain, custom)
Metrics Module - Metric calculators (accuracy, sequence, arguments, required calls, errors, credentials, forbidden calls)
Validators Module - Side-effect validators (HTTP, filesystem, database)
MCP Module - MCP client, session recorder, lint rules and scorecard
Reports Module - Report generators (JSON, HTML, Markdown, CSV)
Quick Reference
Core Functions
|
Evaluate an agent's trace against gold standard. |
Load gold standard specification from JSON file. |
|
Load agent trace from JSON file. |
Adapters
Adapter for OpenAI function call traces. |
|
Adapter for Anthropic Claude tool-use traces. |
|
Adapter for Google Gemini function call traces. |
|
Adapter for Anthropic Model Context Protocol (MCP) traces. |
|
Adapter for OpenTelemetry GenAI tool spans (OTLP JSON exports). |
|
Adapter for custom/generic JSON trace formats. |
|
Represents a single tool call in a trace. |
Metrics
Calculate tool invocation accuracy. |
|
Calculate tool selection accuracy. |
|
Calculate edit distance metrics for tool call sequences. |
|
Calculate F1 score for argument matching. |
|
Calculate redundant call rate. |
|
Share of required calls that the trace completed, counting every required call. |
|
Find credentials in tool arguments. |
|
Report every call that matches a forbidden rule. |
|
Calculate side-effect success rate. |
|
Calculate latency metrics from tool calls. |
|
Calculate cost attribution from tool calls. |
Validators
Validator for HTTP-related side effects. |
|
Validator for filesystem-related side effects. |
|
Validator for SQL/database-related side effects. |
Reports
Generate JSON report from evaluation result. |
|
Generate HTML report from evaluation result. |
|
Generate Markdown report from evaluation result. |