API Reference

This section contains the complete API documentation for Toolscore.

Overview

Toolscore provides a simple, Pythonic API for evaluating LLM tool usage. The main entry point is the evaluate_trace() function, which handles loading traces, computing metrics, and returning results.

Basic Usage

from toolscore import evaluate_trace

result = evaluate_trace(
    gold_file="gold_calls.json",
    trace_file="trace.json",
    format="auto"
)

print(f"Accuracy: {result.metrics['selection_accuracy']:.2%}")

Main Components

  • Core Module - Core evaluation logic

  • Adapters Module - Trace format adapters (OpenAI, Anthropic, Gemini, MCP, OpenTelemetry, LangChain, custom)

  • Metrics Module - Metric calculators (accuracy, sequence, arguments, required calls, errors, credentials, forbidden calls)

  • Validators Module - Side-effect validators (HTTP, filesystem, database)

  • MCP Module - MCP client, session recorder, lint rules and scorecard

  • Reports Module - Report generators (JSON, HTML, Markdown, CSV)

Quick Reference

Core Functions

evaluate_trace

Evaluate an agent's trace against gold standard.

load_gold_standard

Load gold standard specification from JSON file.

load_trace

Load agent trace from JSON file.

Adapters

OpenAIAdapter

Adapter for OpenAI function call traces.

AnthropicAdapter

Adapter for Anthropic Claude tool-use traces.

GeminiAdapter

Adapter for Google Gemini function call traces.

MCPAdapter

Adapter for Anthropic Model Context Protocol (MCP) traces.

OTelAdapter

Adapter for OpenTelemetry GenAI tool spans (OTLP JSON exports).

CustomAdapter

Adapter for custom/generic JSON trace formats.

ToolCall

Represents a single tool call in a trace.

Metrics

calculate_invocation_accuracy

Calculate tool invocation accuracy.

calculate_selection_accuracy

Calculate tool selection accuracy.

calculate_edit_distance

Calculate edit distance metrics for tool call sequences.

calculate_argument_f1

Calculate F1 score for argument matching.

calculate_redundant_call_rate

Calculate redundant call rate.

calculate_required_call_recall

Share of required calls that the trace completed, counting every required call.

find_secrets

Find credentials in tool arguments.

check_forbidden_calls

Report every call that matches a forbidden rule.

calculate_side_effect_success_rate

Calculate side-effect success rate.

calculate_latency

Calculate latency metrics from tool calls.

calculate_cost_attribution

Calculate cost attribution from tool calls.

Validators

HTTPValidator

Validator for HTTP-related side effects.

FileSystemValidator

Validator for filesystem-related side effects.

SQLValidator

Validator for SQL/database-related side effects.

Reports

generate_json_report

Generate JSON report from evaluation result.

generate_html_report

Generate HTML report from evaluation result.

generate_markdown_report

Generate Markdown report from evaluation result.