Core Module

The core module provides the main evaluation functions and result container.

Main Functions

toolscore.core.evaluate(expected, actual, weights=None, strict=False, forbidden=None)[source]

Evaluate tool calls by comparing actual against expected (in-memory).

This is the simplest way to use Toolscore - pass Python dicts directly, no file I/O required. Raw OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, and Claude Agent SDK responses are auto-detected and converted automatically, including list-shaped formats (Claude Agent SDK message lists and bare LangGraph message lists).

Parameters:
  • expected (list[dict[str, Any]]) – List of expected tool calls, each a dict with ‘tool’ and optional ‘args’.

  • actual (list[dict[str, Any]] | Any) – List of actual tool calls from your agent, same format. Also accepts raw OpenAI/Anthropic/Gemini response objects or dicts (auto-detected).

  • weights (dict[str, float] | None) – Optional custom weights for the composite score. Keys: ‘selection_accuracy’, ‘argument_f1’, ‘sequence_accuracy’, ‘redundant_rate’, ‘required_call_recall’ (0 by default; give it a weight to penalize required calls that were skipped or failed). Provided values are merged with the defaults then renormalized so that all weights sum to 1.0 before computing the composite score.

  • strict (bool) – When True, argument comparison uses pure equality (no int/float coercion, no string strip). Default is False (lenient matching).

  • forbidden (list[dict[str, Any]] | None) – Optional calls the agent must never make, as rule dicts with tool, optional args (values or matchers) and optional reason. Violations are reported in metrics["policy_metrics"] and EvaluationResult.policy_violations; they do not change the score. See toolscore.metrics.policy.

Return type:

EvaluationResult

Returns:

EvaluationResult with metrics and a composite .score property.

Example

>>> from toolscore import evaluate
>>> result = evaluate(
...     expected=[{"tool": "search", "args": {"q": "test"}}],
...     actual=[{"tool": "search", "args": {"q": "test"}}],
... )
>>> result.score
1.0
toolscore.core.assert_tools(expected, actual, min_score=0.9, weights=None, strict=False)[source]

Assert that actual tool calls meet a minimum composite score.

Convenience function for use in pytest or any test framework.

Parameters:
  • expected (list[dict[str, Any]]) – List of expected tool calls.

  • actual (list[dict[str, Any]] | Any) – List of actual tool calls.

  • min_score (float) – Minimum composite score required (0.0 to 1.0).

  • weights (dict[str, float] | None) – Optional custom weights for the composite score.

  • strict (bool) – When True, argument comparison uses pure equality (no int/float coercion, no string strip). Default is False.

Return type:

EvaluationResult

Returns:

EvaluationResult if assertion passes.

Raises:
  • ValueError – If min_score is outside [0.0, 1.0].

  • ToolScoreAssertionError – If the composite score is below min_score.

toolscore.core.test_agent(agent, input, expected, min_score=None, weights=None, strict=False)[source]

End-to-end test helper: run an agent, extract tool calls, evaluate.

Calls agent(input), passes the response through auto-detection to extract tool calls, and evaluates against expected.

Parameters:
  • agent (Callable[..., Any]) – Any callable that accepts a string and returns an LLM response (raw provider response or list of tool-call dicts). Must be a synchronous callable; use test_agent_async() for async agents.

  • input (str) – The prompt / user message to send to the agent.

  • expected (list[dict[str, Any]]) – List of expected tool calls.

  • min_score (float | None) – If provided, raises ToolScoreAssertionError when the composite score is below this threshold.

  • weights (dict[str, float] | None) – Optional custom weights for the composite score.

  • strict (bool) – When True, argument comparison uses pure equality.

Return type:

EvaluationResult

Returns:

EvaluationResult with metrics and a composite .score property.

Raises:
  • TypeError – If agent is an async function or returns an awaitable.

  • ValueError – If min_score is outside [0.0, 1.0].

  • ToolScoreAssertionError – If min_score is set and the score is below it.

async toolscore.core.test_agent_async(agent, input, expected, min_score=None, weights=None, strict=False)[source]

Async end-to-end test helper: run an agent, extract tool calls, evaluate.

Works for both synchronous and asynchronous agents. Calls agent(input) and, if the result is awaitable, awaits it.

Parameters:
  • agent (Callable[..., Any]) – Any callable that accepts a string and returns an LLM response (raw provider response or list of tool-call dicts). May be sync or async.

  • input (str) – The prompt / user message to send to the agent.

  • expected (list[dict[str, Any]]) – List of expected tool calls.

  • min_score (float | None) – If provided, raises ToolScoreAssertionError when the composite score is below this threshold.

  • weights (dict[str, float] | None) – Optional custom weights for the composite score.

  • strict (bool) – When True, argument comparison uses pure equality.

Return type:

EvaluationResult

Returns:

EvaluationResult with metrics and a composite .score property.

Raises:
  • ValueError – If min_score is outside [0.0, 1.0].

  • ToolScoreAssertionError – If min_score is set and the score is below it.

toolscore.core.evaluate_trace(gold_file, trace_file, format='auto', validate_side_effects=True, judge=False, forbidden=None, weights=None)[source]

Evaluate an agent’s trace against gold standard.

Parameters:
  • gold_file (str | Path) – Path to gold standard specification.

  • trace_file (str | Path) – Path to agent trace.

  • format (str) – Trace format (‘auto’, ‘openai’, ‘anthropic’, ‘gemini’, ‘mcp’, ‘langchain’, ‘otel’, ‘custom’).

  • validate_side_effects (bool) – Whether to validate side effects.

  • judge (JudgeConfig | str | bool) – LLM-as-a-judge configuration for semantic evaluation. False (default) disables it. True uses a default JudgeConfig(). A string is treated as a model-name shorthand. A JudgeConfig is used as given. Provider is inferred from the model name (or set explicitly via JudgeConfig): claude-* -> Anthropic, gemini-* -> Gemini, a base_url -> any OpenAI-compatible endpoint (Ollama/vLLM/Groq), otherwise OpenAI.

  • forbidden (list[dict[str, Any]] | None) – Optional calls the agent must never make (see evaluate()); violations go to metrics["policy_metrics"].

  • weights (dict[str, float] | None) – Optional custom weights for the composite score (see evaluate()).

Return type:

EvaluationResult

Returns:

EvaluationResult containing all computed metrics.

Raises:
toolscore.core.load_gold_standard(file_path)[source]

Load gold standard specification from JSON file.

Parameters:

file_path (str | Path) – Path to gold_calls.json file.

Return type:

list[ToolCall]

Returns:

List of expected tool calls.

Raises:
toolscore.core.load_trace(file_path, format='auto')[source]

Load agent trace from JSON file.

Parameters:
  • file_path (str | Path) – Path to trace file.

  • format (str) – Trace format (‘auto’, ‘openai’, ‘anthropic’, ‘gemini’, ‘mcp’, ‘langchain’, ‘otel’, ‘custom’).

Return type:

list[ToolCall]

Returns:

List of tool calls from the trace.

Raises:
toolscore.integrations.auto_extract(actual)[source]

Auto-detect the provider format of a response and extract tool calls.

This allows passing raw OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, or Claude Agent SDK responses directly to evaluate() without manually calling the framework-specific helper.

Detection order: 1. Already a list of dicts with "tool" keys → pass through 2. Object with model_dump() → convert to dict, then re-detect 3. Dict with "choices" → OpenAI format 4. Dict with "content" list containing "type" keys → Anthropic format 5. Dict with "candidates" → Gemini format 6. Dict/object with "messages" where some message has tool_calls → LangGraph 7. Object with callable all_messages → Pydantic AI 8. Object/dict with new_items → OpenAI Agents SDK 9. List whose items have content blocks with type == "tool_use" → Claude Agent SDK 10. Bare list of messages where some message has tool_calls → LangGraph

OpenTelemetry GenAI tool spans (an OTLP export with resourceSpans, or a span list containing execute_tool / MCP tools/call spans) are detected before step 2 and converted with from_otel().

Parameters:

actual (Any) – A raw LLM provider response (object or dict), or an already-formatted list of tool-call dicts.

Return type:

list[dict[str, Any]]

Returns:

List of dicts with ‘tool’ and ‘args’ keys.

Raises:
  • TypeError – If the format cannot be detected.

  • ValueError – If a LangGraph/message-list shape with tool_calls is detected but no tool calls can be extracted from it.

toolscore.integrations.from_otel(spans)[source]

Extract tool calls from OpenTelemetry GenAI tool spans.

Works with any framework or platform that records tool executions with the OpenTelemetry semantic conventions for generative AI (execute_tool spans) or for MCP (tools/call spans). Pass an OTLP JSON export ({"resourceSpans": [...]}), a list of span dicts, or a list of OpenTelemetry SDK span objects (for example from an InMemorySpanExporter).

Parameters:

spans (Any) – The exported spans.

Return type:

list[dict[str, Any]]

Returns:

One dict per tool span, in start-time order, with tool, args, result, is_error, error, duration and id. Failures come from error.type or an ERROR span status.

Example

>>> from toolscore import evaluate, from_otel
>>> result = evaluate(expected=[{"tool": "search_orders"}], actual=from_otel(otlp_export))

Result Container

class toolscore.core.EvaluationResult[source]

Container for evaluation results.

DEFAULT_WEIGHTS: ClassVar[dict[str, float]] = {'argument_f1': 0.3, 'redundant_rate': 0.1, 'required_call_recall': 0.0, 'selection_accuracy': 0.4, 'sequence_accuracy': 0.2}
__init__()[source]

Initialize evaluation result.

Return type:

None

property score: float

Compute weighted composite score from key metrics.

Default weights: selection_accuracy=0.4, argument_f1=0.3, sequence_accuracy=0.2, redundant_rate=0.1, required_call_recall=0.0.

Selection accuracy only judges the calls that were made, so by default a trace that skips required calls can still score well. Give required_call_recall a weight (for example weights={"required_call_recall": 0.3}) to make missing and failed required calls lower the score.

Returns:

Composite score between 0.0 and 1.0.

property grade: str

Letter grade ("A" … "F") for score.

property selection_accuracy: float

proportion of calls matching expected tool names.

Type:

Selection accuracy

property argument_f1: float

Argument F1 score across all tool calls.

property sequence_accuracy: float

Sequence accuracy based on edit distance.

property required_call_recall: float | None

Share of required (expected) calls that were made and did not fail, counting repeats.

None when nothing was required. See toolscore.metrics.calculate_required_call_recall().

property policy_violations: list[dict[str, Any]]

Calls that matched a forbidden rule (empty when no rules were given).

RECORD_SCHEMA_VERSION: ClassVar[str] = '2'

Version of the to_dict() record layout. Bumped on any change that could break a consumer; new keys alone do not bump it.

to_dict()[source]

Return a complete, JSON-safe record of the evaluation.

The record is meant to be stored or handed to other tools as-is:

  • schema_version: layout version of this record ("2").

  • score, grade, weights: the composite score, its letter grade and the normalized weights that produced it.

  • required_call_recall: see required_call_recall.

  • metrics: every computed metric.

  • calls: expected and actual calls with their arguments and, for actual calls, result, is_error, error, duration and cost. Values that are not JSON (matcher objects, arbitrary result objects) are stored as their repr.

  • gold_calls_count, trace_calls_count: kept for compatibility.

Return type:

dict[str, Any]

Returns:

Dictionary representation of the evaluation result.