Core Module
The core module provides the main evaluation functions and result container.
Main Functions
- toolscore.core.evaluate(expected, actual, weights=None, strict=False, forbidden=None)[source]
Evaluate tool calls by comparing actual against expected (in-memory).
This is the simplest way to use Toolscore - pass Python dicts directly, no file I/O required. Raw OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, and Claude Agent SDK responses are auto-detected and converted automatically, including list-shaped formats (Claude Agent SDK message lists and bare LangGraph message lists).
- Parameters:
expected (
list[dict[str,Any]]) – List of expected tool calls, each a dict with ‘tool’ and optional ‘args’.actual (
list[dict[str,Any]] |Any) – List of actual tool calls from your agent, same format. Also accepts raw OpenAI/Anthropic/Gemini response objects or dicts (auto-detected).weights (
dict[str,float] |None) – Optional custom weights for the composite score. Keys: ‘selection_accuracy’, ‘argument_f1’, ‘sequence_accuracy’, ‘redundant_rate’, ‘required_call_recall’ (0 by default; give it a weight to penalize required calls that were skipped or failed). Provided values are merged with the defaults then renormalized so that all weights sum to 1.0 before computing the composite score.strict (
bool) – When True, argument comparison uses pure equality (no int/float coercion, no string strip). Default is False (lenient matching).forbidden (
list[dict[str,Any]] |None) – Optional calls the agent must never make, as rule dicts withtool, optionalargs(values or matchers) and optionalreason. Violations are reported inmetrics["policy_metrics"]andEvaluationResult.policy_violations; they do not change the score. Seetoolscore.metrics.policy.
- Return type:
- Returns:
EvaluationResult with metrics and a composite .score property.
Example
>>> from toolscore import evaluate >>> result = evaluate( ... expected=[{"tool": "search", "args": {"q": "test"}}], ... actual=[{"tool": "search", "args": {"q": "test"}}], ... ) >>> result.score 1.0
- toolscore.core.assert_tools(expected, actual, min_score=0.9, weights=None, strict=False)[source]
Assert that actual tool calls meet a minimum composite score.
Convenience function for use in pytest or any test framework.
- Parameters:
expected (
list[dict[str,Any]]) – List of expected tool calls.actual (
list[dict[str,Any]] |Any) – List of actual tool calls.min_score (
float) – Minimum composite score required (0.0 to 1.0).weights (
dict[str,float] |None) – Optional custom weights for the composite score.strict (
bool) – When True, argument comparison uses pure equality (no int/float coercion, no string strip). Default is False.
- Return type:
- Returns:
EvaluationResult if assertion passes.
- Raises:
ValueError – If min_score is outside [0.0, 1.0].
ToolScoreAssertionError – If the composite score is below min_score.
- toolscore.core.test_agent(agent, input, expected, min_score=None, weights=None, strict=False)[source]
End-to-end test helper: run an agent, extract tool calls, evaluate.
Calls
agent(input), passes the response through auto-detection to extract tool calls, and evaluates against expected.- Parameters:
agent (
Callable[...,Any]) – Any callable that accepts a string and returns an LLM response (raw provider response or list of tool-call dicts). Must be a synchronous callable; usetest_agent_async()for async agents.input (
str) – The prompt / user message to send to the agent.expected (
list[dict[str,Any]]) – List of expected tool calls.min_score (
float|None) – If provided, raisesToolScoreAssertionErrorwhen the composite score is below this threshold.weights (
dict[str,float] |None) – Optional custom weights for the composite score.strict (
bool) – When True, argument comparison uses pure equality.
- Return type:
- Returns:
EvaluationResult with metrics and a composite
.scoreproperty.- Raises:
TypeError – If agent is an async function or returns an awaitable.
ValueError – If min_score is outside [0.0, 1.0].
ToolScoreAssertionError – If min_score is set and the score is below it.
- async toolscore.core.test_agent_async(agent, input, expected, min_score=None, weights=None, strict=False)[source]
Async end-to-end test helper: run an agent, extract tool calls, evaluate.
Works for both synchronous and asynchronous agents. Calls
agent(input)and, if the result is awaitable, awaits it.- Parameters:
agent (
Callable[...,Any]) – Any callable that accepts a string and returns an LLM response (raw provider response or list of tool-call dicts). May be sync or async.input (
str) – The prompt / user message to send to the agent.expected (
list[dict[str,Any]]) – List of expected tool calls.min_score (
float|None) – If provided, raisesToolScoreAssertionErrorwhen the composite score is below this threshold.weights (
dict[str,float] |None) – Optional custom weights for the composite score.strict (
bool) – When True, argument comparison uses pure equality.
- Return type:
- Returns:
EvaluationResult with metrics and a composite
.scoreproperty.- Raises:
ValueError – If min_score is outside [0.0, 1.0].
ToolScoreAssertionError – If min_score is set and the score is below it.
- toolscore.core.evaluate_trace(gold_file, trace_file, format='auto', validate_side_effects=True, judge=False, forbidden=None, weights=None)[source]
Evaluate an agent’s trace against gold standard.
- Parameters:
gold_file (
str|Path) – Path to gold standard specification.format (
str) – Trace format (‘auto’, ‘openai’, ‘anthropic’, ‘gemini’, ‘mcp’, ‘langchain’, ‘otel’, ‘custom’).validate_side_effects (
bool) – Whether to validate side effects.judge (
JudgeConfig|str|bool) – LLM-as-a-judge configuration for semantic evaluation.False(default) disables it.Trueuses a defaultJudgeConfig(). A string is treated as a model-name shorthand. AJudgeConfigis used as given. Provider is inferred from the model name (or set explicitly viaJudgeConfig):claude-*-> Anthropic,gemini-*-> Gemini, abase_url-> any OpenAI-compatible endpoint (Ollama/vLLM/Groq), otherwise OpenAI.forbidden (
list[dict[str,Any]] |None) – Optional calls the agent must never make (seeevaluate()); violations go tometrics["policy_metrics"].weights (
dict[str,float] |None) – Optional custom weights for the composite score (seeevaluate()).
- Return type:
- Returns:
EvaluationResult containing all computed metrics.
- Raises:
FileNotFoundError – If files don’t exist.
ValueError – If file formats are invalid.
- toolscore.core.load_gold_standard(file_path)[source]
Load gold standard specification from JSON file.
- Parameters:
- Return type:
- Returns:
List of expected tool calls.
- Raises:
FileNotFoundError – If file doesn’t exist.
ValueError – If file format is invalid.
- toolscore.core.load_trace(file_path, format='auto')[source]
Load agent trace from JSON file.
- Parameters:
- Return type:
- Returns:
List of tool calls from the trace.
- Raises:
FileNotFoundError – If file doesn’t exist.
ValueError – If format is invalid or unsupported.
- toolscore.integrations.auto_extract(actual)[source]
Auto-detect the provider format of a response and extract tool calls.
This allows passing raw OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, or Claude Agent SDK responses directly to
evaluate()without manually calling the framework-specific helper.Detection order: 1. Already a list of dicts with
"tool"keys → pass through 2. Object withmodel_dump()→ convert to dict, then re-detect 3. Dict with"choices"→ OpenAI format 4. Dict with"content"list containing"type"keys → Anthropic format 5. Dict with"candidates"→ Gemini format 6. Dict/object with"messages"where some message hastool_calls→ LangGraph 7. Object with callableall_messages→ Pydantic AI 8. Object/dict withnew_items→ OpenAI Agents SDK 9. List whose items have content blocks withtype == "tool_use"→ Claude Agent SDK 10. Bare list of messages where some message hastool_calls→ LangGraphOpenTelemetry GenAI tool spans (an OTLP export with
resourceSpans, or a span list containingexecute_tool/ MCPtools/callspans) are detected before step 2 and converted withfrom_otel().- Parameters:
actual (
Any) – A raw LLM provider response (object or dict), or an already-formatted list of tool-call dicts.- Return type:
- Returns:
List of dicts with ‘tool’ and ‘args’ keys.
- Raises:
TypeError – If the format cannot be detected.
ValueError – If a LangGraph/message-list shape with
tool_callsis detected but no tool calls can be extracted from it.
- toolscore.integrations.from_otel(spans)[source]
Extract tool calls from OpenTelemetry GenAI tool spans.
Works with any framework or platform that records tool executions with the OpenTelemetry semantic conventions for generative AI (
execute_toolspans) or for MCP (tools/callspans). Pass an OTLP JSON export ({"resourceSpans": [...]}), a list of span dicts, or a list of OpenTelemetry SDK span objects (for example from anInMemorySpanExporter).- Parameters:
spans (
Any) – The exported spans.- Return type:
- Returns:
One dict per tool span, in start-time order, with
tool,args,result,is_error,error,durationandid. Failures come fromerror.typeor an ERROR span status.
Example
>>> from toolscore import evaluate, from_otel >>> result = evaluate(expected=[{"tool": "search_orders"}], actual=from_otel(otlp_export))
Result Container
- class toolscore.core.EvaluationResult[source]
Container for evaluation results.
- DEFAULT_WEIGHTS: ClassVar[dict[str, float]] = {'argument_f1': 0.3, 'redundant_rate': 0.1, 'required_call_recall': 0.0, 'selection_accuracy': 0.4, 'sequence_accuracy': 0.2}
- property score: float
Compute weighted composite score from key metrics.
Default weights: selection_accuracy=0.4, argument_f1=0.3, sequence_accuracy=0.2, redundant_rate=0.1, required_call_recall=0.0.
Selection accuracy only judges the calls that were made, so by default a trace that skips required calls can still score well. Give
required_call_recalla weight (for exampleweights={"required_call_recall": 0.3}) to make missing and failed required calls lower the score.- Returns:
Composite score between 0.0 and 1.0.
- property selection_accuracy: float
proportion of calls matching expected tool names.
- Type:
Selection accuracy
- property required_call_recall: float | None
Share of required (expected) calls that were made and did not fail, counting repeats.
Nonewhen nothing was required. Seetoolscore.metrics.calculate_required_call_recall().
- property policy_violations: list[dict[str, Any]]
Calls that matched a
forbiddenrule (empty when no rules were given).
- RECORD_SCHEMA_VERSION: ClassVar[str] = '2'
Version of the
to_dict()record layout. Bumped on any change that could break a consumer; new keys alone do not bump it.
- to_dict()[source]
Return a complete, JSON-safe record of the evaluation.
The record is meant to be stored or handed to other tools as-is:
schema_version: layout version of this record ("2").score,grade,weights: the composite score, its letter grade and the normalized weights that produced it.required_call_recall: seerequired_call_recall.metrics: every computed metric.calls:expectedandactualcalls with their arguments and, for actual calls,result,is_error,error,durationandcost. Values that are not JSON (matcher objects, arbitrary result objects) are stored as theirrepr.gold_calls_count,trace_calls_count: kept for compatibility.