Metrics Module
The metrics module provides functions to calculate various evaluation metrics.
Accuracy Metrics
- toolscore.metrics.calculate_invocation_accuracy(gold_calls, trace_calls)[source]
Calculate tool invocation accuracy.
Measures whether the agent invoked tools when it was supposed to, and refrained from invoking when not needed.
Sequence Metrics
- toolscore.metrics.calculate_edit_distance(gold_calls, trace_calls)[source]
Calculate edit distance metrics for tool call sequences.
Computes Levenshtein edit distance between the sequence of tool names in the gold standard and the actual trace.
- Parameters:
- Returns:
edit_distance: Raw Levenshtein distance (lower is better)
normalized_distance: Distance normalized by max sequence length (0-1)
sequence_accuracy: 1 - normalized_distance (higher is better, 0-1)
- Return type:
Argument Metrics
- toolscore.metrics.calculate_argument_f1(gold_calls, trace_calls, strict=False)[source]
Calculate F1 score for argument matching.
Evaluates how well the arguments provided to each tool match the expected arguments.
- Parameters:
- Return type:
Gold calls whose
args is Nonecarry a “do not check arguments” expectation (tool-name-only): they are skipped entirely from argument counting and never penalize the score. An explicitargs == {}keeps the strict “expect zero arguments” meaning. When every matched gold call opts out of argument checking, there is nothing to score against, so the result is a perfectf1 == 1.0rather than an undefined 0/0.
Required Calls
- toolscore.metrics.calculate_required_call_recall(gold_calls, trace_calls)[source]
Share of required calls that the trace completed, counting every required call.
Each expected call needs its own actual call with the same tool name that did not fail (see
ToolCall.is_error), so a contract that requiressearchtwice is half met by onesearch, and a required call whose only attempt failed is not met.calculate_tool_correctness()compares sets of tool names and reports both cases as fully correct. Order and arguments are not considered here. Traces without error information count every call as completed.- Parameters:
- Return type:
- Returns:
A value in
[0, 1], orNonewhen nothing is required (the metric does not apply; it is not a zero).
Example
>>> gold = [ToolCall(tool="search"), ToolCall(tool="search")] >>> calculate_required_call_recall(gold, [ToolCall(tool="search")]) 0.5
Efficiency Metrics
- toolscore.metrics.calculate_redundant_call_rate(gold_calls, trace_calls)[source]
Calculate redundant call rate.
Measures inefficiency in tool use by quantifying how many tool calls were unnecessary or redundant.
- Parameters:
- Returns:
redundant_count: Number of redundant calls
total_calls: Total number of calls made
redundant_rate: Proportion of calls that were redundant (0-1)
identical_count: Calls that exactly repeat an earlier call (same tool and same arguments). Unlike
redundant_count, which counts calls beyond the gold’s per-tool expectation, this isolates loop-like repetition from productive repeated use of a tool.identical_rate:
identical_count/total_calls(0-1)error_count: Calls that failed (see
ToolCall.is_error). Zero when the trace carries no error information.error_rate:
error_count/total_calls(0-1)retry_after_error_count: Calls that repeat the immediately preceding call exactly (same tool and arguments) after that call failed, i.e. retrying a failure without changing anything.
- Return type:
Reports error_count, error_rate and retry_after_error_count alongside
the redundancy and loop (identical_rate) metrics.
Safety: Credentials and Forbidden Calls
See Behavior and Safety Checks for the guide, the credential formats and the JSON rule operators.
- toolscore.metrics.redact_secrets(text)[source]
Replace every credential in
textwith its redacted preview.Used for reports that may be published (Markdown job summaries, HTML), so a credential found in a trace is not copied into them in full.
- toolscore.metrics.check_forbidden_calls(trace_calls, rules)[source]
Report every call that matches a forbidden rule.
- Parameters:
- Returns:
one entry per violating call with its
indexin the trace,tool,args, the index of the first matchingruleand that rule’sreason(if any).- Return type:
- Raises:
TypeError – If
rulesis not a list.ValueError – If a rule has no tool name or non-dict
args.
- toolscore.metrics.rules_from_json(data)[source]
Build forbidden rules from JSON data, turning
$operators into matchers.An argument value that is a one-key dict
{"$regex": pattern},{"$contains": item}or{"$one_of": [values]}becomes a matcher:$regexfinds the pattern anywhere in the value (re.search, so a second line or a prefix does not hide it; lists are matched as their items joined with spaces),$containsisContainsand$one_ofisOneOf. Any other value is compared exactly.- Parameters:
data (
Any) – A list of rule dicts, e.g. parsed from a JSON file.- Return type:
- Returns:
Rules ready for
check_forbidden_calls().- Raises:
TypeError – If
datais not a list.ValueError – If a rule or an operator is malformed.
- toolscore.metrics.load_forbidden_rules(path)[source]
Read forbidden rules from a JSON file (see
rules_from_json()).
Side-Effect Metrics
- toolscore.metrics.calculate_side_effect_success_rate(gold_calls, trace_calls, validators=None)[source]
Calculate side-effect success rate.
Evaluates whether tool calls achieved their intended side effects or outcomes.
- Parameters:
- Returns:
total_checks: Total number of side effect checks
passed_checks: Number of passed checks
success_rate: Proportion of checks that passed (0-1)
details: List of detailed check results
- Return type:
Performance Metrics
- toolscore.metrics.calculate_latency(trace_calls)[source]
Calculate latency metrics from tool calls.
- Parameters:
trace_calls (
list[ToolCall]) – Tool calls with timing information.- Returns:
total_duration: Total time spent on all tool calls (seconds)
average_duration: Average duration per call (seconds)
max_duration: Maximum duration of any single call (seconds)
min_duration: Minimum duration of any single call (seconds)
- Return type:
LLM-as-a-judge Metrics (Optional)
These metrics use an optional LLM judge that supports OpenAI, Anthropic, Gemini,
and any OpenAI-compatible endpoint. Configure it with a
JudgeConfig (or a model-name string). See
the LLM-as-a-Judge guide for provider inference, env-var keys, and install
extras.
- toolscore.metrics.llm_judge.calculate_semantic_correctness(gold_calls, trace_calls, *, judge=None)[source]
Calculate semantic correctness using LLM-as-a-judge.
Uses an LLM to evaluate whether the trace calls are semantically equivalent to the gold standard calls, even if they differ syntactically.
By default this issues a single batched request containing every gold/trace pair and asks for a JSON array of per-pair scores. If the batched response cannot be parsed, it falls back to one request per pair.
- Parameters:
trace_calls (
list[ToolCall]) – Actual tool calls from the agent.judge (
JudgeConfig|str|None) – Judge configuration. May be aJudgeConfig, a model-name string shorthand, orNone(defaultJudgeConfig()).
- Returns:
semantic_score: Overall semantic correctness (0.0 to 1.0).per_call_scores: List of scores for each call pair.explanations: List of explanations for each evaluation.model_used: The model that was used.gold_count: Number of gold calls.trace_count: Number of trace calls.
- Return type:
- Raises:
ImportError – If the required provider SDK is not installed.
ValueError – If no API key is available for the resolved provider.
Example
>>> gold = [ToolCall(tool="search", args={"query": "Python"})] >>> trace = [ToolCall(tool="web_search", args={"q": "Python"})] >>> result = calculate_semantic_correctness(gold, trace) >>> result["semantic_score"] 0.95 # High score despite different naming
- toolscore.metrics.llm_judge.calculate_batch_semantic_correctness(evaluations, *, judge=None)[source]
Calculate semantic correctness for multiple independent evaluations.
This is useful for evaluating multiple test cases at once. Each evaluation is judged with
calculate_semantic_correctness()(itself batched).- Parameters:
evaluations (
list[tuple[list[ToolCall],list[ToolCall]]]) – List of(gold_calls, trace_calls)tuples.judge (
JudgeConfig|str|None) – Judge configuration (seecalculate_semantic_correctness()).
- Returns:
average_score: Average semantic score across all evaluations.individual_scores: List of scores for each evaluation.total_evaluations: Number of evaluations performed.
- Return type: