Metrics Module

The metrics module provides functions to calculate various evaluation metrics.

Accuracy Metrics

toolscore.metrics.calculate_invocation_accuracy(gold_calls, trace_calls)[source]

Calculate tool invocation accuracy.

Measures whether the agent invoked tools when it was supposed to, and refrained from invoking when not needed.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls from gold standard.

  • trace_calls (list[ToolCall]) – Actual tool calls from agent trace.

Return type:

float

Returns:

Accuracy score between 0.0 and 1.0. - 1.0: Perfect invocation behavior - 0.0: Completely incorrect invocation behavior

toolscore.metrics.calculate_selection_accuracy(gold_calls, trace_calls)[source]

Calculate tool selection accuracy.

Measures whether the agent selected the correct tools to use, given that it decided to use tools.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls from gold standard.

  • trace_calls (list[ToolCall]) – Actual tool calls from agent trace.

Return type:

float

Returns:

Accuracy score between 0.0 and 1.0. - 1.0: All selected tools were correct - 0.0: No selected tools were correct

Sequence Metrics

toolscore.metrics.calculate_edit_distance(gold_calls, trace_calls)[source]

Calculate edit distance metrics for tool call sequences.

Computes Levenshtein edit distance between the sequence of tool names in the gold standard and the actual trace.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls from gold standard.

  • trace_calls (list[ToolCall]) – Actual tool calls from agent trace.

Returns:

  • edit_distance: Raw Levenshtein distance (lower is better)

  • normalized_distance: Distance normalized by max sequence length (0-1)

  • sequence_accuracy: 1 - normalized_distance (higher is better, 0-1)

Return type:

dict[str, float]

Argument Metrics

toolscore.metrics.calculate_argument_f1(gold_calls, trace_calls, strict=False)[source]

Calculate F1 score for argument matching.

Evaluates how well the arguments provided to each tool match the expected arguments.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls from gold standard.

  • trace_calls (list[ToolCall]) – Actual tool calls from agent trace.

  • strict (bool) – When True, disable int/float coercion and string stripping. Passed through to _compare_values().

Return type:

dict[str, float]

Gold calls whose args is None carry a “do not check arguments” expectation (tool-name-only): they are skipped entirely from argument counting and never penalize the score. An explicit args == {} keeps the strict “expect zero arguments” meaning. When every matched gold call opts out of argument checking, there is nothing to score against, so the result is a perfect f1 == 1.0 rather than an undefined 0/0.

Returns:

  • precision: Proportion of provided arguments that were correct

  • recall: Proportion of expected arguments that were provided

  • f1: Harmonic mean of precision and recall

Return type:

dict[str, float]

Parameters:

Required Calls

toolscore.metrics.calculate_required_call_recall(gold_calls, trace_calls)[source]

Share of required calls that the trace completed, counting every required call.

Each expected call needs its own actual call with the same tool name that did not fail (see ToolCall.is_error), so a contract that requires search twice is half met by one search, and a required call whose only attempt failed is not met. calculate_tool_correctness() compares sets of tool names and reports both cases as fully correct. Order and arguments are not considered here. Traces without error information count every call as completed.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls (the requirements).

  • trace_calls (list[ToolCall]) – Actual tool calls.

Return type:

float | None

Returns:

A value in [0, 1], or None when nothing is required (the metric does not apply; it is not a zero).

Example

>>> gold = [ToolCall(tool="search"), ToolCall(tool="search")]
>>> calculate_required_call_recall(gold, [ToolCall(tool="search")])
0.5

Efficiency Metrics

toolscore.metrics.calculate_redundant_call_rate(gold_calls, trace_calls)[source]

Calculate redundant call rate.

Measures inefficiency in tool use by quantifying how many tool calls were unnecessary or redundant.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls from gold standard.

  • trace_calls (list[ToolCall]) – Actual tool calls from agent trace.

Returns:

  • redundant_count: Number of redundant calls

  • total_calls: Total number of calls made

  • redundant_rate: Proportion of calls that were redundant (0-1)

  • identical_count: Calls that exactly repeat an earlier call (same tool and same arguments). Unlike redundant_count, which counts calls beyond the gold’s per-tool expectation, this isolates loop-like repetition from productive repeated use of a tool.

  • identical_rate: identical_count / total_calls (0-1)

  • error_count: Calls that failed (see ToolCall.is_error). Zero when the trace carries no error information.

  • error_rate: error_count / total_calls (0-1)

  • retry_after_error_count: Calls that repeat the immediately preceding call exactly (same tool and arguments) after that call failed, i.e. retrying a failure without changing anything.

Return type:

dict[str, float]

Reports error_count, error_rate and retry_after_error_count alongside the redundancy and loop (identical_rate) metrics.

Safety: Credentials and Forbidden Calls

See Behavior and Safety Checks for the guide, the credential formats and the JSON rule operators.

toolscore.metrics.find_secrets(trace_calls)[source]

Find credentials in tool arguments.

Parameters:

trace_calls (list[ToolCall]) – The actual tool calls, in order.

Returns:

one entry per credential with the call index, tool, argument path (headers.Authorization, files[0].content), kind and a redacted preview.

Return type:

dict[str, Any]

toolscore.metrics.redact_secrets(text)[source]

Replace every credential in text with its redacted preview.

Used for reports that may be published (Markdown job summaries, HTML), so a credential found in a trace is not copied into them in full.

Parameters:

text (str) – Any text.

Return type:

str

Returns:

The text with each credential replaced by [REDACTED kind: preview].

toolscore.metrics.check_forbidden_calls(trace_calls, rules)[source]

Report every call that matches a forbidden rule.

Parameters:
  • trace_calls (list[ToolCall]) – The actual tool calls, in order.

  • rules (list[dict[str, Any]]) – Forbidden rules (see the module docstring).

Returns:

one entry per violating call with its index in the trace, tool, args, the index of the first matching rule and that rule’s reason (if any).

Return type:

dict[str, Any]

Raises:
  • TypeError – If rules is not a list.

  • ValueError – If a rule has no tool name or non-dict args.

toolscore.metrics.rules_from_json(data)[source]

Build forbidden rules from JSON data, turning $ operators into matchers.

An argument value that is a one-key dict {"$regex": pattern}, {"$contains": item} or {"$one_of": [values]} becomes a matcher: $regex finds the pattern anywhere in the value (re.search, so a second line or a prefix does not hide it; lists are matched as their items joined with spaces), $contains is Contains and $one_of is OneOf. Any other value is compared exactly.

Parameters:

data (Any) – A list of rule dicts, e.g. parsed from a JSON file.

Return type:

list[dict[str, Any]]

Returns:

Rules ready for check_forbidden_calls().

Raises:
toolscore.metrics.load_forbidden_rules(path)[source]

Read forbidden rules from a JSON file (see rules_from_json()).

Parameters:

path (str | Path) – Path to a JSON file holding a list of rules.

Return type:

list[dict[str, Any]]

Returns:

Rules ready for check_forbidden_calls().

toolscore.metrics.policy.call_matches_rule(call_tool, call_args, rule)[source]

Whether one call matches one forbidden rule.

Parameters:
  • call_tool (str) – The called tool’s name.

  • call_args (dict[str, Any] | None) – The call’s arguments (None means none).

  • rule (dict[str, Any]) – A rule dict with tool and optional args.

Return type:

bool

Returns:

True if the tool matches and every argument listed in the rule is present in the call and matches.

Side-Effect Metrics

toolscore.metrics.calculate_side_effect_success_rate(gold_calls, trace_calls, validators=None)[source]

Calculate side-effect success rate.

Evaluates whether tool calls achieved their intended side effects or outcomes.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls with side_effects specifications.

  • trace_calls (list[ToolCall]) – Actual tool calls from agent trace.

  • validators (dict[str, Any] | None) – Optional dict of validator functions for side effects.

Returns:

  • total_checks: Total number of side effect checks

  • passed_checks: Number of passed checks

  • success_rate: Proportion of checks that passed (0-1)

  • details: List of detailed check results

Return type:

dict[str, int | float | list[dict[str, Any]]]

Performance Metrics

toolscore.metrics.calculate_latency(trace_calls)[source]

Calculate latency metrics from tool calls.

Parameters:

trace_calls (list[ToolCall]) – Tool calls with timing information.

Returns:

  • total_duration: Total time spent on all tool calls (seconds)

  • average_duration: Average duration per call (seconds)

  • max_duration: Maximum duration of any single call (seconds)

  • min_duration: Minimum duration of any single call (seconds)

Return type:

dict[str, float]

toolscore.metrics.calculate_cost_attribution(trace_calls)[source]

Calculate cost attribution from tool calls.

Parameters:

trace_calls (list[ToolCall]) – Tool calls with cost information.

Returns:

  • total_cost: Total cost of all tool calls (USD)

  • average_cost: Average cost per call (USD)

  • cost_by_tool: Dictionary mapping tool names to their total costs

Return type:

dict[str, float | dict[str, float]]

LLM-as-a-judge Metrics (Optional)

These metrics use an optional LLM judge that supports OpenAI, Anthropic, Gemini, and any OpenAI-compatible endpoint. Configure it with a JudgeConfig (or a model-name string). See the LLM-as-a-Judge guide for provider inference, env-var keys, and install extras.

toolscore.metrics.llm_judge.calculate_semantic_correctness(gold_calls, trace_calls, *, judge=None)[source]

Calculate semantic correctness using LLM-as-a-judge.

Uses an LLM to evaluate whether the trace calls are semantically equivalent to the gold standard calls, even if they differ syntactically.

By default this issues a single batched request containing every gold/trace pair and asks for a JSON array of per-pair scores. If the batched response cannot be parsed, it falls back to one request per pair.

Parameters:
  • gold_calls (list[ToolCall]) – Expected tool calls.

  • trace_calls (list[ToolCall]) – Actual tool calls from the agent.

  • judge (JudgeConfig | str | None) – Judge configuration. May be a JudgeConfig, a model-name string shorthand, or None (default JudgeConfig()).

Returns:

  • semantic_score: Overall semantic correctness (0.0 to 1.0).

  • per_call_scores: List of scores for each call pair.

  • explanations: List of explanations for each evaluation.

  • model_used: The model that was used.

  • gold_count: Number of gold calls.

  • trace_count: Number of trace calls.

Return type:

dict[str, Any]

Raises:
  • ImportError – If the required provider SDK is not installed.

  • ValueError – If no API key is available for the resolved provider.

Example

>>> gold = [ToolCall(tool="search", args={"query": "Python"})]
>>> trace = [ToolCall(tool="web_search", args={"q": "Python"})]
>>> result = calculate_semantic_correctness(gold, trace)
>>> result["semantic_score"]
0.95  # High score despite different naming
toolscore.metrics.llm_judge.calculate_batch_semantic_correctness(evaluations, *, judge=None)[source]

Calculate semantic correctness for multiple independent evaluations.

This is useful for evaluating multiple test cases at once. Each evaluation is judged with calculate_semantic_correctness() (itself batched).

Parameters:
Returns:

  • average_score: Average semantic score across all evaluations.

  • individual_scores: List of scores for each evaluation.

  • total_evaluations: Number of evaluations performed.

Return type:

dict[str, Any]