Changelog

All notable changes to this project will be documented in this file.

This project adheres to Semantic Versioning and uses Conventional Commits.

Unreleased

1.10.0 - 2026-10-03

Toolscore 1.10 scores what agents really did, not only which tools they named: record real MCP sessions or import OpenTelemetry spans, and see failed calls, required calls that never succeeded, leaked credentials and forbidden calls next to the score. Every new check was validated on real agents and real MCP servers, including GitHub’s official MCP server.

Highlights

  • Record real MCP sessions. toolscore mcp record "<server command>" -o session.json is a transparent stdio proxy: put it in your MCP client config in place of the server, use your agent as usual, and every tool call is saved with its arguments, result, error and duration. toolscore eval gold.json session.json scores it directly.

  • Behavior and safety checks. Every evaluation now reports failed tool calls, retries of a failed call with the same arguments, and credentials passed into tool arguments (API keys, tokens, private keys; reported with a redacted preview). Forbidden-call policies (forbidden= in Python, --forbidden rules.json on the CLI, does_not_call(tool, **args) in expect()) flag calls an agent must never make, and --fail-on-violations turns them into a CI gate.

  • Required calls that never succeeded are visible. required_call_recall counts each required call that was made and did not fail. The CLI prints Required calls completed: X of N when some are missing, and --weight required_call_recall=0.3 (or weights=) makes them lower the score.

  • MCP lint finds broken and hostile tool descriptions. References to tools the server does not expose (in descriptions, parameter descriptions and server instructions), and tool-poisoning patterns: hidden Unicode (tag characters, bidi controls), instructions to hide things from the user or to ignore other instructions (errors), and <IMPORTANT>-style instruction blocks (warnings). On GitHub’s official MCP server (v1.12.2, 254 tool definitions across three toolset configurations) it reports the two real stale references and no false positives.

  • OpenTelemetry import. from_otel(), --format otel and auto-detection read tool calls from OpenTelemetry GenAI spans (execute_tool) and MCP tools/call spans, from OTLP JSON exports or SDK span objects, so any instrumented framework or observability platform can feed Toolscore.

Added

Traces keep what happened

  • Tool calls keep their result, error, is_error, duration, cost, id and metadata when loaded from files or passed to evaluate(); ToolCall.is_error tells whether a call failed.

  • efficiency_metrics adds error_count, error_rate and retry_after_error_count.

  • EvaluationResult.to_dict() returns a complete, JSON-safe, versioned record (schema_version “2”): score, grade, weights, required_call_recall, all metrics, and every expected and actual call. Built for harnesses that store evidence, such as Agent Eval Flow.

Behavior and safety

  • required_call_recall metric and EvaluationResult.required_call_recall: required calls completed, counting repeats (a contract that requires search twice is half met by one search).

  • Opt-in required_call_recall score weight (0 by default), also accepted by evaluate_trace(weights=...), expect(...).with_weights(...) and toolscore eval --weight NAME=VALUE.

  • security_metrics: credentials found in tool arguments at any depth (OpenAI, Anthropic, GitHub, AWS, Google, Slack and Stripe formats and PEM private keys), with the argument path and a redacted preview. OpenAI and Anthropic matches must also look random (digits and both letter cases), so slugs are not reported. No false positives on 861 real agent tool calls. redact_secrets() replaces credentials in any text.

  • Forbidden-call policies: evaluate(..., forbidden=[...]) and evaluate_trace(..., forbidden=[...]) report policy_metrics and EvaluationResult.policy_violations. Rules name a tool and optional argument values or matchers. load_forbidden_rules() / rules_from_json() read JSON rules where values may be {"$regex": ...} (found anywhere in the value, like re.search, so a prefix or a second line does not hide it; lists are matched as their items joined with spaces), {"$contains": ...} or {"$one_of": [...]}. Policies do not change the score.

  • expect(...).does_not_call(tool, **args) accepts arguments and matchers: forbid run_shell only when the command matches rm -rf.

  • Console, Markdown and HTML reports list behavior and safety findings (forbidden calls, credentials, failed calls, blind retries) and the required calls completed. The HTML and Markdown reports now show the score and grade. All three redact credentials wherever they would print them, so a Markdown report posted to a GitHub job summary does not publish a key; the JSON report and to_dict() keep the calls as recorded.

  • The JSON report’s summary adds score, grade, weights, required_call_recall, failed_calls, policy_violations and secrets, so CI scripts can read the verdict without recomputing it.

  • toolscore eval --forbidden FILE, --weight NAME=VALUE and --fail-on-violations.

GitHub Action

  • New inputs forbidden-file and fail-on-violations (off by default) for a safety gate, and new outputs score, grade, required-call-recall and violations. The format input accepts mcp and otel.

Examples

  • examples/guardrails/: a deploy agent’s trace with a destructive command, a credential posted to a webhook and a failed deploy retried unchanged, with forbidden.json rules.

  • examples/otel_genai_spans.json: an OTLP export written by the official OpenTelemetry Python SDK.

  • The quality-gates workflow example gains a safety-check job.

MCP

  • toolscore mcp record (and MCPRecorder): records only tools/call requests and their responses, writes atomically, and exits with the server’s exit code.

  • Lint rules for dangling tool references and tool poisoning; lint_tools(tools, instructions=...) also checks the server’s instructions.

  • The scorecard reports the token cost of the server’s instructions alongside the tool definitions (instructions_tokens, context_tokens).

  • MCPStdioClient.server_instructions and server_capabilities from the handshake; MCPToolResult.text and content_to_text() render every MCP content type (text, embedded resources, resource links, images, audio).

Integrations

  • from_otel(), OTelAdapter and --format otel for OpenTelemetry GenAI tool spans (OTLP JSON exports, span dicts, or SDK ReadableSpan objects). Arguments, results, call ids, errors (error.type or an ERROR status) and durations are kept; calls are ordered by start time.

Fixed

  • MCP traces no longer contain phantom calls. MCPAdapter turned each JSON-RPC response into a separate unknown tool call, which inflated call counts and lowered selection and sequence scores. Responses are now paired with their requests by id, and their result or error is attached to the call. Scores for MCP traces can change; re-approve baselines after upgrading.

  • MCP sessions are auto-detected. With --format auto, a recorded session ({"format": "mcp", "messages": [...]}) was routed to the OpenAI adapter and a plain JSON-RPC message list to the custom adapter; both loaded as zero calls.

  • Embedded resources are kept. Tool results that return MCP resource content were reduced to an empty string.

  • MCP server commands with spaces. toolscore mcp list|lint|test|record re-joined a multi-token command with spaces and split it again, so an argument containing a space (-- npx -y server "/path with spaces", as MCP client configs pass it) was broken in two. Several tokens are now used as given; a single quoted string is still split like a shell command.

Notes

  • The default composite score is unchanged. It judges the calls that were made, so a trace that skips required calls can still score well; use the required_call_recall weight or the new console line to catch that.

  • The weights reported by to_dict() and the JSON report now include required_call_recall (0.0 unless you set it). Passing the four existing weight names works as before.

  • Missing-tool references and <IMPORTANT>-style blocks are lint warnings; hidden Unicode and concealment or override instructions are errors, and lower the MCP scorecard’s lint score.

  • A call counts as failed when it has "is_error": true or a non-empty error; "", null, {} and [] mean no error.

1.9.1 - 2026-10-01

Found by running toolscore mcp test against the official MCP reference servers (modelcontextprotocol/servers). Several failing grades were Toolscore’s mistakes, not the servers’.

Fixed

  • Union schema types. A property whose type is a list (["boolean", "string"], as in the official sequential-thinking server) crashed the scorecard with unhashable type: 'list', and the value generator produced None for it. Values are now generated for the first non-null type, and wrong-type probes match none of the listed types.

  • Reproducible scenarios. Scenario values came from the unseeded global random, so two runs against the same server could differ. Each tool now gets a generator seeded by its name.

  • Optional fields are typed. The linter reported pydantic optional fields (anyOf: [{"type": "string"}, {"type": "null"}]) and enum/const/$ref properties as “missing a ‘type’” (5 false errors on the official git server).

  • Schema hints are used. Happy-path values now come from examples, a non-null default, an example quoted in the description (“e.g., ‘America/New_York’”), or a well-formed value for format (uri, date-time, date, email, uuid, …) before falling back to synthetic values. The official time server previously received "sample_timezone", and its correct rejection was scored as a failure.

Known limitation

  • Servers that restrict paths to allowed roots (filesystem, git) still reject generated paths, which lowers their happy-path rate. Read those failures as “input outside the server’s sandbox”, not as server defects.

1.9.0 - 2026-09-28

Found while evaluating a real research agent (DeerFlow) with Agent Eval Flow + Toolscore: redundant_rate could not tell a looping agent from one running many different searches, and one missing call could zero the argument score of every later call.

Added

  • Identical-call (loop) metric. efficiency_metrics now also reports identical_count and identical_rate: calls that exactly repeat an earlier call (same tool, same arguments, key order ignored). redundant_rate keeps its meaning (calls beyond the gold’s per-tool expectation), so a research agent that runs many different searches is no longer indistinguishable from one stuck repeating the same call. Additive; composite scores are unchanged.

Fixed

  • Argument and side-effect matching now pairs calls one-to-one. Each expected call used to be matched to the first actual call with the same tool name at the same or a later position. One missing, extra or reordered call therefore shifted every later comparison, and a correct call could score 0. Expected calls are now paired with actual calls of the same tool by the best one-to-one assignment (exact for up to 12 calls per tool, greedy above). Examples: a missed earlier call went from argument_f1 0.00 to 0.67; two same-tool calls in swapped order, and a wrong attempt followed by the right one, went from 0.00 to 1.00 (the extra attempt is still penalised by redundant_rate and sequence_accuracy). An agent that correctly calls no tools when none are expected now gets argument_f1 1.0 (was 0.0). Scores can rise for existing baselines and snapshots; re-approve after upgrading.

  • Claude LLM judge works with current anthropic SDKs. Recent SDK releases removed temperature from messages.create, so the Anthropic judge raised TypeError: unexpected keyword argument 'temperature'. temperature is now passed only when the installed SDK accepts it.

1.8.1 - 2026-06-19

This entry covers 1.7.0 through 1.8.1 (released 2026-06-13 to 2026-06-19).

Highlights

  • Snapshot testing — record, approve, replay. Stop hand-writing expected tool calls: toolscore init scaffolds a suite, the first pytest run records your agent’s calls, toolscore approve --all blesses the baseline, and CI replays it forever. Jest snapshots for agents.

  • MCP Scorecard & instant health-check. toolscore mcp test "<server command>" — or toolscore demo for a zero-setup sample server — auto-generates happy-path and edge-case scenarios from tool schemas, runs them, lints definitions, measures each tool’s context-token cost, and prints an A–F grade with a ranked Top issues to fix list and concrete fixes. Export Markdown reports and gate CI with --fail-under or --ci (writes the verdict to $GITHUB_STEP_SUMMARY and fails on blocking issues) — a fast, deterministic, offline quality check for MCP servers.

  • Fluent expect() API, matchers, and rich diffs. expect(agent).on(prompt).calls("tool", arg=ANY).does_not_call(...).with_score(0.9).run(), with ANY/Regex/Approx/Contains/OneOf/IsType matchers and aligned expected-vs-actual failure tables.

  • Native everywhere. Raw responses from OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, and CrewAI (experimental) go straight into evaluate()/expect()/snapshots — zero glue. Async agents via test_agent_async() / .run_async().

  • LLM judge for every provider. Optional semantic judging via OpenAI, Anthropic, Gemini, or any OpenAI-compatible endpoint (Ollama/vLLM/Groq) — one judge= parameter and matching [llm]/[anthropic]/[gemini] extras.

Added

Snapshot Testing

  • snapshot_check(), Snapshot, and SnapshotStore — record-approve-replay state machine backed by plain JSON files under .toolscore/snapshots/

  • toolscore_snapshot pytest fixture: records on first run, replays approved baselines, fails on drift, and prints a Jest-style terminal summary (toolscore: 1 snapshot created (pending approval), 5 passed)

  • Pytest options --toolscore-update (re-record + re-approve), --toolscore-snapshot-dir, and --toolscore-allow-pending; TOOLSCORE_RECORD_UPDATE=1 mirrors --toolscore-update

  • CLI commands: toolscore record (subprocess or --from-trace mode), toolscore approve [NAME|--all], and toolscore snapshots list|show|rm

  • CI-safe by design: snapshots are never created or auto-approved when the CI env var is set — missing/pending snapshots fail the build

MCP Scorecard

  • Zero-dependency MCP stdio client (MCPStdioClient) with config-file support for Claude Desktop style files (--config / --server)

  • Scenario harness: happy-path and edge-case scenarios generated from each tool’s input schema

  • Schema linting (lint_tools) for missing descriptions, malformed schemas, undeclared required lists, and more

  • MCPScorecard with an A–F grade blending happy-path pass rate (60%), edge resilience (20%), and lint cleanliness (20%)

  • CLI commands: toolscore mcp list, toolscore mcp lint, toolscore mcp test with --cases, --no-edge-cases, --report md|json --output, and --fail-under A-F

  • GitHub Action MCP mode: set mcp-command (and optionally mcp-fail-under) to grade an MCP server in CI instead of running the eval path

Fluent API, Matchers & Diffs

  • expect() / Expectation fluent builder: .on(), .calls(), .then_calls(), .does_not_call(), .with_score(), .with_weights(), .with_strict_args(), .run(), .run_async()

  • Argument matchers usable anywhere expected args appear: ANY, Regex, Approx, Contains, OneOf, IsType

  • Rich failure diffs: ToolScoreAssertionError messages now embed an aligned expected-vs-actual table with per-argument mismatches and targeted tips (colored on a TTY, plain text in CI logs)

Framework & Async Support

  • New extractors: from_langgraph(), from_pydantic_ai(), from_openai_agents(), from_claude_agent_sdk(), and from_crewai() (experimental); auto_extract() detects all of them, including list-shaped message formats

  • test_agent_async() and expect(...).run_async() for async agents; sync entry points raise a clear TypeError pointing at the async variant

Multi-Provider LLM Judge

  • JudgeConfig dataclass with provider inference from the model name (claude-* → Anthropic, gemini-* → Gemini, base_url → any OpenAI-compatible endpoint, otherwise OpenAI)

  • CLI flags --llm-model, --llm-provider, and --llm-base-url on toolscore eval (e.g. Ollama: --llm-model llama3.1 --llm-base-url http://localhost:11434/v1)

  • New extras: tool-scorer[anthropic] and tool-scorer[gemini]

Scaffolding & Quality

  • toolscore init reworked into a framework-detecting wizard: detects your agent framework (LangGraph, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, CrewAI, raw SDKs, or generic), writes a pytest suite that passes immediately, and adds a snapshot-replay GitHub Actions workflow

  • evaluate() validates inputs: TypeError for non-list expected, ValueError for unknown, negative, or non-finite weight keys; assert_tools()/test_agent() validate min_score range early

  • assert_score() on ToolscoreAssertions and the toolscore_assert_tools fixture bridge the pytest plugin to the in-memory API

  • py.typed marker (PEP 561) — mypy/pyright now see inline annotations

  • Strict argument comparison mode (strict=True / .with_strict_args()): pure equality, no int/float coercion or string stripping

Changed

This release intentionally changes a few behaviors. Review these before upgrading:

  • evaluate_trace() takes a single judge parameter. The old use_llm_judge, llm_judge_model, and llm_judge_api_key keyword arguments are gone. Pass judge=False (default), judge=True, a model-name string, or a JudgeConfig — e.g. evaluate_trace(gold, trace, judge=JudgeConfig(model="gemini-2.0-flash")).

  • Composite-score weights are renormalized. Custom weights= are merged with the defaults and scaled so they sum to 1.0 before the composite score is computed. Scores produced with partial weight overrides may differ from previous releases; the relative ordering of metrics you emphasized is preserved.

  • Omitted gold arguments now mean “do not check arguments”. An expected call with args omitted (or null) is a tool-name-only expectation: the tool must be called, but any arguments are accepted. An explicit "args": {} keeps its strict meaning — “expect this tool to be called with zero arguments”. This affects argument_f1, tool_correctness, trajectory, the composite score, gold-file loading, the fluent .calls("tool") (no kwargs = don’t check args), and failure-diff rendering.

  • ToolCall.args preserves None. It is no longer coerced to {} on construction, keeping “do not check” (None) distinct from “expect zero args” ({}). Code that assumed .args is always a dict should handle None.

  • The min_accuracy pytest marker is removed. It was dead code; use assert_tools(min_score=...), the toolscore_snapshot fixture, or the toolscore_assert helpers instead.

  • Importing toolscore no longer creates a traces/ directory. Trace capture creates its output directory lazily on first write, so simply importing the package leaves your filesystem untouched.

  • Cost estimator pricing refreshed (February 2026), including a claude-opus-4-6 entry.

Fixed

  • Zero-argument tool calls now score a perfect argument F1 when matched exactly (instead of being penalized as having no overlapping args)

  • evaluate() routes actual through auto_extract(), so list-shaped raw formats (Claude Agent SDK / LangGraph message lists) are detected even when passed as plain lists

  • LLM judge batching falls back to per-pair requests only on parse failures; transport and auth errors propagate instead of fanning out into doomed retries

  • MCPStdioClient.list_tools() guards against runaway nextCursor pagination loops

  • Argument comparison recurses strictly through nested dicts and lists

  • JudgeConfig is exported from the package root

1.6.0

Added - Instant Value: Zero-Friction API

Auto-Detect Provider Responses

  • auto_extract() — auto-detect OpenAI, Anthropic, and Gemini responses and extract tool calls

  • evaluate() and assert_tools() now accept raw provider responses as the actual argument — no need to import or call from_openai() / from_anthropic() / from_gemini() manually

  • Supports response objects (with model_dump()), plain dicts, and already-formatted lists

End-to-End Agent Testing

  • test_agent() — run an agent callable, extract tool calls, evaluate, and optionally assert a minimum score, all in one call

  • Accepts any callable that returns an LLM response (raw or pre-formatted)

Data-Driven Pytest Decorator

  • @toolscore.cases() — parametrize pytest test functions with a list of test-case dicts

  • Thin wrapper around pytest.mark.parametrize with automatic key extraction and test IDs

  • Lazy pytest import avoids breaking non-pytest users

1.5.0

Added - In-Memory API, Integration Helpers & Simplified Output

In-Memory Python API

  • New evaluate() function accepting Python dicts directly - no file I/O required

  • New assert_tools() one-liner for pytest: assert_tools(expected, actual, min_score=0.9)

  • Composite .score property on EvaluationResult (weighted average of key metrics)

  • Convenience properties: .selection_accuracy, .argument_f1, .sequence_accuracy

  • Custom weights support for composite score calculation

  • ToolScoreAssertionError with detailed failure messages

Integration Helpers (toolscore.integrations)

  • from_openai(response) - extract tool calls from OpenAI ChatCompletion responses

  • from_anthropic(response) - extract tool calls from Anthropic Message responses

  • from_gemini(response) - extract tool calls from Google Gemini responses

  • Works with both response objects and plain dicts

  • Supports modern and legacy formats (e.g., tool_calls and function_call for OpenAI)

Simplified CLI Output

  • Default output now shows 4 key metrics: Overall Score, Selection Accuracy, Argument F1, Sequence Accuracy

  • PASS/WARN/FAIL verdict with overall score

  • Full detailed output available with --verbose flag

  • Added toolscore as CLI alias alongside tool-scorer

README & Positioning

  • New tagline: “Lightweight tool-call testing for LLM agents - deterministic, local, zero API cost”

  • Leads with 3-line Python API example

  • Honest comparison table including DeepEval, Ragas, and Inspect AI

  • “When to use Toolscore vs. alternatives” section

  • Advanced features moved to dedicated section

1.4.0 - 2026-01-09

Added - Self-Explaining Metrics, Regression Testing & GitHub Action

Self-Explaining Metrics

  • Know exactly WHY your agent failed with detailed explanations after each metric

  • Automatic detection of tool name mismatches and similar names using SequenceMatcher

  • Actionable tips like “use –llm-judge to catch semantic equivalence”

  • Per-metric breakdowns showing MISSING, EXTRA, and MISMATCH items with severity levels

  • New toolscore/explainer.py module with comprehensive explanation generation

  • Categories: missing tools, extra tools, argument mismatches, type errors, value mismatches

  • Tips tailored to specific failure patterns (low precision vs low recall, etc.)

  • Impact: Users immediately understand what went wrong without manual debugging

Regression Testing (toolscore regression)

  • New CLI command for CI/CD regression detection

  • Save baselines with --save-baseline flag on eval command

  • Automatic PASS/FAIL with configurable thresholds (default: 5%)

  • Detailed delta reports showing improvements and regressions for each metric

  • Exit codes: 0=PASS, 1=FAIL (regression), 2=ERROR

  • Baseline includes:

    • Version tracking

    • Timestamp

    • Gold file hash for verification

    • All core metrics

  • Comparison shows:

    • Baseline vs current values

    • Absolute delta and percentage change

    • Status per metric (REGRESSION/IMPROVED/OK)

  • Impact: CI/CD pipelines can automatically catch agent degradation

GitHub Action

  • Official GitHub Action for one-click CI setup

  • Available on GitHub Actions Marketplace

  • Supports both threshold and regression testing modes

  • Features:

    • Automatic HTML report generation as artifacts

    • Job summary with evaluation results

    • Configurable accuracy thresholds

    • Regression testing against baselines

    • All trace format support

  • Example usage:

    - uses: yotambraun/toolscore@v1
      with:
        gold-file: tests/gold_standard.json
        trace-file: tests/agent_trace.json
        threshold: '0.90'
    
  • Impact: Zero-config CI/CD setup for any repository

Changed

  • Enhanced console output with “What Went Wrong” section showing top issues

  • Tips section now shows actionable suggestions based on detected problems

  • Metrics table now shows self-explaining descriptions (e.g., “3 of 4 correct”)

  • Removed redundant “Suggestions for Improvement” section (replaced by new explainer)

New Files

  • toolscore/explainer.py - Self-explaining metrics generation

  • toolscore/baseline.py - Baseline save/load/compare for regression testing

  • action.yml - GitHub Action definition

1.2.0 - 2025-10-28

Added - 🎯 Multi-Provider Support & Export Formats

🤖 Google Gemini Adapter

  • Full support for Google Gemini function calling traces

  • Auto-detection for Gemini response format

  • Handles multiple Gemini format variations:

    • candidates with functionCall (camelCase)

    • function_call (snake_case) alternative format

    • Direct parts list format

    • Mixed content with text and function calls

  • Comprehensive test coverage (72.82%)

  • Impact: Now supports all major LLM providers (OpenAI, Anthropic, Gemini)

📊 CSV Export (--csv)

  • Export evaluation results to CSV format for Excel/Google Sheets

  • Human-readable formatting with percentage values

  • Flattened metrics structure for easy sorting/filtering

  • Ideal for sharing results with non-technical stakeholders

  • 94.59% test coverage

  • Impact: Business teams can analyze results in familiar spreadsheet tools

📝 Markdown Export (--markdown)

  • Export evaluation results to Markdown format for GitHub/docs

  • Beautiful tables with status indicators (Excellent/Good/Fair/Needs Improvement)

  • Collapsible details sections for clean PR comments

  • Automatic timestamp and metadata

  • Perfect for CI/CD workflows and pull request comments

  • 81.38% test coverage

  • Impact: Seamless integration with GitHub workflows

💰 LLM Cost Estimation

  • Built-in cost tracking with calculate_llm_cost() function

  • Token estimation with estimate_tokens() function

  • Trace-level cost estimation with estimate_trace_cost()

  • Cost savings calculator with calculate_cost_savings()

  • Up-to-date pricing for October 2025 models:

    • OpenAI: GPT-5 (\(1.25/\)10), GPT-5-mini, GPT-5-nano, GPT-4o, GPT-4o-mini

    • Anthropic: Sonnet 4.5 (\(3/\)15), Haiku 4.5 (\(1/\)5), Opus 4.1 (\(15/\)75)

    • Google: Gemini 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, 2.0 Flash

  • Legacy model support for backward compatibility

  • Impact: Quantify ROI and optimize agent costs

🔧 CI/CD Integration Template

  • Ready-to-use GitHub Actions workflow (.github/workflows/toolscore-example.yml)

  • Automatic PR comments with evaluation results

  • Quality gates with configurable thresholds

  • CSV/Markdown/HTML report generation

  • Impact: Zero-config CI/CD setup

🛤️ Trajectory Evaluation

  • Multi-step path analysis with calculate_trajectory_accuracy()

  • Evaluates reasoning PATH taken by agent, not just final result

  • Step-by-step comparison with detailed trajectory analysis

  • Path efficiency metrics (penalizes unnecessary detours)

  • Partial trajectory matching for flexible evaluation

  • 89.29% test coverage with 13 comprehensive tests

  • Impact: Industry-standard evaluation (matches BFCL V3/V4 capabilities)

🔌 MCP (Model Context Protocol) Adapter

  • Full support for Anthropic’s Model Context Protocol (JSON-RPC 2.0)

  • Auto-detection for MCP message format

  • Handles tool requests, results, and error responses

  • Batch call support

  • 93.10% test coverage with 20 comprehensive tests

  • Impact: Future-proof support for emerging open standard

📦 Production Trace Capture

  • @capture_trace decorator for capturing real agent executions

  • Auto-save production traces as JSON test cases

  • Convert captured traces to gold standard format

  • Manual tool capture API with TraceCapture class

  • Impact: Closes production→testing feedback loop

🔍 Enhanced State Validators

  • FileSystemValidator: Content validation (file contains text, size constraints)

  • SQLValidator: Row-level validation (WHERE conditions, specific row matching)

  • Deeper validation beyond simple existence checks

  • Impact: State-based evaluation like BFCL V4

Changed

  • Enhanced README: Complete repositioning as “pytest for LLM agents”

  • Improved SEO: Expanded PyPI keywords from 16 to 40+ for better discoverability

  • Updated Comparison Table: Now compares against real competitors (LangSmith, OpenAI Evals, W&B)

  • CLI Enhancement: Added --csv and --markdown flags to eval command

  • Format Support: Added “gemini” option to all format-related commands

Testing

  • Increased test coverage from 47.80% to 55.05% (+7.25%)

  • Added 63 comprehensive tests across new features (31 from initial features, 32 from advanced features)

  • All 215 tests passing with zero bugs

  • Strict mypy type checking compliance maintained

  • New test suites:

    • 13 trajectory evaluation tests

    • 20 MCP adapter tests

    • Enhanced validator tests

1.1.0 - 2025-10-18

Added - 🎯 Week 2: Major UX & Developer Experience Improvements

🚀 Zero-Friction Onboarding (toolscore init)

  • Interactive CLI command for project setup in 30 seconds

  • Choose from 5 pre-built agent types (Weather, E-commerce, Code, RAG, Multi-tool)

  • Automatically generates gold standard templates, README, and example files

  • Impact: Onboarding time: 1+ hour → 30 seconds

⚡ Synthetic Test Generator (toolscore generate)

  • Generate comprehensive test cases from OpenAI function schemas

  • Automatic edge case and boundary value generation (60% normal, 20% boundary, 20% edge)

  • Smart value generation based on parameter names (email, url, query, etc.)

  • Schema validation metadata extraction

  • Impact: Test creation time: hours → 30 seconds

📊 Quick Compare (toolscore compare)

  • Compare multiple model traces side-by-side in single command

  • Color-coded performance comparison table

  • Automatic ranking per metric with best model highlighted

  • Overall winner calculation with weighted average scores

  • JSON comparison report export

  • Impact: Multi-model benchmarking: manual spreadsheet → 1 command

🔍 Interactive Debug Mode (--debug)

  • Step-by-step failure analysis with --debug flag on eval command

  • Interactive navigation through mismatches (next/previous/quit)

  • Side-by-side expected vs actual comparison tables

  • Context-specific fix suggestions for each failure type

  • Detailed failure categorization:

    • Missing tools (expected but never called)

    • Extra tools (called but not expected)

    • Tool name mismatches

    • Argument mismatches with field-by-field comparison

    • Missing calls (trace ended early)

    • Extra calls (redundant invocations)

  • Impact: Debugging: manual JSON diffing → guided walkthrough

💡 Actionable Error Messages

  • Automatic detection of common failure patterns during evaluation

  • Specific fix suggestions displayed in console output

  • Suggests --llm-judge for tool name mismatches

  • Suggests --verbose for schema validation errors

  • Suggests reviewing logic for missing tools

  • Suggests checking arguments for type/value errors

  • Impact: Users know exactly what to fix instead of guessing

Added - 🎯 Week 1: Core Evaluation Features

Tool Correctness Metric

  • New metric measuring whether ALL expected tools were called

  • Complements selection accuracy by checking coverage vs per-call matching

  • Reports missing_tools and extra_tools lists

  • Deterministic evaluation without LLM needed

Integrated LLM-as-a-Judge

  • Semantic evaluation now built into core evaluation engine

  • Simple --llm-judge CLI flag (no separate script needed)

  • Catches semantically equivalent tool names (e.g., “search” vs “web_search”)

  • Configurable model selection with --llm-model option

  • Previously: Separate script in metrics directory

  • Now: Integrated with single flag

Parameter Schema Validation

  • Validate argument types (string, integer, number, boolean, array, object)

  • Numeric constraints (minimum, maximum)

  • String constraints (minLength, maxLength, pattern regex)

  • Enum validation for allowed values

  • Required field checking

  • Detailed error reporting with field-level validation failures

  • Define schemas in gold standard metadata:

    {
      "metadata": {
        "schema": {
          "query": {"type": "string", "minLength": 1},
          "limit": {"type": "integer", "minimum": 1, "maximum": 100}
        }
      }
    }
    

Example Datasets

  • 5 realistic gold standard templates included:

    • weather_agent.json (Beginner - API lookup)

    • ecommerce_agent.json (Intermediate - Shopping workflow)

    • code_assistant.json (Intermediate - Code search/edit)

    • rag_agent.json (Advanced - Retrieval pipeline)

    • multi_tool_agent.json (Advanced - Research workflow)

  • Each includes schema validation examples and descriptions

  • Located in examples/datasets/ directory

Added - 🎯 Previous Features

  • LLM-as-a-judge metrics (now integrated): Optional semantic correctness evaluation using OpenAI API

    • calculate_semantic_correctness() function for single evaluations

    • calculate_batch_semantic_correctness() for batch processing

    • Detects semantically equivalent tool calls beyond exact string matching

    • Supports custom models (default: gpt-4o-mini)

  • LangChain adapter: Support for LangChain agent traces

    • Legacy AgentAction format (tool, tool_input, log)

    • Modern ToolCall format (name, args, id)

    • Alternative action/action_input format

    • Auto-detection support

  • Comprehensive test suite increasing coverage from 37% to 80%+

  • Complete Sphinx documentation with ReadTheDocs integration

  • Interactive Jupyter notebook tutorials (quickstart, custom formats, advanced metrics)

  • Rich-formatted console output with color-coded metrics tables

  • Pytest plugin for seamless test integration

  • Python-semantic-release configuration for automated versioning

  • Optional dependency groups: llm, langchain, all

Changed

  • Enhanced console output with new tables:

    • Tool Correctness metrics table

    • Schema Validation metrics table

    • Actionable suggestions section

  • Improved CLI help text for all commands

  • Updated CLI to include all new commands: init, generate, compare

  • Enhanced documentation with comprehensive examples

  • Improved README with “What’s New in v1.1” section

  • Better command organization in help output

Fixed

  • Windows console Unicode compatibility (removed emoji characters causing UnicodeEncodeError)

  • SQLValidator now supports multiple database field name variations

  • CLI tests updated for correct Click exit codes

  • Moved manual test files to proper tests/manual/ directory

0.1.0 - 2025-10-13

Added

  • Initial release of Toolscore

  • Core evaluation engine for LLM tool usage

  • Support for OpenAI, Anthropic, and custom trace formats

  • Comprehensive metrics:

    • Invocation accuracy

    • Selection accuracy

    • Sequence edit distance

    • Argument F1 score

    • Redundant call rate

    • Side-effect validation (HTTP, filesystem, database)

    • Performance metrics (latency, cost)

  • CLI with eval and validate commands

  • JSON and HTML report generation

  • Side-effect validators for HTTP, filesystem, and SQL operations

  • Format auto-detection

  • Adapters for multiple LLM providers

Documentation

  • README with quick start guide

  • Example files for all supported formats

  • API documentation

  • Usage examples