Changelog
All notable changes to this project will be documented in this file.
This project adheres to Semantic Versioning and uses Conventional Commits.
Unreleased
1.10.0 - 2026-10-03
Toolscore 1.10 scores what agents really did, not only which tools they named: record real MCP sessions or import OpenTelemetry spans, and see failed calls, required calls that never succeeded, leaked credentials and forbidden calls next to the score. Every new check was validated on real agents and real MCP servers, including GitHub’s official MCP server.
Highlights
Record real MCP sessions.
toolscore mcp record "<server command>" -o session.jsonis a transparent stdio proxy: put it in your MCP client config in place of the server, use your agent as usual, and every tool call is saved with its arguments, result, error and duration.toolscore eval gold.json session.jsonscores it directly.Behavior and safety checks. Every evaluation now reports failed tool calls, retries of a failed call with the same arguments, and credentials passed into tool arguments (API keys, tokens, private keys; reported with a redacted preview). Forbidden-call policies (
forbidden=in Python,--forbidden rules.jsonon the CLI,does_not_call(tool, **args)inexpect()) flag calls an agent must never make, and--fail-on-violationsturns them into a CI gate.Required calls that never succeeded are visible.
required_call_recallcounts each required call that was made and did not fail. The CLI printsRequired calls completed: X of Nwhen some are missing, and--weight required_call_recall=0.3(orweights=) makes them lower the score.MCP lint finds broken and hostile tool descriptions. References to tools the server does not expose (in descriptions, parameter descriptions and server instructions), and tool-poisoning patterns: hidden Unicode (tag characters, bidi controls), instructions to hide things from the user or to ignore other instructions (errors), and
<IMPORTANT>-style instruction blocks (warnings). On GitHub’s official MCP server (v1.12.2, 254 tool definitions across three toolset configurations) it reports the two real stale references and no false positives.OpenTelemetry import.
from_otel(),--format oteland auto-detection read tool calls from OpenTelemetry GenAI spans (execute_tool) and MCPtools/callspans, from OTLP JSON exports or SDK span objects, so any instrumented framework or observability platform can feed Toolscore.
Added
Traces keep what happened
Tool calls keep their
result,error,is_error,duration,cost,idand metadata when loaded from files or passed toevaluate();ToolCall.is_errortells whether a call failed.efficiency_metricsaddserror_count,error_rateandretry_after_error_count.EvaluationResult.to_dict()returns a complete, JSON-safe, versioned record (schema_version“2”): score, grade, weights,required_call_recall, all metrics, and every expected and actual call. Built for harnesses that store evidence, such as Agent Eval Flow.
Behavior and safety
required_call_recallmetric andEvaluationResult.required_call_recall: required calls completed, counting repeats (a contract that requiressearchtwice is half met by onesearch).Opt-in
required_call_recallscore weight (0 by default), also accepted byevaluate_trace(weights=...),expect(...).with_weights(...)andtoolscore eval --weight NAME=VALUE.security_metrics: credentials found in tool arguments at any depth (OpenAI, Anthropic, GitHub, AWS, Google, Slack and Stripe formats and PEM private keys), with the argument path and a redacted preview. OpenAI and Anthropic matches must also look random (digits and both letter cases), so slugs are not reported. No false positives on 861 real agent tool calls.redact_secrets()replaces credentials in any text.Forbidden-call policies:
evaluate(..., forbidden=[...])andevaluate_trace(..., forbidden=[...])reportpolicy_metricsandEvaluationResult.policy_violations. Rules name a tool and optional argument values or matchers.load_forbidden_rules()/rules_from_json()read JSON rules where values may be{"$regex": ...}(found anywhere in the value, likere.search, so a prefix or a second line does not hide it; lists are matched as their items joined with spaces),{"$contains": ...}or{"$one_of": [...]}. Policies do not change the score.expect(...).does_not_call(tool, **args)accepts arguments and matchers: forbidrun_shellonly when the command matchesrm -rf.Console, Markdown and HTML reports list behavior and safety findings (forbidden calls, credentials, failed calls, blind retries) and the required calls completed. The HTML and Markdown reports now show the score and grade. All three redact credentials wherever they would print them, so a Markdown report posted to a GitHub job summary does not publish a key; the JSON report and
to_dict()keep the calls as recorded.The JSON report’s
summaryaddsscore,grade,weights,required_call_recall,failed_calls,policy_violationsandsecrets, so CI scripts can read the verdict without recomputing it.toolscore eval --forbidden FILE,--weight NAME=VALUEand--fail-on-violations.
GitHub Action
New inputs
forbidden-fileandfail-on-violations(off by default) for a safety gate, and new outputsscore,grade,required-call-recallandviolations. Theformatinput acceptsmcpandotel.
Examples
examples/guardrails/: a deploy agent’s trace with a destructive command, a credential posted to a webhook and a failed deploy retried unchanged, withforbidden.jsonrules.examples/otel_genai_spans.json: an OTLP export written by the official OpenTelemetry Python SDK.The quality-gates workflow example gains a safety-check job.
MCP
toolscore mcp record(andMCPRecorder): records onlytools/callrequests and their responses, writes atomically, and exits with the server’s exit code.Lint rules for dangling tool references and tool poisoning;
lint_tools(tools, instructions=...)also checks the server’sinstructions.The scorecard reports the token cost of the server’s instructions alongside the tool definitions (
instructions_tokens,context_tokens).MCPStdioClient.server_instructionsandserver_capabilitiesfrom the handshake;MCPToolResult.textandcontent_to_text()render every MCP content type (text, embedded resources, resource links, images, audio).
Integrations
from_otel(),OTelAdapterand--format otelfor OpenTelemetry GenAI tool spans (OTLP JSON exports, span dicts, or SDKReadableSpanobjects). Arguments, results, call ids, errors (error.typeor an ERROR status) and durations are kept; calls are ordered by start time.
Fixed
MCP traces no longer contain phantom calls.
MCPAdapterturned each JSON-RPC response into a separateunknowntool call, which inflated call counts and lowered selection and sequence scores. Responses are now paired with their requests by id, and their result or error is attached to the call. Scores for MCP traces can change; re-approve baselines after upgrading.MCP sessions are auto-detected. With
--format auto, a recorded session ({"format": "mcp", "messages": [...]}) was routed to the OpenAI adapter and a plain JSON-RPC message list to the custom adapter; both loaded as zero calls.Embedded resources are kept. Tool results that return MCP
resourcecontent were reduced to an empty string.MCP server commands with spaces.
toolscore mcp list|lint|test|recordre-joined a multi-token command with spaces and split it again, so an argument containing a space (-- npx -y server "/path with spaces", as MCP client configs pass it) was broken in two. Several tokens are now used as given; a single quoted string is still split like a shell command.
Notes
The default composite score is unchanged. It judges the calls that were made, so a trace that skips required calls can still score well; use the
required_call_recallweight or the new console line to catch that.The
weightsreported byto_dict()and the JSON report now includerequired_call_recall(0.0 unless you set it). Passing the four existing weight names works as before.Missing-tool references and
<IMPORTANT>-style blocks are lint warnings; hidden Unicode and concealment or override instructions are errors, and lower the MCP scorecard’s lint score.A call counts as failed when it has
"is_error": trueor a non-emptyerror;"",null,{}and[]mean no error.
1.9.1 - 2026-10-01
Found by running toolscore mcp test against the official MCP reference servers (modelcontextprotocol/servers). Several failing grades were Toolscore’s mistakes, not the servers’.
Fixed
Union schema types. A property whose
typeis a list (["boolean", "string"], as in the official sequential-thinking server) crashed the scorecard withunhashable type: 'list', and the value generator producedNonefor it. Values are now generated for the first non-null type, and wrong-type probes match none of the listed types.Reproducible scenarios. Scenario values came from the unseeded global
random, so two runs against the same server could differ. Each tool now gets a generator seeded by its name.Optional fields are typed. The linter reported pydantic optional fields (
anyOf: [{"type": "string"}, {"type": "null"}]) andenum/const/$refproperties as “missing a ‘type’” (5 false errors on the official git server).Schema hints are used. Happy-path values now come from
examples, a non-nulldefault, an example quoted in the description (“e.g., ‘America/New_York’”), or a well-formed value forformat(uri, date-time, date, email, uuid, …) before falling back to synthetic values. The official time server previously received"sample_timezone", and its correct rejection was scored as a failure.
Known limitation
Servers that restrict paths to allowed roots (filesystem, git) still reject generated paths, which lowers their happy-path rate. Read those failures as “input outside the server’s sandbox”, not as server defects.
1.9.0 - 2026-09-28
Found while evaluating a real research agent (DeerFlow) with Agent Eval Flow + Toolscore: redundant_rate could not tell a looping agent from one running many different searches, and one missing call could zero the argument score of every later call.
Added
Identical-call (loop) metric.
efficiency_metricsnow also reportsidentical_countandidentical_rate: calls that exactly repeat an earlier call (same tool, same arguments, key order ignored).redundant_ratekeeps its meaning (calls beyond the gold’s per-tool expectation), so a research agent that runs many different searches is no longer indistinguishable from one stuck repeating the same call. Additive; composite scores are unchanged.
Fixed
Argument and side-effect matching now pairs calls one-to-one. Each expected call used to be matched to the first actual call with the same tool name at the same or a later position. One missing, extra or reordered call therefore shifted every later comparison, and a correct call could score 0. Expected calls are now paired with actual calls of the same tool by the best one-to-one assignment (exact for up to 12 calls per tool, greedy above). Examples: a missed earlier call went from
argument_f10.00 to 0.67; two same-tool calls in swapped order, and a wrong attempt followed by the right one, went from 0.00 to 1.00 (the extra attempt is still penalised byredundant_rateandsequence_accuracy). An agent that correctly calls no tools when none are expected now getsargument_f11.0 (was 0.0). Scores can rise for existing baselines and snapshots; re-approve after upgrading.Claude LLM judge works with current
anthropicSDKs. Recent SDK releases removedtemperaturefrommessages.create, so the Anthropic judge raisedTypeError: unexpected keyword argument 'temperature'.temperatureis now passed only when the installed SDK accepts it.
1.8.1 - 2026-06-19
This entry covers 1.7.0 through 1.8.1 (released 2026-06-13 to 2026-06-19).
Highlights
Snapshot testing — record, approve, replay. Stop hand-writing expected tool calls:
toolscore initscaffolds a suite, the firstpytestrun records your agent’s calls,toolscore approve --allblesses the baseline, and CI replays it forever. Jest snapshots for agents.MCP Scorecard & instant health-check.
toolscore mcp test "<server command>"— ortoolscore demofor a zero-setup sample server — auto-generates happy-path and edge-case scenarios from tool schemas, runs them, lints definitions, measures each tool’s context-token cost, and prints an A–F grade with a ranked Top issues to fix list and concrete fixes. Export Markdown reports and gate CI with--fail-underor--ci(writes the verdict to$GITHUB_STEP_SUMMARYand fails on blocking issues) — a fast, deterministic, offline quality check for MCP servers.Fluent
expect()API, matchers, and rich diffs.expect(agent).on(prompt).calls("tool", arg=ANY).does_not_call(...).with_score(0.9).run(), withANY/Regex/Approx/Contains/OneOf/IsTypematchers and aligned expected-vs-actual failure tables.Native everywhere. Raw responses from OpenAI, Anthropic, Gemini, LangGraph, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, and CrewAI (experimental) go straight into
evaluate()/expect()/snapshots — zero glue. Async agents viatest_agent_async()/.run_async().LLM judge for every provider. Optional semantic judging via OpenAI, Anthropic, Gemini, or any OpenAI-compatible endpoint (Ollama/vLLM/Groq) — one
judge=parameter and matching[llm]/[anthropic]/[gemini]extras.
Added
Snapshot Testing
snapshot_check(),Snapshot, andSnapshotStore— record-approve-replay state machine backed by plain JSON files under.toolscore/snapshots/toolscore_snapshotpytest fixture: records on first run, replays approved baselines, fails on drift, and prints a Jest-style terminal summary (toolscore: 1 snapshot created (pending approval), 5 passed)Pytest options
--toolscore-update(re-record + re-approve),--toolscore-snapshot-dir, and--toolscore-allow-pending;TOOLSCORE_RECORD_UPDATE=1mirrors--toolscore-updateCLI commands:
toolscore record(subprocess or--from-tracemode),toolscore approve [NAME|--all], andtoolscore snapshots list|show|rmCI-safe by design: snapshots are never created or auto-approved when the
CIenv var is set — missing/pending snapshots fail the build
MCP Scorecard
Zero-dependency MCP stdio client (
MCPStdioClient) with config-file support for Claude Desktop style files (--config/--server)Scenario harness: happy-path and edge-case scenarios generated from each tool’s input schema
Schema linting (
lint_tools) for missing descriptions, malformed schemas, undeclared required lists, and moreMCPScorecardwith an A–F grade blending happy-path pass rate (60%), edge resilience (20%), and lint cleanliness (20%)CLI commands:
toolscore mcp list,toolscore mcp lint,toolscore mcp testwith--cases,--no-edge-cases,--report md|json --output, and--fail-under A-FGitHub Action MCP mode: set
mcp-command(and optionallymcp-fail-under) to grade an MCP server in CI instead of running the eval path
Fluent API, Matchers & Diffs
expect()/Expectationfluent builder:.on(),.calls(),.then_calls(),.does_not_call(),.with_score(),.with_weights(),.with_strict_args(),.run(),.run_async()Argument matchers usable anywhere expected args appear:
ANY,Regex,Approx,Contains,OneOf,IsTypeRich failure diffs:
ToolScoreAssertionErrormessages now embed an aligned expected-vs-actual table with per-argument mismatches and targeted tips (colored on a TTY, plain text in CI logs)
Framework & Async Support
New extractors:
from_langgraph(),from_pydantic_ai(),from_openai_agents(),from_claude_agent_sdk(), andfrom_crewai()(experimental);auto_extract()detects all of them, including list-shaped message formatstest_agent_async()andexpect(...).run_async()for async agents; sync entry points raise a clearTypeErrorpointing at the async variant
Multi-Provider LLM Judge
JudgeConfigdataclass with provider inference from the model name (claude-*→ Anthropic,gemini-*→ Gemini,base_url→ any OpenAI-compatible endpoint, otherwise OpenAI)CLI flags
--llm-model,--llm-provider, and--llm-base-urlontoolscore eval(e.g. Ollama:--llm-model llama3.1 --llm-base-url http://localhost:11434/v1)New extras:
tool-scorer[anthropic]andtool-scorer[gemini]
Scaffolding & Quality
toolscore initreworked into a framework-detecting wizard: detects your agent framework (LangGraph, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, CrewAI, raw SDKs, or generic), writes a pytest suite that passes immediately, and adds a snapshot-replay GitHub Actions workflowevaluate()validates inputs:TypeErrorfor non-listexpected,ValueErrorfor unknown, negative, or non-finite weight keys;assert_tools()/test_agent()validatemin_scorerange earlyassert_score()onToolscoreAssertionsand thetoolscore_assert_toolsfixture bridge the pytest plugin to the in-memory APIpy.typedmarker (PEP 561) — mypy/pyright now see inline annotationsStrict argument comparison mode (
strict=True/.with_strict_args()): pure equality, no int/float coercion or string stripping
Changed
This release intentionally changes a few behaviors. Review these before upgrading:
evaluate_trace()takes a singlejudgeparameter. The olduse_llm_judge,llm_judge_model, andllm_judge_api_keykeyword arguments are gone. Passjudge=False(default),judge=True, a model-name string, or aJudgeConfig— e.g.evaluate_trace(gold, trace, judge=JudgeConfig(model="gemini-2.0-flash")).Composite-score weights are renormalized. Custom
weights=are merged with the defaults and scaled so they sum to 1.0 before the composite score is computed. Scores produced with partial weight overrides may differ from previous releases; the relative ordering of metrics you emphasized is preserved.Omitted gold arguments now mean “do not check arguments”. An expected call with
argsomitted (ornull) is a tool-name-only expectation: the tool must be called, but any arguments are accepted. An explicit"args": {}keeps its strict meaning — “expect this tool to be called with zero arguments”. This affectsargument_f1,tool_correctness, trajectory, the composite score, gold-file loading, the fluent.calls("tool")(no kwargs = don’t check args), and failure-diff rendering.ToolCall.argspreservesNone. It is no longer coerced to{}on construction, keeping “do not check” (None) distinct from “expect zero args” ({}). Code that assumed.argsis always a dict should handleNone.The
min_accuracypytest marker is removed. It was dead code; useassert_tools(min_score=...), thetoolscore_snapshotfixture, or thetoolscore_asserthelpers instead.Importing
toolscoreno longer creates atraces/directory. Trace capture creates its output directory lazily on first write, so simply importing the package leaves your filesystem untouched.Cost estimator pricing refreshed (February 2026), including a
claude-opus-4-6entry.
Fixed
Zero-argument tool calls now score a perfect argument F1 when matched exactly (instead of being penalized as having no overlapping args)
evaluate()routesactualthroughauto_extract(), so list-shaped raw formats (Claude Agent SDK / LangGraph message lists) are detected even when passed as plain listsLLM judge batching falls back to per-pair requests only on parse failures; transport and auth errors propagate instead of fanning out into doomed retries
MCPStdioClient.list_tools()guards against runawaynextCursorpagination loopsArgument comparison recurses strictly through nested dicts and lists
JudgeConfigis exported from the package root
1.6.0
Added - Instant Value: Zero-Friction API
Auto-Detect Provider Responses
auto_extract()— auto-detect OpenAI, Anthropic, and Gemini responses and extract tool callsevaluate()andassert_tools()now accept raw provider responses as theactualargument — no need to import or callfrom_openai()/from_anthropic()/from_gemini()manuallySupports response objects (with
model_dump()), plain dicts, and already-formatted lists
End-to-End Agent Testing
test_agent()— run an agent callable, extract tool calls, evaluate, and optionally assert a minimum score, all in one callAccepts any callable that returns an LLM response (raw or pre-formatted)
Data-Driven Pytest Decorator
@toolscore.cases()— parametrize pytest test functions with a list of test-case dictsThin wrapper around
pytest.mark.parametrizewith automatic key extraction and test IDsLazy pytest import avoids breaking non-pytest users
1.5.0
Added - In-Memory API, Integration Helpers & Simplified Output
In-Memory Python API
New
evaluate()function accepting Python dicts directly - no file I/O requiredNew
assert_tools()one-liner for pytest:assert_tools(expected, actual, min_score=0.9)Composite
.scoreproperty onEvaluationResult(weighted average of key metrics)Convenience properties:
.selection_accuracy,.argument_f1,.sequence_accuracyCustom weights support for composite score calculation
ToolScoreAssertionErrorwith detailed failure messages
Integration Helpers (toolscore.integrations)
from_openai(response)- extract tool calls from OpenAI ChatCompletion responsesfrom_anthropic(response)- extract tool calls from Anthropic Message responsesfrom_gemini(response)- extract tool calls from Google Gemini responsesWorks with both response objects and plain dicts
Supports modern and legacy formats (e.g.,
tool_callsandfunction_callfor OpenAI)
Simplified CLI Output
Default output now shows 4 key metrics: Overall Score, Selection Accuracy, Argument F1, Sequence Accuracy
PASS/WARN/FAIL verdict with overall score
Full detailed output available with
--verboseflagAdded
toolscoreas CLI alias alongsidetool-scorer
README & Positioning
New tagline: “Lightweight tool-call testing for LLM agents - deterministic, local, zero API cost”
Leads with 3-line Python API example
Honest comparison table including DeepEval, Ragas, and Inspect AI
“When to use Toolscore vs. alternatives” section
Advanced features moved to dedicated section
1.4.0 - 2026-01-09
Added - Self-Explaining Metrics, Regression Testing & GitHub Action
Self-Explaining Metrics
Know exactly WHY your agent failed with detailed explanations after each metric
Automatic detection of tool name mismatches and similar names using SequenceMatcher
Actionable tips like “use –llm-judge to catch semantic equivalence”
Per-metric breakdowns showing MISSING, EXTRA, and MISMATCH items with severity levels
New
toolscore/explainer.pymodule with comprehensive explanation generationCategories: missing tools, extra tools, argument mismatches, type errors, value mismatches
Tips tailored to specific failure patterns (low precision vs low recall, etc.)
Impact: Users immediately understand what went wrong without manual debugging
Regression Testing (toolscore regression)
New CLI command for CI/CD regression detection
Save baselines with
--save-baselineflag onevalcommandAutomatic PASS/FAIL with configurable thresholds (default: 5%)
Detailed delta reports showing improvements and regressions for each metric
Exit codes: 0=PASS, 1=FAIL (regression), 2=ERROR
Baseline includes:
Version tracking
Timestamp
Gold file hash for verification
All core metrics
Comparison shows:
Baseline vs current values
Absolute delta and percentage change
Status per metric (REGRESSION/IMPROVED/OK)
Impact: CI/CD pipelines can automatically catch agent degradation
GitHub Action
Official GitHub Action for one-click CI setup
Available on GitHub Actions Marketplace
Supports both threshold and regression testing modes
Features:
Automatic HTML report generation as artifacts
Job summary with evaluation results
Configurable accuracy thresholds
Regression testing against baselines
All trace format support
Example usage:
- uses: yotambraun/toolscore@v1 with: gold-file: tests/gold_standard.json trace-file: tests/agent_trace.json threshold: '0.90'
Impact: Zero-config CI/CD setup for any repository
Changed
Enhanced console output with “What Went Wrong” section showing top issues
Tips section now shows actionable suggestions based on detected problems
Metrics table now shows self-explaining descriptions (e.g., “3 of 4 correct”)
Removed redundant “Suggestions for Improvement” section (replaced by new explainer)
New Files
toolscore/explainer.py- Self-explaining metrics generationtoolscore/baseline.py- Baseline save/load/compare for regression testingaction.yml- GitHub Action definition
1.2.0 - 2025-10-28
Added - 🎯 Multi-Provider Support & Export Formats
🤖 Google Gemini Adapter
Full support for Google Gemini function calling traces
Auto-detection for Gemini response format
Handles multiple Gemini format variations:
candidateswithfunctionCall(camelCase)function_call(snake_case) alternative formatDirect
partslist formatMixed content with text and function calls
Comprehensive test coverage (72.82%)
Impact: Now supports all major LLM providers (OpenAI, Anthropic, Gemini)
📊 CSV Export (--csv)
Export evaluation results to CSV format for Excel/Google Sheets
Human-readable formatting with percentage values
Flattened metrics structure for easy sorting/filtering
Ideal for sharing results with non-technical stakeholders
94.59% test coverage
Impact: Business teams can analyze results in familiar spreadsheet tools
📝 Markdown Export (--markdown)
Export evaluation results to Markdown format for GitHub/docs
Beautiful tables with status indicators (Excellent/Good/Fair/Needs Improvement)
Collapsible details sections for clean PR comments
Automatic timestamp and metadata
Perfect for CI/CD workflows and pull request comments
81.38% test coverage
Impact: Seamless integration with GitHub workflows
💰 LLM Cost Estimation
Built-in cost tracking with
calculate_llm_cost()functionToken estimation with
estimate_tokens()functionTrace-level cost estimation with
estimate_trace_cost()Cost savings calculator with
calculate_cost_savings()Up-to-date pricing for October 2025 models:
OpenAI: GPT-5 (\(1.25/\)10), GPT-5-mini, GPT-5-nano, GPT-4o, GPT-4o-mini
Anthropic: Sonnet 4.5 (\(3/\)15), Haiku 4.5 (\(1/\)5), Opus 4.1 (\(15/\)75)
Google: Gemini 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, 2.0 Flash
Legacy model support for backward compatibility
Impact: Quantify ROI and optimize agent costs
🔧 CI/CD Integration Template
Ready-to-use GitHub Actions workflow (
.github/workflows/toolscore-example.yml)Automatic PR comments with evaluation results
Quality gates with configurable thresholds
CSV/Markdown/HTML report generation
Impact: Zero-config CI/CD setup
🛤️ Trajectory Evaluation
Multi-step path analysis with
calculate_trajectory_accuracy()Evaluates reasoning PATH taken by agent, not just final result
Step-by-step comparison with detailed trajectory analysis
Path efficiency metrics (penalizes unnecessary detours)
Partial trajectory matching for flexible evaluation
89.29% test coverage with 13 comprehensive tests
Impact: Industry-standard evaluation (matches BFCL V3/V4 capabilities)
🔌 MCP (Model Context Protocol) Adapter
Full support for Anthropic’s Model Context Protocol (JSON-RPC 2.0)
Auto-detection for MCP message format
Handles tool requests, results, and error responses
Batch call support
93.10% test coverage with 20 comprehensive tests
Impact: Future-proof support for emerging open standard
📦 Production Trace Capture
@capture_tracedecorator for capturing real agent executionsAuto-save production traces as JSON test cases
Convert captured traces to gold standard format
Manual tool capture API with
TraceCaptureclassImpact: Closes production→testing feedback loop
🔍 Enhanced State Validators
FileSystemValidator: Content validation (file contains text, size constraints)
SQLValidator: Row-level validation (WHERE conditions, specific row matching)
Deeper validation beyond simple existence checks
Impact: State-based evaluation like BFCL V4
Changed
Enhanced README: Complete repositioning as “pytest for LLM agents”
Improved SEO: Expanded PyPI keywords from 16 to 40+ for better discoverability
Updated Comparison Table: Now compares against real competitors (LangSmith, OpenAI Evals, W&B)
CLI Enhancement: Added
--csvand--markdownflags toevalcommandFormat Support: Added “gemini” option to all format-related commands
Testing
Increased test coverage from 47.80% to 55.05% (+7.25%)
Added 63 comprehensive tests across new features (31 from initial features, 32 from advanced features)
All 215 tests passing with zero bugs
Strict mypy type checking compliance maintained
New test suites:
13 trajectory evaluation tests
20 MCP adapter tests
Enhanced validator tests
1.1.0 - 2025-10-18
Added - 🎯 Week 2: Major UX & Developer Experience Improvements
🚀 Zero-Friction Onboarding (toolscore init)
Interactive CLI command for project setup in 30 seconds
Choose from 5 pre-built agent types (Weather, E-commerce, Code, RAG, Multi-tool)
Automatically generates gold standard templates, README, and example files
Impact: Onboarding time: 1+ hour → 30 seconds
⚡ Synthetic Test Generator (toolscore generate)
Generate comprehensive test cases from OpenAI function schemas
Automatic edge case and boundary value generation (60% normal, 20% boundary, 20% edge)
Smart value generation based on parameter names (email, url, query, etc.)
Schema validation metadata extraction
Impact: Test creation time: hours → 30 seconds
📊 Quick Compare (toolscore compare)
Compare multiple model traces side-by-side in single command
Color-coded performance comparison table
Automatic ranking per metric with best model highlighted
Overall winner calculation with weighted average scores
JSON comparison report export
Impact: Multi-model benchmarking: manual spreadsheet → 1 command
🔍 Interactive Debug Mode (--debug)
Step-by-step failure analysis with
--debugflag on eval commandInteractive navigation through mismatches (next/previous/quit)
Side-by-side expected vs actual comparison tables
Context-specific fix suggestions for each failure type
Detailed failure categorization:
Missing tools (expected but never called)
Extra tools (called but not expected)
Tool name mismatches
Argument mismatches with field-by-field comparison
Missing calls (trace ended early)
Extra calls (redundant invocations)
Impact: Debugging: manual JSON diffing → guided walkthrough
💡 Actionable Error Messages
Automatic detection of common failure patterns during evaluation
Specific fix suggestions displayed in console output
Suggests
--llm-judgefor tool name mismatchesSuggests
--verbosefor schema validation errorsSuggests reviewing logic for missing tools
Suggests checking arguments for type/value errors
Impact: Users know exactly what to fix instead of guessing
Added - 🎯 Week 1: Core Evaluation Features
Tool Correctness Metric
New metric measuring whether ALL expected tools were called
Complements selection accuracy by checking coverage vs per-call matching
Reports
missing_toolsandextra_toolslistsDeterministic evaluation without LLM needed
Integrated LLM-as-a-Judge
Semantic evaluation now built into core evaluation engine
Simple
--llm-judgeCLI flag (no separate script needed)Catches semantically equivalent tool names (e.g., “search” vs “web_search”)
Configurable model selection with
--llm-modeloptionPreviously: Separate script in metrics directory
Now: Integrated with single flag
Parameter Schema Validation
Validate argument types (string, integer, number, boolean, array, object)
Numeric constraints (minimum, maximum)
String constraints (minLength, maxLength, pattern regex)
Enum validation for allowed values
Required field checking
Detailed error reporting with field-level validation failures
Define schemas in gold standard metadata:
{ "metadata": { "schema": { "query": {"type": "string", "minLength": 1}, "limit": {"type": "integer", "minimum": 1, "maximum": 100} } } }
Example Datasets
5 realistic gold standard templates included:
weather_agent.json(Beginner - API lookup)ecommerce_agent.json(Intermediate - Shopping workflow)code_assistant.json(Intermediate - Code search/edit)rag_agent.json(Advanced - Retrieval pipeline)multi_tool_agent.json(Advanced - Research workflow)
Each includes schema validation examples and descriptions
Located in
examples/datasets/directory
Added - 🎯 Previous Features
LLM-as-a-judge metrics (now integrated): Optional semantic correctness evaluation using OpenAI API
calculate_semantic_correctness()function for single evaluationscalculate_batch_semantic_correctness()for batch processingDetects semantically equivalent tool calls beyond exact string matching
Supports custom models (default: gpt-4o-mini)
LangChain adapter: Support for LangChain agent traces
Legacy AgentAction format (
tool,tool_input,log)Modern ToolCall format (
name,args,id)Alternative action/action_input format
Auto-detection support
Comprehensive test suite increasing coverage from 37% to 80%+
Complete Sphinx documentation with ReadTheDocs integration
Interactive Jupyter notebook tutorials (quickstart, custom formats, advanced metrics)
Rich-formatted console output with color-coded metrics tables
Pytest plugin for seamless test integration
Python-semantic-release configuration for automated versioning
Optional dependency groups:
llm,langchain,all
Changed
Enhanced console output with new tables:
Tool Correctness metrics table
Schema Validation metrics table
Actionable suggestions section
Improved CLI help text for all commands
Updated CLI to include all new commands:
init,generate,compareEnhanced documentation with comprehensive examples
Improved README with “What’s New in v1.1” section
Better command organization in help output
Fixed
Windows console Unicode compatibility (removed emoji characters causing UnicodeEncodeError)
SQLValidator now supports multiple database field name variations
CLI tests updated for correct Click exit codes
Moved manual test files to proper
tests/manual/directory
0.1.0 - 2025-10-13
Added
Initial release of Toolscore
Core evaluation engine for LLM tool usage
Support for OpenAI, Anthropic, and custom trace formats
Comprehensive metrics:
Invocation accuracy
Selection accuracy
Sequence edit distance
Argument F1 score
Redundant call rate
Side-effect validation (HTTP, filesystem, database)
Performance metrics (latency, cost)
CLI with
evalandvalidatecommandsJSON and HTML report generation
Side-effect validators for HTTP, filesystem, and SQL operations
Format auto-detection
Adapters for multiple LLM providers
Documentation
README with quick start guide
Example files for all supported formats
API documentation
Usage examples