Toolscore Logo

Getting Started

  • Installation
    • Requirements
    • From PyPI
    • From Source
    • Development Installation
    • Optional Dependencies
      • HTTP Validators
      • LLM-as-a-judge Metrics
      • LangChain Support
      • All Optional Features
      • Documentation
    • Verification
    • Next Steps
  • Quick Start
    • Install Toolscore
    • The fastest path: toolscore init + snapshots
      • 1. Scaffold a test suite
      • 2. Run pytest — the first run records a snapshot
      • 3. Review and approve the snapshot
      • The one-line assertion
    • Evaluating captured traces
    • Basic Usage
      • Command Line Interface
      • Python API (In-Memory)
      • Python API (File-Based)
    • Creating Gold Standards
    • Supported Trace Formats
      • OpenAI Format
      • Anthropic Format
      • Custom Format
      • Other Trace Sources
    • Check behavior and safety
    • Next Steps

Guides

  • User Guide
    • Understanding Metrics
      • Tool Invocation Accuracy
      • Tool Selection Accuracy
      • Sequence Edit Distance
      • Argument Match F1 Score
      • Redundant Call Rate
      • Required Call Recall
      • Failed Calls and Blind Retries
      • Credentials and Forbidden Calls
      • Side-Effect Success Rate
      • LLM-as-a-judge Semantic Correctness (Optional)
    • Scoring Semantics
      • Omitted args vs {}
      • Custom weights are renormalized
      • Strict mode
      • Calls are paired one-to-one
      • Loops vs. exploration
    • Working with Trace Formats
      • Auto-Detection
      • Explicit Format
    • Capturing Traces
      • From OpenAI
      • From Anthropic
      • From LangChain
    • Creating Effective Gold Standards
      • Best Practices
      • Example Gold Standard
    • Side-Effect Validation
      • HTTP Validation
      • Filesystem Validation
      • Database Validation
    • Generating Reports
      • JSON Reports
      • HTML Reports
    • Batch Evaluation
    • End-to-End Agent Testing
    • Data-Driven Testing with @toolscore.cases()
    • Pytest Integration
      • Using Fixtures
      • Available Fixtures
      • Assertion Helpers
      • Example Test Suite
    • Interactive Tutorials
    • Tips and Tricks
    • Troubleshooting
      • Common Issues
      • Getting Help
  • Snapshot Testing
    • The record → approve → replay story
    • The toolscore_snapshot fixture
      • Call signature
      • Pytest options
      • Environment variables
    • CLI commands
      • toolscore record — capture snapshots
      • toolscore approve — promote a recording to a baseline
      • toolscore snapshots — list / show / rm
    • Snapshot file format
    • CI behavior
    • The update flow
    • Programmatic API
  • Argument Matchers
    • The matchers
      • ANY — match any value
      • Regex — full-match a string
      • Approx — numeric closeness
      • Contains — membership
      • OneOf — value is one of
      • IsType — type check
    • The omitted-args vs {} contract
    • Strict mode interplay
    • API reference
  • Framework Integration Guide
    • LangGraph
    • Pydantic AI
    • OpenAI Agents SDK
    • Claude Agent SDK
    • CrewAI (experimental)
    • OpenTelemetry (any instrumented framework)
    • MCP sessions
  • Fluent API — expect()
    • Hero Example
    • Quick Start
    • Chain-Method Reference
    • Using Matchers
    • Async Agents
    • Forbidden-Only Contract
    • API Reference
      • expect()
      • Expectation
        • Expectation.__init__()
        • Expectation.on()
        • Expectation.calls()
        • Expectation.then_calls()
        • Expectation.does_not_call()
        • Expectation.with_score()
        • Expectation.with_weights()
        • Expectation.with_strict_args()
        • Expectation.run()
        • Expectation.run_async()
  • Behavior and Safety Checks
    • Failed calls and blind retries
    • Required calls completed
    • Credentials in tool arguments
    • Forbidden calls
      • Rules in a JSON file
    • Failing a build on violations
    • A complete record for harnesses
  • LLM-as-a-Judge
    • JudgeConfig
    • Provider inference
    • Environment-variable keys
    • Local / OpenAI-compatible endpoints (Ollama, vLLM, Groq)
    • Batching behavior
    • Extras install matrix
    • Using the judge from evaluate_trace
    • CLI flags
    • API reference
      • JudgeConfig
        • JudgeConfig.model
        • JudgeConfig.provider
        • JudgeConfig.api_key
        • JudgeConfig.base_url
        • JudgeConfig.temperature
        • JudgeConfig.max_retries
        • JudgeConfig.__init__()
      • infer_provider()
  • Testing MCP Servers
    • Test your MCP server in 60 seconds
    • The subcommands
    • What the linter checks
    • Recording real sessions
    • What the grade means
    • Tuning the run
    • Using it in CI
    • Embedding the scorecard
    • Python API
  • Extending Toolscore
    • When you need a custom adapter
    • The BaseAdapter contract
    • Using a custom adapter
    • Request an adapter
  • Toolscore vs. Other Tools
    • The short version
    • What the others are
    • Feature comparison
    • When to reach for Toolscore
    • When to reach for a platform instead (or as well)

API Reference

  • API Reference
    • Core Module
      • Main Functions
        • evaluate()
        • assert_tools()
        • test_agent()
        • test_agent_async()
        • evaluate_trace()
        • load_gold_standard()
        • load_trace()
        • auto_extract()
        • from_otel()
      • Result Container
        • EvaluationResult
    • Adapters Module
      • Base Classes
        • ToolCall
        • BaseAdapter
      • Adapter Implementations
        • OpenAI Adapter
        • Anthropic Adapter
        • LangChain Adapter
        • Gemini Adapter
        • MCP Adapter
        • OpenTelemetry Adapter
        • Custom Adapter
    • Matchers Module
      • ANY
      • Approx
        • Approx.__init__()
        • Approx.matches()
      • Contains
        • Contains.__init__()
        • Contains.matches()
      • IsType
        • IsType.__init__()
        • IsType.matches()
      • Matcher
        • Matcher.matches()
      • OneOf
        • OneOf.__init__()
        • OneOf.matches()
      • Regex
        • Regex.__init__()
        • Regex.matches()
    • Metrics Module
      • Accuracy Metrics
        • calculate_invocation_accuracy()
        • calculate_selection_accuracy()
      • Sequence Metrics
        • calculate_edit_distance()
      • Argument Metrics
        • calculate_argument_f1()
      • Required Calls
        • calculate_required_call_recall()
      • Efficiency Metrics
        • calculate_redundant_call_rate()
      • Safety: Credentials and Forbidden Calls
        • find_secrets()
        • redact_secrets()
        • check_forbidden_calls()
        • rules_from_json()
        • load_forbidden_rules()
        • call_matches_rule()
      • Side-Effect Metrics
        • calculate_side_effect_success_rate()
      • Performance Metrics
        • calculate_latency()
        • calculate_cost_attribution()
      • LLM-as-a-judge Metrics (Optional)
        • calculate_semantic_correctness()
        • calculate_batch_semantic_correctness()
    • Snapshots Module
      • snapshot_check()
      • Snapshot
        • Snapshot.name
        • Snapshot.calls
        • Snapshot.approved
        • Snapshot.source
        • Snapshot.created_at
        • Snapshot.updated_at
        • Snapshot.schema_version
        • Snapshot.toolscore_version
        • Snapshot.__post_init__()
        • Snapshot.to_dict()
        • Snapshot.from_dict()
        • Snapshot.__init__()
      • SnapshotStore
        • SnapshotStore.__init__()
        • SnapshotStore.path_for()
        • SnapshotStore.exists()
        • SnapshotStore.load()
        • SnapshotStore.save()
        • SnapshotStore.approve()
        • SnapshotStore.delete()
        • SnapshotStore.list()
        • SnapshotStore.pending()
    • MCP Module
      • Client
        • MCPStdioClient
        • MCPToolDef
        • MCPToolResult
        • MCPError
        • MCPTimeoutError
        • content_to_text()
      • Recording sessions
        • MCPRecorder
      • Server configuration
        • MCPServerSpec
        • load_mcp_config()
      • Scorecard harness
        • generate_scenarios()
        • run_scenarios()
        • lint_tools()
        • dangling_reference_issues()
        • poisoning_issues()
        • Scenario
        • ScenarioResult
        • LintIssue
      • Scorecard
        • MCPScorecard
        • grade_meets()
        • print_scorecard()
        • scorecard_to_json()
        • scorecard_to_markdown()
    • Validators Module
      • HTTP Validator
        • HTTPValidator
      • Filesystem Validator
        • FileSystemValidator
      • SQL Validator
        • SQLValidator
    • Reports Module
      • JSON Reports
        • generate_json_report()
      • HTML Reports
        • generate_html_report()
      • Markdown and CSV Reports
        • generate_markdown_report()
        • generate_csv_report()
      • Behavior and Safety Findings
        • behavior_findings()
    • Verdict Module
      • letter_grade()
      • FixSuggestion
        • FixSuggestion.tool
        • FixSuggestion.problem
        • FixSuggestion.fix
        • FixSuggestion.priority
        • FixSuggestion.__init__()
    • Overview
    • Basic Usage
    • Main Components
    • Quick Reference
      • Core Functions
      • Adapters
      • Metrics
      • Validators
      • Reports
  • Core Module
    • Main Functions
      • evaluate()
      • assert_tools()
      • test_agent()
      • test_agent_async()
      • evaluate_trace()
      • load_gold_standard()
      • load_trace()
      • auto_extract()
      • from_otel()
    • Result Container
      • EvaluationResult
        • EvaluationResult.DEFAULT_WEIGHTS
        • EvaluationResult.__init__()
        • EvaluationResult.score
        • EvaluationResult.grade
        • EvaluationResult.selection_accuracy
        • EvaluationResult.argument_f1
        • EvaluationResult.sequence_accuracy
        • EvaluationResult.required_call_recall
        • EvaluationResult.policy_violations
        • EvaluationResult.RECORD_SCHEMA_VERSION
        • EvaluationResult.to_dict()
  • Adapters Module
    • Base Classes
      • ToolCall
        • ToolCall.tool
        • ToolCall.args
        • ToolCall.result
        • ToolCall.timestamp
        • ToolCall.duration
        • ToolCall.cost
        • ToolCall.metadata
        • ToolCall.__post_init__()
        • ToolCall.is_error
        • ToolCall.__init__()
      • BaseAdapter
        • BaseAdapter.parse()
    • Adapter Implementations
      • OpenAI Adapter
        • OpenAIAdapter
      • Anthropic Adapter
        • AnthropicAdapter
      • LangChain Adapter
        • LangChainAdapter
      • Gemini Adapter
        • GeminiAdapter
      • MCP Adapter
        • MCPAdapter
      • OpenTelemetry Adapter
        • OTelAdapter
        • tool_calls_from_otel()
      • Custom Adapter
        • CustomAdapter
  • Matchers Module
    • ANY
    • Approx
      • Approx.__init__()
      • Approx.matches()
    • Contains
      • Contains.__init__()
      • Contains.matches()
    • IsType
      • IsType.__init__()
      • IsType.matches()
    • Matcher
      • Matcher.matches()
    • OneOf
      • OneOf.__init__()
      • OneOf.matches()
    • Regex
      • Regex.__init__()
      • Regex.matches()
  • Metrics Module
    • Accuracy Metrics
      • calculate_invocation_accuracy()
      • calculate_selection_accuracy()
    • Sequence Metrics
      • calculate_edit_distance()
    • Argument Metrics
      • calculate_argument_f1()
    • Required Calls
      • calculate_required_call_recall()
    • Efficiency Metrics
      • calculate_redundant_call_rate()
    • Safety: Credentials and Forbidden Calls
      • find_secrets()
      • redact_secrets()
      • check_forbidden_calls()
      • rules_from_json()
      • load_forbidden_rules()
      • call_matches_rule()
    • Side-Effect Metrics
      • calculate_side_effect_success_rate()
    • Performance Metrics
      • calculate_latency()
      • calculate_cost_attribution()
    • LLM-as-a-judge Metrics (Optional)
      • calculate_semantic_correctness()
      • calculate_batch_semantic_correctness()
  • Snapshots Module
    • snapshot_check()
    • Snapshot
      • Snapshot.name
      • Snapshot.calls
      • Snapshot.approved
      • Snapshot.source
      • Snapshot.created_at
      • Snapshot.updated_at
      • Snapshot.schema_version
      • Snapshot.toolscore_version
      • Snapshot.__post_init__()
      • Snapshot.to_dict()
      • Snapshot.from_dict()
      • Snapshot.__init__()
    • SnapshotStore
      • SnapshotStore.__init__()
      • SnapshotStore.path_for()
      • SnapshotStore.exists()
      • SnapshotStore.load()
      • SnapshotStore.save()
      • SnapshotStore.approve()
      • SnapshotStore.delete()
      • SnapshotStore.list()
      • SnapshotStore.pending()
  • MCP Module
    • Client
      • MCPStdioClient
        • MCPStdioClient.__init__()
        • MCPStdioClient.server_messages
        • MCPStdioClient.protocol_version
        • MCPStdioClient.server_info
        • MCPStdioClient.server_instructions
        • MCPStdioClient.server_capabilities
        • MCPStdioClient.__enter__()
        • MCPStdioClient.__exit__()
        • MCPStdioClient.start()
        • MCPStdioClient.close()
        • MCPStdioClient.list_tools()
        • MCPStdioClient.call_tool()
      • MCPToolDef
        • MCPToolDef.name
        • MCPToolDef.description
        • MCPToolDef.input_schema
        • MCPToolDef.__init__()
      • MCPToolResult
        • MCPToolResult.content
        • MCPToolResult.is_error
        • MCPToolResult.raw
        • MCPToolResult.duration
        • MCPToolResult.text
        • MCPToolResult.__init__()
      • MCPError
      • MCPTimeoutError
      • content_to_text()
    • Recording sessions
      • MCPRecorder
        • MCPRecorder.__init__()
        • MCPRecorder.calls_recorded
        • MCPRecorder.write()
        • MCPRecorder.run()
    • Server configuration
      • MCPServerSpec
        • MCPServerSpec.name
        • MCPServerSpec.command
        • MCPServerSpec.env
        • MCPServerSpec.__init__()
      • load_mcp_config()
    • Scorecard harness
      • generate_scenarios()
      • run_scenarios()
      • lint_tools()
      • dangling_reference_issues()
      • poisoning_issues()
      • Scenario
        • Scenario.tool
        • Scenario.arguments
        • Scenario.kind
        • Scenario.description
        • Scenario.__init__()
      • ScenarioResult
        • ScenarioResult.scenario
        • ScenarioResult.ok
        • ScenarioResult.is_error
        • ScenarioResult.duration
        • ScenarioResult.detail
        • ScenarioResult.__init__()
      • LintIssue
        • LintIssue.tool
        • LintIssue.severity
        • LintIssue.message
        • LintIssue.fix
        • LintIssue.__init__()
    • Scorecard
      • MCPScorecard
        • MCPScorecard.server_info
        • MCPScorecard.tools
        • MCPScorecard.results
        • MCPScorecard.lint
        • MCPScorecard.instructions
        • MCPScorecard.happy_pass_rate
        • MCPScorecard.edge_resilience_rate
        • MCPScorecard.lint_error_count
        • MCPScorecard.lint_warning_count
        • MCPScorecard.lint_score
        • MCPScorecard.score
        • MCPScorecard.grade
        • MCPScorecard.total_tool_tokens
        • MCPScorecard.instructions_tokens
        • MCPScorecard.context_tokens
        • MCPScorecard.__init__()
      • grade_meets()
      • print_scorecard()
      • scorecard_to_json()
      • scorecard_to_markdown()
  • Validators Module
    • HTTP Validator
      • HTTPValidator
        • HTTPValidator.__init__()
        • HTTPValidator.validate()
    • Filesystem Validator
      • FileSystemValidator
        • FileSystemValidator.__init__()
        • FileSystemValidator.validate()
    • SQL Validator
      • SQLValidator
        • SQLValidator.__init__()
        • SQLValidator.validate()
  • Reports Module
    • JSON Reports
      • generate_json_report()
    • HTML Reports
      • generate_html_report()
    • Markdown and CSV Reports
      • generate_markdown_report()
      • generate_csv_report()
    • Behavior and Safety Findings
      • behavior_findings()

Development

  • Contributing
    • Development Setup
    • Code Style
    • Type Hints
    • Testing
    • Commit Messages
    • Pull Request Process
    • Adding New Features
      • Adding a New Trace Adapter
      • Adding a New Metric
      • Adding a New Validator
    • Documentation
    • Questions?
    • Code of Conduct
    • License
  • Changelog
    • Unreleased
    • 1.10.0 - 2026-10-03
      • Highlights
      • Added
        • Traces keep what happened
        • Behavior and safety
        • GitHub Action
        • Examples
        • MCP
        • Integrations
      • Fixed
      • Notes
    • 1.9.1 - 2026-10-01
      • Fixed
      • Known limitation
    • 1.9.0 - 2026-09-28
      • Added
      • Fixed
    • 1.8.1 - 2026-06-19
      • Highlights
      • Added
        • Snapshot Testing
        • MCP Scorecard
        • Fluent API, Matchers & Diffs
        • Framework & Async Support
        • Multi-Provider LLM Judge
        • Scaffolding & Quality
      • Changed
      • Fixed
    • 1.6.0
      • Added - Instant Value: Zero-Friction API
        • Auto-Detect Provider Responses
        • End-to-End Agent Testing
        • Data-Driven Pytest Decorator
    • 1.5.0
      • Added - In-Memory API, Integration Helpers & Simplified Output
        • In-Memory Python API
        • Integration Helpers (toolscore.integrations)
        • Simplified CLI Output
        • README & Positioning
    • 1.4.0 - 2026-01-09
      • Added - Self-Explaining Metrics, Regression Testing & GitHub Action
        • Self-Explaining Metrics
        • Regression Testing (toolscore regression)
        • GitHub Action
      • Changed
      • New Files
    • 1.2.0 - 2025-10-28
      • Added - 🎯 Multi-Provider Support & Export Formats
        • 🤖 Google Gemini Adapter
        • 📊 CSV Export (--csv)
        • 📝 Markdown Export (--markdown)
        • 💰 LLM Cost Estimation
        • 🔧 CI/CD Integration Template
        • 🛤️ Trajectory Evaluation
        • 🔌 MCP (Model Context Protocol) Adapter
        • 📦 Production Trace Capture
        • 🔍 Enhanced State Validators
      • Changed
      • Testing
    • 1.1.0 - 2025-10-18
      • Added - 🎯 Week 2: Major UX & Developer Experience Improvements
        • 🚀 Zero-Friction Onboarding (toolscore init)
        • ⚡ Synthetic Test Generator (toolscore generate)
        • 📊 Quick Compare (toolscore compare)
        • 🔍 Interactive Debug Mode (--debug)
        • 💡 Actionable Error Messages
      • Added - 🎯 Week 1: Core Evaluation Features
        • Tool Correctness Metric
        • Integrated LLM-as-a-Judge
        • Parameter Schema Validation
        • Example Datasets
      • Added - 🎯 Previous Features
      • Changed
      • Fixed
    • 0.1.0 - 2025-10-13
      • Added
      • Documentation
Toolscore
  • Search


© Copyright 2025, Yotam Braun.

Built with Sphinx using a theme provided by Read the Docs.