Framework Integration Guide
Toolscore ships with extractors for the most popular agent frameworks so you
can pass raw responses directly to toolscore.evaluate() or
toolscore.test_agent() without manually building lists of tool-call
dicts.
All extractors return list[dict] with "tool" and "args" keys — the
same format accepted by toolscore.evaluate(). Where the source records
them, calls also carry result, error, is_error and duration,
which feed the Behavior and Safety Checks.
LangGraph
Use toolscore.from_langgraph() (or auto_extract — it is
auto-detected) to extract tool calls from a LangGraph agent’s final state.
Accepted inputs:
A final-state dict or object with a
messageskey/attributeA plain list of messages
Each message whose tool_calls attribute/key is non-empty contributes its
entries. Order across messages is preserved.
from langgraph.prebuilt import create_react_agent
from toolscore import evaluate
from toolscore.integrations import from_langgraph
agent = create_react_agent(model, tools)
state = agent.invoke({"messages": [("user", "What's the weather in NYC?")]})
expected = [{"tool": "get_weather", "args": {"city": "NYC"}}]
result = evaluate(expected=expected, actual=from_langgraph(state))
print(result.score)
You can also use toolscore.test_agent() directly because
toolscore.auto_extract() recognises the LangGraph state shape
automatically:
import toolscore
result = toolscore.test_agent(
agent=lambda prompt: agent.invoke({"messages": [("user", prompt)]}),
input="What's the weather in NYC?",
expected=[{"tool": "get_weather", "args": {"city": "NYC"}}],
min_score=0.9,
)
Pydantic AI
Use toolscore.from_pydantic_ai() to extract tool calls from a Pydantic
AI AgentRunResult or a list of messages.
Accepted inputs:
An AgentRunResult (duck-typed: must have a callable
all_messages())A plain list of messages
Parts with part_kind == "tool-call" (or class name ToolCallPart) are
collected; args_as_dict() is called when available, otherwise args is
used (JSON string is parsed automatically).
from pydantic_ai import Agent
from toolscore import evaluate
from toolscore.integrations import from_pydantic_ai
agent = Agent("openai:gpt-4o", tools=[get_weather])
run_result = agent.run_sync("What's the weather in NYC?")
expected = [{"tool": "get_weather", "args": {"city": "NYC"}}]
result = evaluate(expected=expected, actual=from_pydantic_ai(run_result))
print(result.score)
Or use toolscore.test_agent_async() for async Pydantic AI agents:
import toolscore
async def my_agent(prompt: str):
result = await agent.run(prompt)
return result # auto_extract will detect it
result = await toolscore.test_agent_async(
agent=my_agent,
input="What's the weather in NYC?",
expected=[{"tool": "get_weather", "args": {"city": "NYC"}}],
min_score=0.9,
)
OpenAI Agents SDK
Use toolscore.from_openai_agents() to extract tool calls from an OpenAI
Agents SDK RunResult.
Accepted inputs:
A RunResult (duck-typed: must have a
new_itemsattribute)A plain list of items
Items with type == "tool_call_item" are collected; each item’s
raw_item.name and raw_item.arguments (JSON string) provide the tool
name and parsed arguments.
from agents import Agent, Runner
from toolscore import evaluate
from toolscore.integrations import from_openai_agents
agent = Agent(name="MyAgent", tools=[get_weather])
run_result = Runner.run_sync(agent, "What's the weather in NYC?")
expected = [{"tool": "get_weather", "args": {"city": "NYC"}}]
result = evaluate(expected=expected, actual=from_openai_agents(run_result))
print(result.score)
For async usage:
import toolscore
async def my_agent(prompt: str):
return await Runner.run(agent, prompt)
result = await toolscore.test_agent_async(
agent=my_agent,
input="What's the weather in NYC?",
expected=[{"tool": "get_weather", "args": {"city": "NYC"}}],
min_score=0.9,
)
Claude Agent SDK
Use toolscore.from_claude_agent_sdk() to extract tool calls from a list
of Claude Agent SDK messages.
Accepted input:
A list of messages (dicts or objects) — each message’s
contentblocks are inspected; blocks withtype == "tool_use"yield calls vianame+input.
import anthropic
from toolscore import evaluate
from toolscore.integrations import from_claude_agent_sdk
client = anthropic.Anthropic()
messages = []
# … run your agent loop, appending messages …
expected = [{"tool": "web_search", "args": {"query": "python async"}}]
result = evaluate(expected=expected, actual=from_claude_agent_sdk(messages))
print(result.score)
Note
toolscore.from_anthropic() handles a single Anthropic
Message response. Use from_claude_agent_sdk() when
you have a list of messages collected from a multi-turn agent loop.
CrewAI (experimental)
Use toolscore.from_crewai() to extract tool calls from a CrewAI result.
Warning
This extractor is experimental. CrewAI does not expose a stable structured trace API; extraction is best-effort and may not cover all CrewAI versions or configurations.
Accepted inputs:
A list of tool-result entries with
tool_name+tool_argskeys/attrsAn object or dict with a
tools_resultsattribute/key
from crewai import Crew, Agent, Task
from toolscore import evaluate
from toolscore.integrations import from_crewai
crew = Crew(agents=[...], tasks=[...])
crew_result = crew.kickoff()
# crew_result.tools_results is a list of tool invocation records
expected = [{"tool": "search", "args": {"q": "AI agents"}}]
result = evaluate(expected=expected, actual=from_crewai(crew_result))
print(result.score)
OpenTelemetry (any instrumented framework)
Frameworks and observability platforms that follow the OpenTelemetry semantic
conventions for generative AI record each tool execution as a span. Use
toolscore.from_otel() (or pass the spans straight to
toolscore.evaluate(); they are auto-detected) to turn those spans into
tool calls, so any instrumented agent can be scored without a dedicated
extractor.
Recognized spans:
GenAI
execute_toolspans:gen_ai.operation.name == "execute_tool"withgen_ai.tool.nameand, when recorded,gen_ai.tool.call.arguments,gen_ai.tool.call.resultandgen_ai.tool.call.idMCP client spans:
mcp.method.name == "tools/call"withgen_ai.tool.name(the call id comes fromjsonrpc.request.id)
Other spans (agent, chat, embeddings) are skipped. A call failed when the span
has error.type or an ERROR status. Calls are returned in start-time order
with tool, args, result, is_error, error, duration
(seconds) and id; arguments and results recorded as JSON strings are
decoded.
Accepted inputs:
An OTLP JSON export (
{"resourceSpans": [...]}), as written by OTLP/JSON exporters and the collector’s file exporterA list of span dicts
A list of SDK span objects (
ReadableSpan), for example from anInMemorySpanExporterin tests
import json
from toolscore import evaluate, from_otel
spans = json.load(open("examples/otel_genai_spans.json"))
for call in from_otel(spans):
print(call["tool"], call["is_error"], call["error"])
result = evaluate(
expected=[{"tool": "search_orders"}, {"tool": "issue_refund"}, {"tool": "send_email"}],
actual=spans,
)
print(result.required_call_recall) # 0.67: the refund call failed
From the command line, OTLP JSON files are auto-detected, or pass
--format otel:
toolscore eval gold.json spans.json --format otel
Note
Tool arguments and results are opt-in attributes in the conventions (they may contain sensitive data). Without them, Toolscore still scores tool selection, order and failures, but not arguments.
MCP sessions
To score what an agent did against an MCP server, record the session with
toolscore mcp record (see Recording real sessions) and evaluate the file;
recorded sessions and plain JSON-RPC 2.0 message lists are auto-detected:
toolscore eval gold.json session.json