MCP Module
A standard-library-only client for Model Context Protocol (MCP) servers over the
stdio transport, plus the scorecard harness (scenario generation, execution,
linting, and A–F scoring). See Testing MCP Servers for the CLI guide and
the demo script for
programmatic usage.
Client
- class toolscore.mcp.MCPStdioClient(command, env=None, cwd=None, timeout=30.0)[source]
Synchronous MCP client communicating with a server over stdio.
The client spawns the server as a subprocess, performs the MCP initialize handshake, and exposes
list_tools/call_toolhelpers. A daemon thread reads stdout lines into a queue and another drains stderr, so neither stream can deadlock the protocol.The instance is a context manager: entering calls
start()and exiting callsclose().- __init__(command, env=None, cwd=None, timeout=30.0)[source]
Initialize the client.
- Parameters:
command (
list[str]) – The command vector used to launch the server, e.g.["python", "server.py", "--flag"].env (
dict[str,str] |None) – Environment variables to overlay on top of the current process environment (os.environ);PATHand friends are preserved unless explicitly overridden.cwd (
str|Path|None) – Working directory for the server process.timeout (
float) – Default per-request timeout in seconds.
- Return type:
None
- server_messages: list[dict[str, Any]]
Notifications/requests initiated by the server that we do not handle.
- server_instructions: str | None
The server’s
instructionsfrom the initialize result, orNone. MCP clients pass these to the model with every session, so they are part of what the model reads (and of its context cost).
- __enter__()[source]
Start the server and perform the handshake.
- Return type:
- Returns:
This client instance.
- close()[source]
Shut down the server process gracefully.
Closes stdin, waits briefly for the process to exit, then terminates and finally kills it if necessary. Safe to call multiple times.
- Return type:
- list_tools()[source]
List the tools advertised by the server.
Follows pagination via the
nextCursorfield until exhausted. Guards against a misbehaving server that paginates forever: a repeated cursor or more than_MAX_LIST_TOOLS_PAGESpages raisesMCPError.- Return type:
- Returns:
The advertised tools as
MCPToolDefobjects.- Raises:
MCPError – On a protocol or transport failure, or if the server paginates in a loop (repeated cursor) or beyond the page cap.
MCPTimeoutError – If the server does not respond in time.
- call_tool(name, arguments)[source]
Invoke a tool on the server.
- Parameters:
- Return type:
- Returns:
The
MCPToolResult.is_errorisTrueboth when the server reportsisErrorand when the call yields a JSON-RPC error.- Raises:
MCPTimeoutError – If the server does not respond in time.
MCPError – On a transport failure (e.g. the process dies).
- class toolscore.mcp.MCPToolDef(name, description, input_schema)[source]
A tool advertised by an MCP server.
- Variables:
name – The tool’s unique name.
description – Human-readable description (may be empty).
input_schema – The JSON Schema for the tool’s arguments, taken verbatim from the server’s
inputSchemafield.
- Parameters:
- class toolscore.mcp.MCPToolResult(content, is_error, raw, duration)[source]
The outcome of a
tools/callinvocation.- Variables:
content – The parsed
contentlist from the call result (or the raw error payload when the call failed at the JSON-RPC level).is_error –
Trueif the result reportedisErroror the response was a JSON-RPC error.raw – The full JSON-RPC response object.
duration – Wall-clock seconds measured client-side for the round trip.
- Parameters:
- exception toolscore.mcp.MCPTimeoutError[source]
Raised when a request does not receive a response within the timeout.
- toolscore.mcp.content_to_text(content)[source]
Render MCP
contentas the text an MCP client would give to a model.- Parameters:
content (
Any) – Thecontentvalue of atools/callresult (normally a list of content items). A plain string is returned unchanged, and a JSON-RPC error object yields itsmessage.- Returns:
textas-is; an embedded textresourceas[resource <uri>]followed by its body; a binary resource,imageoraudioas a bracketed placeholder with its MIME type; aresource_linkas[resource link <uri> <name>]. Unknown item types are skipped.- Return type:
Example
>>> content_to_text([{"type": "text", "text": "ok"}, ... {"type": "resource", "resource": {"uri": "file:///a.md", "text": "# A"}}]) 'ok\n[resource file:///a.md]\n# A'
Recording sessions
- class toolscore.mcp.MCPRecorder(command, output, env=None, cwd=None)[source]
A transparent stdio proxy that records
tools/calltraffic.- Parameters:
Example
>>> recorder = MCPRecorder(["python", "server.py"], Path("session.json")) >>> exit_code = recorder.run() # relays stdin/stdout until the client disconnects
- calls_recorded
Number of completed tool calls recorded so far.
- run(stdin=None, stdout=None, stderr=None)[source]
Relay a whole session and return the server’s exit code.
Reads client messages from
stdinuntil it closes, relays server output tostdoutuntil the server exits, and passes the server’s stderr through. The trace is written at the end even if no tool was called.- Parameters:
- Return type:
- Returns:
The server process exit code.
Server configuration
- class toolscore.mcp.MCPServerSpec(name, command, env=<factory>)[source]
A single MCP server entry resolved from a config file.
- Variables:
name – The key under
mcpServersidentifying this server.command – The full command vector, i.e.
[command, *args].env – Environment variable overrides for the server process.
- Parameters:
- toolscore.mcp.load_mcp_config(path, server=None)[source]
Load an MCP server specification from a Claude Desktop style config file.
- Parameters:
server (
str|None) – Name of the server to load. IfNoneand exactly one server is defined, that server is used. IfNoneand multiple servers exist, aValueErroris raised listing the available names.
- Return type:
- Returns:
The resolved
MCPServerSpec.- Raises:
FileNotFoundError – If
pathdoes not exist.ValueError – If the file is malformed, the
mcpServerskey is missing or empty, the requestedserveris not found, orserverisNonewhile multiple servers are defined.
Scorecard harness
- toolscore.mcp.generate_scenarios(tools, cases_per_tool=3, include_edge_cases=True)[source]
Generate test scenarios for every advertised tool.
For each tool with a usable object schema this produces up to
cases_per_toolhappy-path scenarios (well-formed arguments generated viagenerate_value_from_schema(), with all required properties present) and, wheninclude_edge_casesis set and the schema permits, a handful of edge cases (missing required argument, wrong-type value, empty/zero values).Tools whose schema is empty or invalid cannot be introspected, so they each receive a single no-arguments scenario whose description flags the problem.
- toolscore.mcp.run_scenarios(client, scenarios)[source]
Execute every scenario against the server, never aborting on failure.
Each scenario is called in turn. Transport problems (
MCPError) and timeouts (MCPTimeoutError) are caught per scenario and recorded as a failed result rather than propagated, so one misbehaving tool cannot sink the whole run.The
okdetermination depends on the scenario kind:"happy"scenarios areokonly if the call completed without an error result."edge"scenarios areokwhenever the server responded at all (an error payload is an acceptable, expected reaction to bad input); they fail only on a timeout or transport crash.
- Parameters:
client (
MCPStdioClient) – A started MCP client connected to the server under test.
- Return type:
- Returns:
One
ScenarioResultper input scenario, in order.
- toolscore.mcp.lint_tools(tools, instructions=None)[source]
Statically lint tool schemas and the text a model reads.
Errors (real defects):
missing or empty
inputSchema,schema missing a
type,a declared property that lacks a
type,a name in
requiredthat is absent fromproperties.
Warnings (ergonomic nits):
missing or very short (< 10 chars) description,
a tool name that is not
snake_case,properties present but no
requiredlist declared.
Text checks over tool names, descriptions, parameter descriptions and the server
instructions(seetoolscore.mcp.lint_rules):a reference to a tool the server does not expose (warning),
hidden Unicode tag characters or bidirectional controls (error),
tool-poisoning directives such as “do not tell the user” (error),
other invisible format characters (warning).
Each issue carries a concrete
fixhint. Issues found in the instructions are reported under the tool name"<instructions>".- Parameters:
tools (
list[MCPToolDef]) – The tool definitions to lint.instructions (
str|None) – The server’sinstructionsfrom the initialize result (MCPStdioClient.server_instructions), if any.
- Return type:
- Returns:
Every issue found, across all tools.
- toolscore.mcp.lint_rules.dangling_reference_issues(tools, instructions=None)[source]
Find text that tells the model to use a tool the server does not expose.
A name counts as a tool reference only in phrasing that names a tool: a quoted name followed by tool; a quoted name after use/call/invoke/run/try that is shaped like a tool name (
snake_caseorkebab-case, so fields such asendCursorand values such asaddare not read as tools); or an unquotedsnake_casename between use/call/invoke and tool. Names that exist as a tool, a parameter, or an allowed parameter value anywhere in the server are not reported (methods such asget_commentsare often values of amethodparameter).
- toolscore.mcp.lint_rules.poisoning_issues(tools, instructions=None)[source]
Find hidden text and concealment directives in what the model reads.
Errors:
Unicode tag characters (invisible text a model still reads),
bidirectional override/isolate controls (text that displays differently than it is stored),
directives to hide things from the user (“do not tell the user”, “do not mention this to the user”, “without telling the user”) or to ignore other instructions.
Warnings:
<IMPORTANT>-,<system>- or<instructions>-style blocks, which some servers use for emphasis and published attacks use to wrap hidden instructions;other invisible format characters (zero-width characters, direction marks, BOM). They are normal in some scripts, so they are only surfaced.
Subdivision flag emoji, which are built from tag characters, are not reported.
Tool names are checked for hidden characters too.
- class toolscore.mcp.Scenario(tool, arguments, kind, description)[source]
A single planned tool invocation.
- Variables:
tool – The name of the tool to call.
arguments – The arguments to pass to the tool.
kind – Either
"happy"(well-formed, expected to succeed) or"edge"(intentionally malformed/boundary input where an error response from the server is acceptable).description – Human-readable summary of what the scenario exercises.
- Parameters:
- class toolscore.mcp.ScenarioResult(scenario, ok, is_error, duration, detail)[source]
The outcome of executing a single
Scenario.- Variables:
scenario – The scenario that produced this result.
ok – Whether the scenario is considered a pass. For
"happy"scenarios this means the call completed withoutis_error. For"edge"scenarios this means the server responded (with or without an error payload) rather than crashing or timing out – error responses to bad input are acceptable and expected.is_error – Whether the tool result reported an error.
duration – Wall-clock seconds the call took.
detail – A short outcome string (e.g. an error message excerpt).
- Parameters:
- class toolscore.mcp.LintIssue(tool, severity, message, fix='')[source]
A single problem found while statically linting a tool schema.
- Variables:
tool – The name of the offending tool.
severity – Either
"error"(a real schema defect) or"warning"(a quality nit that hurts agent ergonomics).message – A human-readable description of the issue.
fix – A short, concrete suggestion for how to resolve the issue. May be empty for issues raised without a specific remedy.
- Parameters:
Scorecard
- class toolscore.mcp.MCPScorecard(server_info, tools, results, lint, instructions=None)[source]
An aggregated quality report for a single MCP server.
- Variables:
server_info – The
serverInfodict reported during the handshake.tools – The tools advertised by the server.
results – The executed scenario results.
lint – The lint issues found in the tool schemas and the text a model reads.
instructions – The server
instructionsfrom the handshake, if any. MCP clients send them to the model with every session, so they count toward the context cost.
- Parameters:
tools (list[MCPToolDef])
results (list[ScenarioResult])
instructions (str | None)
- tools: list[MCPToolDef]
- results: list[ScenarioResult]
- property total_tool_tokens: int
Estimated total context tokens consumed by all tool definitions.
Tool definitions are sent to the model on every request, so this is a proxy for how much of the context window the server’s tools occupy.
- toolscore.mcp.grade_meets(grade, threshold)[source]
Report whether
gradeis at least as good asthreshold.
- toolscore.mcp.print_scorecard(card, console=None)[source]
Render the scorecard to a rich console.
Prints a header panel (server name/version and the big letter grade), a per-tool table (scenarios passed/total and average latency), and a lint section listing any issues.
- Parameters:
card (
MCPScorecard) – The scorecard to render.console (
Console|None) – The console to print to; a fresh one is created if omitted.
- Return type:
- toolscore.mcp.scorecard_to_json(card)[source]
Render the scorecard as a JSON-serializable dict.
- Parameters:
card (
MCPScorecard) – The scorecard to render.- Return type:
- Returns:
A plain dict (lists/dicts/strings/numbers only) capturing the grade, component scores, per-tool summaries, every scenario result, and every lint issue.
- toolscore.mcp.scorecard_to_markdown(card)[source]
Render the scorecard as Markdown for READMEs or PR comments.
- Parameters:
card (
MCPScorecard) – The scorecard to render.- Return type:
- Returns:
A Markdown document with a grade line, component scores (including the tool-definition token cost), a per-tool summary table, and a ranked “Top issues to fix” list.