MCP Module

A standard-library-only client for Model Context Protocol (MCP) servers over the stdio transport, plus the scorecard harness (scenario generation, execution, linting, and A–F scoring). See Testing MCP Servers for the CLI guide and the demo script for programmatic usage.

Client

class toolscore.mcp.MCPStdioClient(command, env=None, cwd=None, timeout=30.0)[source]

Synchronous MCP client communicating with a server over stdio.

The client spawns the server as a subprocess, performs the MCP initialize handshake, and exposes list_tools / call_tool helpers. A daemon thread reads stdout lines into a queue and another drains stderr, so neither stream can deadlock the protocol.

The instance is a context manager: entering calls start() and exiting calls close().

Parameters:
__init__(command, env=None, cwd=None, timeout=30.0)[source]

Initialize the client.

Parameters:
  • command (list[str]) – The command vector used to launch the server, e.g. ["python", "server.py", "--flag"].

  • env (dict[str, str] | None) – Environment variables to overlay on top of the current process environment (os.environ); PATH and friends are preserved unless explicitly overridden.

  • cwd (str | Path | None) – Working directory for the server process.

  • timeout (float) – Default per-request timeout in seconds.

Return type:

None

server_messages: list[dict[str, Any]]

Notifications/requests initiated by the server that we do not handle.

protocol_version: str

Protocol version negotiated during the handshake.

server_info: dict[str, Any]

serverInfo returned by the initialize handshake.

server_instructions: str | None

The server’s instructions from the initialize result, or None. MCP clients pass these to the model with every session, so they are part of what the model reads (and of its context cost).

server_capabilities: dict[str, Any]

The capabilities object from the initialize result.

__enter__()[source]

Start the server and perform the handshake.

Return type:

MCPStdioClient

Returns:

This client instance.

__exit__(*exc)[source]

Close the server process on context exit.

Return type:

None

Parameters:

exc (object)

start()[source]

Spawn the server process and run the initialize handshake.

Return type:

dict[str, Any]

Returns:

The serverInfo dictionary reported by the server (may be empty if the server omits it).

Raises:

MCPError – If the process cannot be spawned or the handshake fails.

close()[source]

Shut down the server process gracefully.

Closes stdin, waits briefly for the process to exit, then terminates and finally kills it if necessary. Safe to call multiple times.

Return type:

None

list_tools()[source]

List the tools advertised by the server.

Follows pagination via the nextCursor field until exhausted. Guards against a misbehaving server that paginates forever: a repeated cursor or more than _MAX_LIST_TOOLS_PAGES pages raises MCPError.

Return type:

list[MCPToolDef]

Returns:

The advertised tools as MCPToolDef objects.

Raises:
  • MCPError – On a protocol or transport failure, or if the server paginates in a loop (repeated cursor) or beyond the page cap.

  • MCPTimeoutError – If the server does not respond in time.

call_tool(name, arguments)[source]

Invoke a tool on the server.

Parameters:
  • name (str) – The tool name.

  • arguments (dict[str, Any]) – Arguments matching the tool’s input schema.

Return type:

MCPToolResult

Returns:

The MCPToolResult. is_error is True both when the server reports isError and when the call yields a JSON-RPC error.

Raises:
  • MCPTimeoutError – If the server does not respond in time.

  • MCPError – On a transport failure (e.g. the process dies).

class toolscore.mcp.MCPToolDef(name, description, input_schema)[source]

A tool advertised by an MCP server.

Variables:
  • name – The tool’s unique name.

  • description – Human-readable description (may be empty).

  • input_schema – The JSON Schema for the tool’s arguments, taken verbatim from the server’s inputSchema field.

Parameters:
name: str
description: str
input_schema: dict[str, Any]
__init__(name, description, input_schema)
Parameters:
Return type:

None

class toolscore.mcp.MCPToolResult(content, is_error, raw, duration)[source]

The outcome of a tools/call invocation.

Variables:
  • content – The parsed content list from the call result (or the raw error payload when the call failed at the JSON-RPC level).

  • is_error – True if the result reported isError or the response was a JSON-RPC error.

  • raw – The full JSON-RPC response object.

  • duration – Wall-clock seconds measured client-side for the round trip.

Parameters:
content: Any
is_error: bool
raw: dict[str, Any]
duration: float
property text: str

The result as the text an MCP client would show a model.

Renders every content item, not only text: embedded resources (how many servers return file contents), resource links, and placeholders for images and audio. See toolscore.mcp.content.content_to_text().

__init__(content, is_error, raw, duration)
Parameters:
Return type:

None

exception toolscore.mcp.MCPError[source]

Raised on MCP protocol or transport failures.

exception toolscore.mcp.MCPTimeoutError[source]

Raised when a request does not receive a response within the timeout.

toolscore.mcp.content_to_text(content)[source]

Render MCP content as the text an MCP client would give to a model.

Parameters:

content (Any) – The content value of a tools/call result (normally a list of content items). A plain string is returned unchanged, and a JSON-RPC error object yields its message.

Returns:

text as-is; an embedded text resource as [resource <uri>] followed by its body; a binary resource, image or audio as a bracketed placeholder with its MIME type; a resource_link as [resource link <uri> <name>]. Unknown item types are skipped.

Return type:

str

Example

>>> content_to_text([{"type": "text", "text": "ok"},
...                  {"type": "resource", "resource": {"uri": "file:///a.md", "text": "# A"}}])
'ok\n[resource file:///a.md]\n# A'

Recording sessions

class toolscore.mcp.MCPRecorder(command, output, env=None, cwd=None)[source]

A transparent stdio proxy that records tools/call traffic.

Parameters:
  • command (list[str]) – The server launch command vector.

  • output (str | Path) – Path of the trace file to write.

  • env (dict[str, str] | None) – Extra environment variables for the server process.

  • cwd (str | Path | None) – Working directory for the server process.

Example

>>> recorder = MCPRecorder(["python", "server.py"], Path("session.json"))
>>> exit_code = recorder.run()  # relays stdin/stdout until the client disconnects
__init__(command, output, env=None, cwd=None)[source]
Parameters:
Return type:

None

calls_recorded

Number of completed tool calls recorded so far.

write()[source]

Write the trace file atomically (a temporary file, then a rename).

Return type:

None

run(stdin=None, stdout=None, stderr=None)[source]

Relay a whole session and return the server’s exit code.

Reads client messages from stdin until it closes, relays server output to stdout until the server exits, and passes the server’s stderr through. The trace is written at the end even if no tool was called.

Parameters:
  • stdin (Optional[IO[str]]) – Client-to-server stream (default: sys.stdin).

  • stdout (Optional[IO[str]]) – Server-to-client stream (default: sys.stdout).

  • stderr (Optional[IO[str]]) – Where the server’s stderr goes (default: sys.stderr).

Return type:

int

Returns:

The server process exit code.

Server configuration

class toolscore.mcp.MCPServerSpec(name, command, env=<factory>)[source]

A single MCP server entry resolved from a config file.

Variables:
  • name – The key under mcpServers identifying this server.

  • command – The full command vector, i.e. [command, *args].

  • env – Environment variable overrides for the server process.

Parameters:
name: str
command: list[str]
env: dict[str, str]
__init__(name, command, env=<factory>)
Parameters:
Return type:

None

toolscore.mcp.load_mcp_config(path, server=None)[source]

Load an MCP server specification from a Claude Desktop style config file.

Parameters:
  • path (str | Path) – Path to the JSON configuration file.

  • server (str | None) – Name of the server to load. If None and exactly one server is defined, that server is used. If None and multiple servers exist, a ValueError is raised listing the available names.

Return type:

MCPServerSpec

Returns:

The resolved MCPServerSpec.

Raises:
  • FileNotFoundError – If path does not exist.

  • ValueError – If the file is malformed, the mcpServers key is missing or empty, the requested server is not found, or server is None while multiple servers are defined.

Scorecard harness

toolscore.mcp.generate_scenarios(tools, cases_per_tool=3, include_edge_cases=True)[source]

Generate test scenarios for every advertised tool.

For each tool with a usable object schema this produces up to cases_per_tool happy-path scenarios (well-formed arguments generated via generate_value_from_schema(), with all required properties present) and, when include_edge_cases is set and the schema permits, a handful of edge cases (missing required argument, wrong-type value, empty/zero values).

Tools whose schema is empty or invalid cannot be introspected, so they each receive a single no-arguments scenario whose description flags the problem.

Parameters:
  • tools (list[MCPToolDef]) – The tool definitions to plan scenarios for.

  • cases_per_tool (int) – Number of happy-path scenarios per tool.

  • include_edge_cases (bool) – Whether to also generate edge-case scenarios.

Return type:

list[Scenario]

Returns:

The planned scenarios across all tools.

toolscore.mcp.run_scenarios(client, scenarios)[source]

Execute every scenario against the server, never aborting on failure.

Each scenario is called in turn. Transport problems (MCPError) and timeouts (MCPTimeoutError) are caught per scenario and recorded as a failed result rather than propagated, so one misbehaving tool cannot sink the whole run.

The ok determination depends on the scenario kind:

  • "happy" scenarios are ok only if the call completed without an error result.

  • "edge" scenarios are ok whenever the server responded at all (an error payload is an acceptable, expected reaction to bad input); they fail only on a timeout or transport crash.

Parameters:
  • client (MCPStdioClient) – A started MCP client connected to the server under test.

  • scenarios (list[Scenario]) – The scenarios to execute.

Return type:

list[ScenarioResult]

Returns:

One ScenarioResult per input scenario, in order.

toolscore.mcp.lint_tools(tools, instructions=None)[source]

Statically lint tool schemas and the text a model reads.

Errors (real defects):

  • missing or empty inputSchema,

  • schema missing a type,

  • a declared property that lacks a type,

  • a name in required that is absent from properties.

Warnings (ergonomic nits):

  • missing or very short (< 10 chars) description,

  • a tool name that is not snake_case,

  • properties present but no required list declared.

Text checks over tool names, descriptions, parameter descriptions and the server instructions (see toolscore.mcp.lint_rules):

  • a reference to a tool the server does not expose (warning),

  • hidden Unicode tag characters or bidirectional controls (error),

  • tool-poisoning directives such as “do not tell the user” (error),

  • other invisible format characters (warning).

Each issue carries a concrete fix hint. Issues found in the instructions are reported under the tool name "<instructions>".

Parameters:
Return type:

list[LintIssue]

Returns:

Every issue found, across all tools.

toolscore.mcp.lint_rules.dangling_reference_issues(tools, instructions=None)[source]

Find text that tells the model to use a tool the server does not expose.

A name counts as a tool reference only in phrasing that names a tool: a quoted name followed by tool; a quoted name after use/call/invoke/run/try that is shaped like a tool name (snake_case or kebab-case, so fields such as endCursor and values such as add are not read as tools); or an unquoted snake_case name between use/call/invoke and tool. Names that exist as a tool, a parameter, or an allowed parameter value anywhere in the server are not reported (methods such as get_comments are often values of a method parameter).

Parameters:
  • tools (list[MCPToolDef]) – The server’s tools.

  • instructions (str | None) – The server’s instructions, if any.

Return type:

list[LintIssue]

Returns:

One warning per distinct missing name per owner (tool or <instructions>).

toolscore.mcp.lint_rules.poisoning_issues(tools, instructions=None)[source]

Find hidden text and concealment directives in what the model reads.

Errors:

  • Unicode tag characters (invisible text a model still reads),

  • bidirectional override/isolate controls (text that displays differently than it is stored),

  • directives to hide things from the user (“do not tell the user”, “do not mention this to the user”, “without telling the user”) or to ignore other instructions.

Warnings:

  • <IMPORTANT>-, <system>- or <instructions>-style blocks, which some servers use for emphasis and published attacks use to wrap hidden instructions;

  • other invisible format characters (zero-width characters, direction marks, BOM). They are normal in some scripts, so they are only surfaced.

Subdivision flag emoji, which are built from tag characters, are not reported.

Tool names are checked for hidden characters too.

Parameters:
  • tools (list[MCPToolDef]) – The server’s tools.

  • instructions (str | None) – The server’s instructions, if any.

Return type:

list[LintIssue]

Returns:

The issues found, at most one per kind per owner.

class toolscore.mcp.Scenario(tool, arguments, kind, description)[source]

A single planned tool invocation.

Variables:
  • tool – The name of the tool to call.

  • arguments – The arguments to pass to the tool.

  • kind – Either "happy" (well-formed, expected to succeed) or "edge" (intentionally malformed/boundary input where an error response from the server is acceptable).

  • description – Human-readable summary of what the scenario exercises.

Parameters:
tool: str
arguments: dict[str, Any]
kind: str
description: str
__init__(tool, arguments, kind, description)
Parameters:
Return type:

None

class toolscore.mcp.ScenarioResult(scenario, ok, is_error, duration, detail)[source]

The outcome of executing a single Scenario.

Variables:
  • scenario – The scenario that produced this result.

  • ok – Whether the scenario is considered a pass. For "happy" scenarios this means the call completed without is_error. For "edge" scenarios this means the server responded (with or without an error payload) rather than crashing or timing out – error responses to bad input are acceptable and expected.

  • is_error – Whether the tool result reported an error.

  • duration – Wall-clock seconds the call took.

  • detail – A short outcome string (e.g. an error message excerpt).

Parameters:
scenario: Scenario
ok: bool
is_error: bool
duration: float
detail: str
__init__(scenario, ok, is_error, duration, detail)
Parameters:
Return type:

None

class toolscore.mcp.LintIssue(tool, severity, message, fix='')[source]

A single problem found while statically linting a tool schema.

Variables:
  • tool – The name of the offending tool.

  • severity – Either "error" (a real schema defect) or "warning" (a quality nit that hurts agent ergonomics).

  • message – A human-readable description of the issue.

  • fix – A short, concrete suggestion for how to resolve the issue. May be empty for issues raised without a specific remedy.

Parameters:
tool: str
severity: str
message: str
fix: str = ''
__init__(tool, severity, message, fix='')
Parameters:
Return type:

None

Scorecard

class toolscore.mcp.MCPScorecard(server_info, tools, results, lint, instructions=None)[source]

An aggregated quality report for a single MCP server.

Variables:
  • server_info – The serverInfo dict reported during the handshake.

  • tools – The tools advertised by the server.

  • results – The executed scenario results.

  • lint – The lint issues found in the tool schemas and the text a model reads.

  • instructions – The server instructions from the handshake, if any. MCP clients send them to the model with every session, so they count toward the context cost.

Parameters:
server_info: dict[str, Any]
tools: list[MCPToolDef]
results: list[ScenarioResult]
lint: list[LintIssue]
instructions: str | None = None
property happy_pass_rate: float

Fraction of happy-path scenarios that passed (1.0 if none).

property edge_resilience_rate: float

Fraction of edge scenarios the server survived (1.0 if none).

property lint_error_count: int

Number of error-severity lint issues.

property lint_warning_count: int

Number of warning-severity lint issues.

property lint_score: float

Schema cleanliness score in [0, 1] (1.0 is spotless).

property score: float

The overall blended score in [0, 1].

property grade: str

The letter grade ("A" … "F") for score.

property total_tool_tokens: int

Estimated total context tokens consumed by all tool definitions.

Tool definitions are sent to the model on every request, so this is a proxy for how much of the context window the server’s tools occupy.

property instructions_tokens: int

Estimated context tokens of the server instructions (0 if none).

property context_tokens: int

tools plus instructions.

Type:

Estimated context the server costs on every request

__init__(server_info, tools, results, lint, instructions=None)
Parameters:
Return type:

None

toolscore.mcp.grade_meets(grade, threshold)[source]

Report whether grade is at least as good as threshold.

Parameters:
  • grade (str) – The achieved grade (case-insensitive).

  • threshold (str) – The minimum acceptable grade (case-insensitive).

Return type:

bool

Returns:

True if grade ranks the same as or better than threshold. Unknown grades are treated as the worst possible.

toolscore.mcp.print_scorecard(card, console=None)[source]

Render the scorecard to a rich console.

Prints a header panel (server name/version and the big letter grade), a per-tool table (scenarios passed/total and average latency), and a lint section listing any issues.

Parameters:
  • card (MCPScorecard) – The scorecard to render.

  • console (Console | None) – The console to print to; a fresh one is created if omitted.

Return type:

None

toolscore.mcp.scorecard_to_json(card)[source]

Render the scorecard as a JSON-serializable dict.

Parameters:

card (MCPScorecard) – The scorecard to render.

Return type:

dict[str, Any]

Returns:

A plain dict (lists/dicts/strings/numbers only) capturing the grade, component scores, per-tool summaries, every scenario result, and every lint issue.

toolscore.mcp.scorecard_to_markdown(card)[source]

Render the scorecard as Markdown for READMEs or PR comments.

Parameters:

card (MCPScorecard) – The scorecard to render.

Return type:

str

Returns:

A Markdown document with a grade line, component scores (including the tool-definition token cost), a per-tool summary table, and a ranked “Top issues to fix” list.