Fluent API — expect()

The expect() function is the easiest way to assert on an agent’s tool-calling behavior. It returns an Expectation builder that you configure with a fluent chain of method calls, then execute with run() (or run_async() for async agents).

Hero Example

from toolscore import expect, ANY, Regex

expect(agent).on("book me a flight to NYC") \
    .calls("search_flights", origin=ANY, destination="NYC") \
    .then_calls("book_flight", flight_id=Regex(r"FL-\d+")) \
    .does_not_call("cancel_booking") \
    .with_score(0.9) \
    .run()

If the assertion fails you get a rich diff table printed inline — no manual digging through raw response objects required.

Quick Start

Assert on an already-produced result (list of call dicts)

actual = [
    {"tool": "search_flights", "args": {"origin": "JFK", "destination": "NYC"}},
    {"tool": "book_flight",    "args": {"flight_id": "FL-456"}},
]

expect(actual) \
    .calls("search_flights", origin="JFK", destination="NYC") \
    .then_calls("book_flight", flight_id=Regex(r"FL-\d+")) \
    .with_score(0.9) \
    .run()

Assert on a sync agent callable

def my_agent(prompt: str) -> list[dict]:
    ...  # calls your LLM, returns tool-call dicts

expect(my_agent) \
    .on("find flights from JFK to LAX") \
    .calls("search_flights", origin="JFK", destination="LAX") \
    .run()

Assert on a raw provider response (auto-extracted)

expect() automatically calls auto_extract() so you can pass a raw OpenAI, Anthropic, or Gemini response dict directly:

raw_response = openai_client.chat.completions.create(...)  # returns dict-like
expect(raw_response) \
    .calls("search_flights", origin=ANY, destination="NYC") \
    .run()

Chain-Method Reference

Method

Description

.on(prompt)

Set the input prompt for a callable subject. Required when the subject is an agent function; forbidden when the subject is a result/list.

.calls(tool, **args)

Append an expected tool call. **args may contain matcher objects such as ANY or Regex. Calling with no kwargs means “match the tool name but do not check arguments”.

.then_calls(tool, **args)

Alias for calls() — reads more naturally when describing sequences.

.does_not_call(tool, **args)

Assert that tool must not be called. With keyword arguments, only calls whose arguments match them (values or Argument Matchers) are forbidden, e.g. does_not_call("run_shell", command=Contains("rm -rf")). Can be combined with calls() or used alone (forbidden-only contract).

.with_score(min_score)

Set the minimum composite score (default: 0.9). Raises ToolScoreAssertionError with a diff table when not met.

.with_weights(**weights)

Override composite-score weights. Valid keys: selection_accuracy, argument_f1, sequence_accuracy, redundant_rate, required_call_recall (0 by default; weight it to penalize required calls that were skipped or failed).

.with_strict_args()

Enable strict argument comparison: no int/float coercion, no string strip.

.run()

Execute the assertion synchronously. Returns EvaluationResult on success.

.run_async()

Execute the assertion asynchronously. Works for both sync and async agent callables. Returns EvaluationResult on success.

Using Matchers

Import matchers from toolscore:

from toolscore import ANY, Approx, Contains, IsType, OneOf, Regex

Place them as argument values inside calls():

expect(actual) \
    .calls("book_flight",
           flight_id=Regex(r"FL-\d+"),  # regex match
           seats=Approx(2),             # numeric closeness
           origin=ANY,                  # matches anything
           cabin=OneOf("economy", "business")) \
    .run()

See Core Module for the full matcher reference.

Async Agents

Use run_async() for async agents:

import asyncio
from toolscore import expect, ANY

async def async_agent(prompt: str) -> list[dict]:
    ...  # async LLM call

async def test_booking():
    result = await expect(async_agent) \
        .on("book me a flight") \
        .calls("search_flights", destination=ANY) \
        .run_async()
    assert result.score >= 0.9

# or with asyncio.run in a plain test:
asyncio.run(test_booking())

Calling .run() on an async agent function raises TypeError with a message directing you to use run_async() instead.

Forbidden-Only Contract

You may use does_not_call() without any calls() declarations. In this mode the evaluation runs but the minimum-score threshold is not enforced — only the forbidden-call check is applied:

# Assert the agent never calls the dangerous tool, regardless of what else it does
expect(actual).does_not_call("delete_all_files").run()

Forbid only some calls to a tool by passing argument conditions. A safe run_shell call passes; a destructive one fails the assertion and the message lists the offending call:

from toolscore import Contains, expect

expect(actual).does_not_call("run_shell", command=Contains("rm -rf")).run()

Contains finds the text anywhere, including on a second line. Regex matches the whole string and . does not cross newlines, so for forbidden calls prefer Contains or Regex(pattern, re.DOTALL). For policies over whole traces, files and CI, see Behavior and Safety Checks.

API Reference

toolscore.expect.expect(subject)[source]

Create a fluent Expectation for subject.

subject may be:

  • An agent callable (sync or async) — pair with .on(prompt) then call .run() or .run_async().

  • An already-produced result — a raw OpenAI/Anthropic/Gemini response, a LangGraph state, or a list of call dicts. Do not call .on().

Parameters:

subject (Any) – An agent callable or a raw LLM response / list of call dicts.

Return type:

Expectation

Returns:

An Expectation builder.

Example:

from toolscore import expect, ANY, Regex

expect(agent).on("book me a flight to NYC") \
    .calls("search_flights", origin=ANY, destination="NYC") \
    .then_calls("book_flight", flight_id=Regex(r"FL-\d+")) \
    .does_not_call("cancel_booking") \
    .with_score(0.9) \
    .run()
class toolscore.expect.Expectation(subject)[source]

Fluent builder for tool-call assertions.

Do not instantiate directly — use expect() instead.

Parameters:

subject (Any)

__init__(subject)[source]
Parameters:

subject (Any)

Return type:

None

on(prompt)[source]

Set the input prompt for a callable subject.

Parameters:

prompt (str) – The string to pass to the agent callable.

Return type:

Expectation

Returns:

self for chaining.

calls(tool, **args)[source]

Append an expected tool call.

Calling with no keyword arguments — calls("tool") — means do not check arguments: the tool name must be called, but whatever arguments the agent passed are accepted (equivalent to using toolscore.ANY for every argument). This is the common case and keeps casual assertions from failing just because the agent supplied arguments.

Calling with keyword arguments — calls("tool", q="x") — checks those arguments. Use toolscore.ANY, toolscore.Regex, etc. as values for individual arguments when you want flexible matching.

Note

There is intentionally no fluent way to assert “the tool was called with exactly zero arguments”, because calls("tool", **{}) is indistinguishable in Python from calls("tool"). For that rare expectation, use toolscore.evaluate() directly with an explicit empty dict:

evaluate([{"tool": "t", "args": {}}], actual)
Parameters:
  • tool (str) – Expected tool name.

  • **args (Any) – Expected argument key/value pairs (may contain Matcher instances). Omit entirely to skip argument checking.

Return type:

Expectation

Returns:

self for chaining.

then_calls(tool, **args)[source]

Alias for calls(); reads naturally in sequences.

Parameters:
  • tool (str) – Expected tool name.

  • **args (Any) – Expected argument key/value pairs.

Return type:

Expectation

Returns:

self for chaining.

does_not_call(tool, **args)[source]

Assert that the agent must NOT call tool (optionally: with these arguments).

Without args any call to tool fails the expectation. With args only a call whose listed arguments all match does (values or matchers), for example does_not_call("run_shell", command=Regex(r".*\brm\s+-rf\b.*")).

Parameters:
  • tool (str) – Tool name that must not be called.

  • **args (Any) – Optional argument values or matchers that make a call forbidden.

Return type:

Expectation

Returns:

self for chaining.

with_score(min_score)[source]

Set the minimum composite score required (default: 0.9).

Parameters:

min_score (float) – Float in [0.0, 1.0].

Return type:

Expectation

Returns:

self for chaining.

with_weights(**weights)[source]

Override composite-score weights.

Valid keys: selection_accuracy, argument_f1, sequence_accuracy, redundant_rate, required_call_recall (0 by default; weight it to penalize skipped or failed required calls).

Parameters:

**weights (float) – Weight key/value pairs.

Return type:

Expectation

Returns:

self for chaining.

with_strict_args()[source]

Enable strict argument comparison (no int/float coercion, no string strip).

Return type:

Expectation

Returns:

self for chaining.

run()[source]

Execute the assertion synchronously.

Raises:
  • ValueError – If no expectations are declared, or prompt is missing for a callable subject, or prompt is set for a non-callable subject.

  • TypeError – If the subject is an async function (use run_async()).

  • ToolScoreAssertionError – If a forbidden tool was called, or the composite score is below with_score().

Return type:

EvaluationResult

Returns:

EvaluationResult on success.

async run_async()[source]

Execute the assertion asynchronously.

Works for both sync and async agent callables, as well as already-produced results.

Raises:
  • ValueError – If no expectations are declared, or prompt is missing for a callable subject, or prompt is set for a non-callable subject.

  • ToolScoreAssertionError – If a forbidden tool was called, or the composite score is below with_score().

Return type:

EvaluationResult

Returns:

EvaluationResult on success.