Fluent API — expect()
The expect() function is the easiest way to assert on an agent’s tool-calling
behavior. It returns an Expectation builder that you
configure with a fluent chain of method calls, then execute with
run() (or
run_async() for async agents).
Hero Example
from toolscore import expect, ANY, Regex
expect(agent).on("book me a flight to NYC") \
.calls("search_flights", origin=ANY, destination="NYC") \
.then_calls("book_flight", flight_id=Regex(r"FL-\d+")) \
.does_not_call("cancel_booking") \
.with_score(0.9) \
.run()
If the assertion fails you get a rich diff table printed inline — no manual digging through raw response objects required.
Quick Start
Assert on an already-produced result (list of call dicts)
actual = [
{"tool": "search_flights", "args": {"origin": "JFK", "destination": "NYC"}},
{"tool": "book_flight", "args": {"flight_id": "FL-456"}},
]
expect(actual) \
.calls("search_flights", origin="JFK", destination="NYC") \
.then_calls("book_flight", flight_id=Regex(r"FL-\d+")) \
.with_score(0.9) \
.run()
Assert on a sync agent callable
def my_agent(prompt: str) -> list[dict]:
... # calls your LLM, returns tool-call dicts
expect(my_agent) \
.on("find flights from JFK to LAX") \
.calls("search_flights", origin="JFK", destination="LAX") \
.run()
Assert on a raw provider response (auto-extracted)
expect() automatically calls auto_extract() so
you can pass a raw OpenAI, Anthropic, or Gemini response dict directly:
raw_response = openai_client.chat.completions.create(...) # returns dict-like
expect(raw_response) \
.calls("search_flights", origin=ANY, destination="NYC") \
.run()
Chain-Method Reference
Method |
Description |
|---|---|
|
Set the input prompt for a callable subject. Required when the subject is an agent function; forbidden when the subject is a result/list. |
|
Append an expected tool call. |
|
Alias for |
|
Assert that tool must not be called. With keyword arguments, only
calls whose arguments match them (values or Argument Matchers) are
forbidden, e.g. |
|
Set the minimum composite score (default: 0.9). Raises
|
|
Override composite-score weights. Valid keys: |
|
Enable strict argument comparison: no int/float coercion, no string strip. |
|
Execute the assertion synchronously. Returns
|
|
Execute the assertion asynchronously. Works for both sync and async agent
callables. Returns |
Using Matchers
Import matchers from toolscore:
from toolscore import ANY, Approx, Contains, IsType, OneOf, Regex
Place them as argument values inside calls():
expect(actual) \
.calls("book_flight",
flight_id=Regex(r"FL-\d+"), # regex match
seats=Approx(2), # numeric closeness
origin=ANY, # matches anything
cabin=OneOf("economy", "business")) \
.run()
See Core Module for the full matcher reference.
Async Agents
Use run_async() for async agents:
import asyncio
from toolscore import expect, ANY
async def async_agent(prompt: str) -> list[dict]:
... # async LLM call
async def test_booking():
result = await expect(async_agent) \
.on("book me a flight") \
.calls("search_flights", destination=ANY) \
.run_async()
assert result.score >= 0.9
# or with asyncio.run in a plain test:
asyncio.run(test_booking())
Calling .run() on an async agent function raises TypeError with a
message directing you to use run_async() instead.
Forbidden-Only Contract
You may use does_not_call() without any
calls() declarations. In this mode the
evaluation runs but the minimum-score threshold is not enforced — only the
forbidden-call check is applied:
# Assert the agent never calls the dangerous tool, regardless of what else it does
expect(actual).does_not_call("delete_all_files").run()
Forbid only some calls to a tool by passing argument conditions. A safe
run_shell call passes; a destructive one fails the assertion and the
message lists the offending call:
from toolscore import Contains, expect
expect(actual).does_not_call("run_shell", command=Contains("rm -rf")).run()
Contains finds the text anywhere, including on a second line. Regex
matches the whole string and . does not cross newlines, so for forbidden
calls prefer Contains or Regex(pattern, re.DOTALL). For policies over whole traces, files and CI, see
Behavior and Safety Checks.
API Reference
- toolscore.expect.expect(subject)[source]
Create a fluent
Expectationfor subject.subject may be:
An agent callable (sync or async) — pair with
.on(prompt)then call.run()or.run_async().An already-produced result — a raw OpenAI/Anthropic/Gemini response, a LangGraph state, or a list of call dicts. Do not call
.on().
- Parameters:
subject (
Any) – An agent callable or a raw LLM response / list of call dicts.- Return type:
- Returns:
An
Expectationbuilder.
Example:
from toolscore import expect, ANY, Regex expect(agent).on("book me a flight to NYC") \ .calls("search_flights", origin=ANY, destination="NYC") \ .then_calls("book_flight", flight_id=Regex(r"FL-\d+")) \ .does_not_call("cancel_booking") \ .with_score(0.9) \ .run()
- class toolscore.expect.Expectation(subject)[source]
Fluent builder for tool-call assertions.
Do not instantiate directly — use
expect()instead.- Parameters:
subject (Any)
- on(prompt)[source]
Set the input prompt for a callable subject.
- Parameters:
prompt (
str) – The string to pass to the agent callable.- Return type:
- Returns:
selffor chaining.
- calls(tool, **args)[source]
Append an expected tool call.
Calling with no keyword arguments —
calls("tool")— means do not check arguments: the tool name must be called, but whatever arguments the agent passed are accepted (equivalent to usingtoolscore.ANYfor every argument). This is the common case and keeps casual assertions from failing just because the agent supplied arguments.Calling with keyword arguments —
calls("tool", q="x")— checks those arguments. Usetoolscore.ANY,toolscore.Regex, etc. as values for individual arguments when you want flexible matching.Note
There is intentionally no fluent way to assert “the tool was called with exactly zero arguments”, because
calls("tool", **{})is indistinguishable in Python fromcalls("tool"). For that rare expectation, usetoolscore.evaluate()directly with an explicit empty dict:evaluate([{"tool": "t", "args": {}}], actual)
- Parameters:
- Return type:
- Returns:
selffor chaining.
- then_calls(tool, **args)[source]
Alias for
calls(); reads naturally in sequences.- Parameters:
- Return type:
- Returns:
selffor chaining.
- does_not_call(tool, **args)[source]
Assert that the agent must NOT call tool (optionally: with these arguments).
Without
argsany call to tool fails the expectation. Withargsonly a call whose listed arguments all match does (values or matchers), for exampledoes_not_call("run_shell", command=Regex(r".*\brm\s+-rf\b.*")).- Parameters:
- Return type:
- Returns:
selffor chaining.
- with_score(min_score)[source]
Set the minimum composite score required (default: 0.9).
- Parameters:
min_score (
float) – Float in[0.0, 1.0].- Return type:
- Returns:
selffor chaining.
- with_weights(**weights)[source]
Override composite-score weights.
Valid keys:
selection_accuracy,argument_f1,sequence_accuracy,redundant_rate,required_call_recall(0 by default; weight it to penalize skipped or failed required calls).- Parameters:
**weights (
float) – Weight key/value pairs.- Return type:
- Returns:
selffor chaining.
- with_strict_args()[source]
Enable strict argument comparison (no int/float coercion, no string strip).
- Return type:
- Returns:
selffor chaining.
- run()[source]
Execute the assertion synchronously.
- Raises:
ValueError – If no expectations are declared, or prompt is missing for a callable subject, or prompt is set for a non-callable subject.
TypeError – If the subject is an async function (use
run_async()).ToolScoreAssertionError – If a forbidden tool was called, or the composite score is below
with_score().
- Return type:
- Returns:
EvaluationResulton success.
- async run_async()[source]
Execute the assertion asynchronously.
Works for both sync and async agent callables, as well as already-produced results.
- Raises:
ValueError – If no expectations are declared, or prompt is missing for a callable subject, or prompt is set for a non-callable subject.
ToolScoreAssertionError – If a forbidden tool was called, or the composite score is below
with_score().
- Return type:
- Returns:
EvaluationResulton success.