LLM-as-a-Judge
Toolscore’s core metrics are deterministic and syntactic — they compare tool
names and arguments exactly. Sometimes that is too strict: search vs
web_search, query vs q, or two phrasings of the same value should
count as equivalent. The optional LLM judge uses a language model to score
semantic equivalence, multiplexed across providers.
Note
The LLM judge is opt-in and non-deterministic by nature. Toolscore’s default gate is the deterministic metric suite; reach for the judge as an additional signal, not a replacement. See Toolscore vs. Other Tools for where this fits.
JudgeConfig
A single dataclass configures every provider:
from toolscore.metrics.llm_judge import JudgeConfig
JudgeConfig(
model="gpt-4o-mini", # model name; drives provider inference
provider=None, # None = infer from model / base_url
api_key=None, # None = read the per-provider env var
base_url=None, # set → OpenAI-compatible endpoint (Ollama/vLLM/...)
temperature=0.0, # 0.0 for determinism
max_retries=2, # passed to the underlying SDK
)
Anywhere a judge is accepted you can also pass a bare model-name string
(shorthand for JudgeConfig(model=...)) or None (the default config).
Provider inference
When provider is None it is resolved from the model name (or
base_url):
Condition |
Provider |
Example model |
|---|---|---|
|
|
|
model starts with |
|
|
model starts with |
|
|
anything else |
|
|
You can always set provider explicitly. Two combinations are rejected:
provider="openai_compatible" requires a base_url, and any other
explicit provider must not be combined with a base_url (a base URL
implies an OpenAI-compatible endpoint, so the pairing is ambiguous).
Environment-variable keys
When api_key is omitted, the key is read from a per-provider environment
variable:
Provider |
Environment variable |
|---|---|
|
|
|
|
|
|
|
|
For Gemini, GOOGLE_API_KEY is checked first and GEMINI_API_KEY second,
matching the newer google-genai SDK which accepts either.
Local / OpenAI-compatible endpoints (Ollama, vLLM, Groq)
Any server that speaks the OpenAI chat-completions API works through the
openai_compatible provider. Just set base_url:
from toolscore.metrics.llm_judge import JudgeConfig, calculate_semantic_correctness
# Ollama running locally — no real API key needed.
judge = JudgeConfig(model="llama3.1", base_url="http://localhost:11434/v1")
result = calculate_semantic_correctness(gold_calls, trace_calls, judge=judge)
print(result["semantic_score"])
OpenAI-compatible endpoints often need no real key; the SDK still requires a
non-empty value, so Toolscore supplies a "not-needed" placeholder when no key
is found and a base_url is set. This also means the judge does not raise a
“missing API key” error for openai_compatible — it lets the local server
decide.
Batching behavior
By default the judge issues a single batched request containing every gold/trace pair and asks for a JSON array of per-pair scores. This keeps cost and latency low.
Only the first
min(len(gold), len(trace))pairs are judged. A length mismatch additionally applies a proportional length penalty to the overallsemantic_score.If the batched response cannot be parsed into the expected number of pairs, the judge transparently falls back to one request per pair.
Transport, authentication, and other SDK errors are not retried per-pair — they propagate so a bad key or network failure surfaces clearly instead of fanning out into N doomed requests.
The return value is a dict with semantic_score, per_call_scores,
explanations, model_used, gold_count, and trace_count.
Extras install matrix
Each provider’s SDK is an optional extra, imported lazily so the module loads even with no SDK installed:
Provider |
Install |
SDK |
|---|---|---|
OpenAI / OpenAI-compatible |
|
|
Anthropic |
|
|
Gemini |
|
|
Everything |
|
all of the above |
If you call the judge without the required SDK, you get an ImportError whose
message includes the exact install command.
Using the judge from evaluate_trace
The file-based toolscore.evaluate_trace() accepts a judge= argument
that folds semantic correctness into the result:
from toolscore import evaluate_trace
from toolscore.metrics.llm_judge import JudgeConfig
# judge=True → default JudgeConfig() (gpt-4o-mini, OpenAI)
result = evaluate_trace("gold.json", "trace.json", judge=True)
# judge="claude-3-5-haiku-latest" → string shorthand, inferred as Anthropic
result = evaluate_trace("gold.json", "trace.json", judge="claude-3-5-haiku-latest")
# judge=JudgeConfig(...) → full control
result = evaluate_trace(
"gold.json", "trace.json",
judge=JudgeConfig(model="gemini-2.0-flash"),
)
judge=False (the default) skips the judge entirely. The in-memory
toolscore.evaluate() does not run the judge — call
calculate_semantic_correctness() directly when
working in memory.
Note
Earlier versions exposed use_llm_judge=... keyword arguments. Those are
gone — the single judge= argument (bool, string, or JudgeConfig) is
the supported surface.
CLI flags
The toolscore eval command exposes the judge:
# OpenAI (default model gpt-4o-mini, reads OPENAI_API_KEY):
toolscore eval gold.json trace.json --llm-judge
# Pick a model; the provider is inferred from the name:
toolscore eval gold.json trace.json --llm-judge --llm-model claude-3-5-haiku-latest
toolscore eval gold.json trace.json --llm-judge --llm-model gemini-2.0-flash
# Force a provider:
toolscore eval gold.json trace.json --llm-judge --llm-provider anthropic --llm-model ...
# Local OpenAI-compatible server (Ollama):
toolscore eval gold.json trace.json --llm-judge \
--llm-model llama3.1 --llm-base-url http://localhost:11434/v1
Flag |
Meaning |
|---|---|
|
Enable the semantic judge (off by default). |
|
Judge model (default |
|
Force |
|
Custom OpenAI-compatible endpoint; forces |
API reference
The calculate_semantic_correctness and calculate_batch_semantic_correctness
functions are autodocumented under Metrics Module.
- class toolscore.metrics.llm_judge.JudgeConfig(model='gpt-4o-mini', provider=None, api_key=None, base_url=None, temperature=0.0, max_retries=2)[source]
Configuration for the LLM judge.
- Parameters:
model (
str) – Model name (e.g.gpt-4o-mini,claude-3-5-haiku,gemini-2.0-flash,llama3.1).provider (
Optional[Literal['openai','anthropic','gemini','openai_compatible']]) – Provider to use.Noneinfers it frommodel/base_url(see module docstring).api_key (
str|None) – API key.Nonefalls back to the per-provider env var (OPENAI_API_KEY/ANTHROPIC_API_KEY/GOOGLE_API_KEY).base_url (
str|None) – Custom endpoint. When set, the provider becomesopenai_compatible(covers Ollama/vLLM/Groq/etc.).temperature (
float) – Sampling temperature (default0.0for determinism).max_retries (
int) – Max retries passed to the underlying SDK client.
- __init__(model='gpt-4o-mini', provider=None, api_key=None, base_url=None, temperature=0.0, max_retries=2)
- toolscore.metrics.llm_judge.infer_provider(config)[source]
Resolve the effective provider for a config.
Resolution rules:
An explicit
providerwins. It is validated againstbase_url:openai_compatiblerequires abase_url; any other explicit provider must not be combined with abase_url(the combination is ambiguous).With no explicit provider, a
base_urlforcesopenai_compatible.Otherwise the provider is inferred from the model name.
- Parameters:
config (
JudgeConfig) – Judge configuration.- Return type:
Literal['openai','anthropic','gemini','openai_compatible']- Returns:
The resolved provider.
- Raises:
ValueError – If an explicit
providerconflicts withbase_url.