eval_tool_call_accuracy
Check whether an agent called the expected tool or sequence by comparing expected versus actual calls, with optional argument, order, and unexpected-call penalization.
Instructions
Evaluate whether an agent called the expected tool or tool sequence.
Pure deterministic — no LLM judge needed. Two compatible modes:
Single-call mode compares
expected_tool/actual_tooland optional argument dictionaries exactly.Trace mode consumes
expected_tool_callsplus the canonicalagent_tracereturned byeval_ingest_trace. It can require order and optionally penalize unexpected calls.
Args:
expected_tool: Single tool name the agent should have called.
actual_tool: Single tool name the agent actually called.
expected_arguments: Expected arguments for single-call mode.
actual_arguments: Actual arguments for single-call mode.
expected_tool_calls: Expected names for trace mode. An empty list
explicitly asserts that the agent should call no tools.
agent_trace: Canonical step dictionaries returned by
eval_ingest_trace.
require_order: In trace mode, require expected names in order.
penalize_unexpected: In trace mode, lower the score for calls not
present in expected_tool_calls.
Returns:
{"score": float, "passed": bool, "reason": str, "evaluator": "tool_call_accuracy"}, or an error dict when
the arguments do not form either mode.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| actual_tool | No | ||
| agent_trace | No | ||
| expected_tool | No | ||
| require_order | No | ||
| actual_arguments | No | ||
| expected_arguments | No | ||
| expected_tool_calls | No | ||
| penalize_unexpected | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||