Agentguard47
This server provides read-only access to hosted AgentGuard data—querying traces, alerts, usage, costs, and budget health without modifying anything or enforcing guards.
Query trace summaries with search, filtering, and pagination (
query_traces)Retrieve the full event tree of a specific trace (
get_trace)Get normalized decision events from a trace (
get_trace_decisions)List recent guard alerts and errors (
get_alerts)Check hosted event quota usage (
get_usage)View current-month hosted cost breakdown (
get_costs)Perform a read-only budget/quota health check (
check_budget)
Provides runtime guards for CrewAI agents, including loop, retry, and budget limits via optional integration extra.
Integrates runtime guards with LangChain agents to enforce budgets, loops, and retry limits via optional extra.
Adds runtime control to LangGraph agents, including hard budget caps and loop detection via optional integration extra.
Auto-patches OpenAI SDK to trace and enforce budget limits on chat completions, stopping overspend in-process.
AgentGuard
Stop runaway agents with runtime checks in Python.
AgentGuard checks budgets, repeated tool calls, retries, and elapsed time in instrumented Python code. Guards raise exceptions so your application can stop the next operation. The base SDK has no runtime dependencies and needs no account.
Names: this repository is agent47, the PyPI package is agentguard47,
and the Python import is agentguard. Requires Python 3.9 or newer.
Getting started
Install in a virtual environment, then run the offline checks:
python -m pip install agentguard47
agentguard doctor
agentguard demodoctor checks the installation and local trace writing. demo exercises
budget, loop, and retry stops without provider keys or network access. Follow
the trace path printed by the command to inspect its output.
agentguard demo --feedback prints a local redacted report; nothing is sent.
agentguard receipt agentguard_demo_traces.jsonl prints a receipt of each stop
with the trace's SHA-256 drawn as a barcode. Add --format markdown to paste it
into a PR or issue. The hash identifies the trace file; it is not a signature.
Guard a Claude Code session
agentguard hook claude-code --install --writeThis installs a Claude Code hook that refuses the third identical tool call in
a row and a call that already failed twice. Refusals go to
.agentguard/claude-code/trace.jsonl. It checks tool calls, not tokens or
subscription quota. See the Claude Code hook guide.
Guard a script without editing it
agentguard run --budget-usd 5 agent.pyThis patches the OpenAI and Anthropic clients, then runs agent.py in the same
interpreter. Settings come from flags, then environment variables, then
.agentguard.json. A guard stop exits 1. Every run ends with the trace path on
stderr, ready for agentguard receipt. agentguard run python -m mypkg works too. The bounds are
the same as patching the client yourself; see
enforcement boundary.
Stop before a third call
Save this as budget_demo.py and run python budget_demo.py. It makes no
network requests.
from agentguard import BudgetExceeded, BudgetGuard
budget = BudgetGuard(max_calls=2)
completed = 0
for _ in range(3):
try:
budget.check() # Check before the operation.
# Put your provider or tool call here.
completed += 1
budget.consume(calls=1) # Record the completed operation.
except BudgetExceeded:
print(f"Stopped before call {completed + 1}")
assert completed == 2Expected output: Stopped before call 3.
Connect a provider
Install the provider's client separately. For OpenAI:
python -m pip install openaifrom agentguard import BudgetGuard, JsonlFileSink, Tracer, patch_openai
budget = BudgetGuard(max_cost_usd=5.00)
tracer = Tracer(
service="my-agent",
sink=JsonlFileSink(".agentguard/traces.jsonl"),
)
patch_openai(tracer, budget_guard=budget)
# Make your OpenAI chat.completions.create or responses.create calls after this setup.The patch checks recorded usage before dispatch and records response usage
afterward, including streamed calls once the final usage arrives. A response
can exceed the remaining cost or token allowance. Concurrent requests do not
reserve capacity. Chat Completions streams request include_usage unless the
caller already set it.
The OpenAI Agents SDK runs on responses.create, so agentguard.init() before
the Runner puts every model call under the budget
(example). Hosted tools and
background=True responses are not covered. See the getting started guide
for setup, traces, and framework starters.
Related MCP server: Langfuse MCP Server
How enforcement works
flowchart TD
accTitle: AgentGuard operation checks
accDescr: Check a limit before an operation, then record usage.
A[Instrumented operation] --> B{Guard check}
B -->|Limit reached| C[Raise exception]
B -->|Allowed| D[Run operation]
D --> E[Record usage and trace]
E --> AText equivalent: check before an operation, run it if allowed, then record usage. A guard exception returns control to your application's error handler.
Guard | Checks | Raises |
| Recorded calls, tokens, or estimated cost |
|
| Repeated tool calls |
|
| Tool frequency and alternating patterns |
|
| Retries per tool |
|
| Elapsed time when checked |
|
| Calls within a sliding minute |
|
| Payment amounts before the payment callback |
|
For task budgets, use BudgetGuard.goal(...). For signatures and defaults,
read the guard source and
public exports.
Limits and security
Guards cover operations you instrument. Installing the package does not intercept every action in Cursor, Claude Code, or another agent.
A guard is not a sandbox or permission system. A permitted operation can still be destructive.
Timeout checks do not interrupt an already blocked function or cancel an agent running on a provider's server.
Cost estimates are not invoices. Supply reported cost or use strict cost resolution when an estimate is insufficient.
Recorded-budget preflight refuses the next instrumented call when stored usage is already at a cap. It does not reserve concurrent in-flight requests, predict the next response, or cap a provider subscription. See the enforcement boundary.
The base SDK uses the standard library. Optional framework extras install third-party dependencies and need their own security review.
The optional
[crewai]extra pulls ChromaDB. The 2026-09-12 audit found four unresolved advisories, including PYSEC-2026-311 / CVE-2026-45829. Review that exposure before installing the extra. Base SDK installs do not include ChromaDB.Trace content can contain application data. Review it before sharing or configuring a remote sink.
See security reporting, the dated dependency audit, and release notes. Audit results describe their recorded date, not a permanent clean bill of health.
Local traces and optional hosted ingest
The SDK is the free local proof path. Start local. Add hosted ingest only when you need retained history, alerts, team visibility, spend trends, hosted decision history, or dashboard-managed remote kill signals.
Local guards remain authoritative. HttpSink mirrors trace and decision events;
it does not execute remote kill signals by itself. See the
dashboard contract before configuring it.
Local use has no hosted event quota, retention period, or API-key allocation.
Network egress requires an integration you configure, such as HttpSink or
an OpenTelemetry exporter.
Nothing in the local SDK phones home. The AgentGuard website describes the optional hosted service.
Documentation
You want to | Start here |
See which paths actually stop a call | |
Install and trace a first run | |
Find guides and source references | |
Try a runnable example | |
Connect LangChain, LangGraph, or CrewAI | |
Inspect hosted data through MCP | |
Use local budget tools through MCP | |
Navigate with an AI assistant | |
Contribute a fix | |
Check what changed |
Help and maintenance
Maintained by Patrick Hughes. Report a bug with the package version, a minimal reproduction, and the expected result. Report vulnerabilities through SECURITY.md.
The source metadata defines the branch version. The PyPI badge links to the published version. Documentation examples and local links are tested in CI. The PyPI README is generated from this README and the changelog.
Available Tools
7 toolscheck_budgetARead-onlyIdempotent
Read-only hosted event-quota health check. Combines dashboard event quota and recorded cost summaries. This is not SDK BudgetGuard, not a provider invoice cap, and it does not refuse the next model call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds useful context beyond those annotations, notably that it does not refuse the next model call and that it combines two data sources. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. It front-loads the core purpose and then efficiently adds exclusions that prevent misuse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument read-only health check, the description provides enough context for correct selection and invocation. It lacks an explicit description of the return format, but this is a minor gap given the tool's simplicity and the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantic burden on the description. The schema coverage is effectively complete, and no parameter explanation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read-only health check for hosted event quota and states it combines quota and cost summary data. It is specific about what it does, though it does not explicitly distinguish itself from sibling tools like get_usage or get_costs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking hosted event quota health and includes negative guidance (not SDK BudgetGuard, not a provider invoice cap). However, it does not explicitly state when to use this tool over alternatives or name sibling tools as fallbacks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_alertsARead-onlyIdempotent
Read-only recent guard alerts (loop detection, budget exceeded) and errors from the hosted Read API. This reports stored alerts; it does not stop a running agent.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max alerts to return (default 50) | |
| since | No | ISO timestamp — only alerts after this time |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, idempotent, and non-destructive. The description adds the valuable clarification that it only reports stored alerts and never stops a running agent, which is not fully captured by the annotation hints. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundancy: the first states purpose and scope, the second clarifies a key side-effect. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with two optional, fully documented parameters and safety annotations already covering behavior, the description is largely complete. It could hint at ordering or default behavior, but nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes both limit and since with their defaults and meanings, so schema coverage is 100%. The description itself adds no parameter-level detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves recent guard alerts and errors from the hosted Read API in a read-only manner. It distinguishes the resource (alerts) from siblings like traces, usage, costs, and budget, though it does not explicitly name an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes it reports stored alerts and does not stop a running agent, giving one useful exclusion. However, it does not explicitly explain when to choose this over related tools such as check_budget or get_trace_decisions, leaving usage implications to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_costsARead-onlyIdempotent
Read-only hosted cost breakdown for the current month from the AgentGuard Read API. Estimated savings are recorded guard events, not an invoice credit.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds useful context about estimated savings being recorded guard events, not an invoice credit, which helps set expectations about the data's meaning. It doesn't mention pagination or response format, but that's acceptable for a zero-parameter tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, each earning its place: the first states the scope and source, the second clarifies an important caveat about the data's meaning. It's concise and front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with rich annotations, the description is largely complete. It clarifies the nature of the cost breakdown and the caveat about estimated savings. The only minor gap is not detailing the output structure, but since there's no output schema, it could be more explicit, but the tool is simple enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since there are 0 parameters, the schema provides full coverage at 100%, and the description doesn't need to explain any parameters. A baseline of 4 is appropriate because there is nothing to clarify, and the description correctly focuses on the output's semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (get) and resource (hosted cost breakdown for the current month), and clarifies it is read-only. It does not explicitly differentiate from siblings like get_usage or check_budget, but the resource is distinct enough for an agent to understand its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for viewing current month's costs but provides no explicit guidance on when to use this tool versus siblings like get_usage or check_budget. It mentions the source (AgentGuard Read API) but doesn't state exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_traceARead-onlyIdempotent
Read-only full event tree for one retained trace. Shows spans, tool calls, LLM calls, guard triggers, and errors. Does not enforce BudgetGuard or intercept provider calls.
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | The trace ID to look up |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, open-world, and non-destructive behavior, so the description adds value beyond them by enumerating the event types returned and explicitly negating BudgetGuard enforcement and provider-call interception. No behavioral surprises remain undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences lead with the read-only scope and core output, then list return contents and end with an important non-behavior. There is no filler or repetition of annotation metadata.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with rich annotations, the description covers what the tool returns and what side effects it avoids. The absence of an output schema is acceptable because the description enumerates the event categories the returned tree contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the only parameter, trace_id, is already described as 'The trace ID to look up'. The description reinforces that it targets a single retained trace but does not add format, source, or lookup semantics beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (get) and resource (one retained trace), and describes the output as a 'full event tree' that includes spans, tool calls, LLM calls, guard triggers, and errors. This clearly differentiates it from the sibling tools, which target listings, decisions, alerts, usage, costs, or budget checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use this for read-only inspection of a single retained trace and the full event tree. It also states exclusions ('does not enforce BudgetGuard or intercept provider calls'), helping an agent avoid expecting side effects, though it does not explicitly name sibling alternatives or when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trace_decisionsARead-onlyIdempotent
Read-only normalized decision.* events from one retained trace. Use this when a workflow includes proposal, override, approval, or binding steps. Does not approve, bind, or enforce those decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | The trace ID to inspect for decision events |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true. The description adds value by stating 'Read-only' at the beginning and explicitly saying it does not approve, bind, or enforce decisions, which reinforces the non-mutating nature beyond the annotation flags. This context about scope (retained trace, normalized events) is useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with zero fluff. It front-loads the core purpose, then gives usage guidance, then clarifies a potential misassumption. Every sentence earns its place, and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with rich annotations and no output schema, the description covers the main points: what it returns (decision.* events), the scope (one retained trace), and what it does not do (approve/bind/enforce). It is nearly complete, though it does not elaborate on what 'normalized' means or whether there are limits on event counts, which is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter trace_id already has a clear description ('The trace ID to inspect for decision events'). The tool description adds only marginal detail about the parameter ('from one retained trace'), which hints at a constraint but does not significantly expand the schema's meaning. A baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read-only normalized decision.* events from one retained trace.' This clearly distinguishes it from sibling tools like get_trace or query_traces, and the mention of proposal, override, approval, and binding steps further scopes what kind of decisions are covered.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage context: 'Use this when a workflow includes proposal, override, approval, or binding steps.' It also states what the tool does not do ('Does not approve, bind, or enforce those decisions'), which gives a when-not. However, it does not name specific alternative sibling tools, so it falls slightly short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageARead-onlyIdempotent
Read-only hosted event quota usage and plan limits from the AgentGuard Read API. This is dashboard event quota, not SDK BudgetGuard and not a provider invoice cap.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds useful context beyond annotations by clarifying the data source and scope (dashboard event quota vs SDK BudgetGuard vs provider invoice cap), which helps the agent understand exactly what this read operation covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundancy. The primary purpose is front-loaded, and the second sentence earns its place by disambiguating from related concepts. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with rich annotations, the description is complete. It states the resource, the source API, and the return scope (quota usage and plan limits), and clarifies what it is not, so an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is no parameter documentation burden on the description. The description appropriately focuses on what the tool returns rather than parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read-only'), a clear resource ('hosted event quota usage and plan limits'), and the source ('AgentGuard Read API'). It also explicitly distinguishes this from SDK BudgetGuard and provider invoice caps, which separates it from sibling tools like check_budget and get_costs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context that this tool is for dashboard event quota, not SDK BudgetGuard or provider invoice cap. It gives when-not guidance but does not explicitly name alternative sibling tools to use in those cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_tracesARead-onlyIdempotent
Read-only search for retained AgentGuard trace summaries from the AgentGuard Read API. Requires AGENTGUARD_API_KEY with read access; create keys in the AgentGuard dashboard. Returns JSON with a traces array, newest traces first when the API supports ordering; items include trace_id, service, root_name, event_count, error_count, duration_ms, started_at, API key metadata, and total_cost when available. Defaults to a small page, accepts offset pagination, exact service filtering, and ISO 8601 since/until bounds. Use this to find candidate trace_id values; use get_trace for the full event tree of one trace or get_trace_decisions for decision.* events from a known trace.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum trace summaries to return. Defaults to 20; API maximum is 500. | |
| since | No | ISO 8601 timestamp; include only traces that started at or after this time. | |
| until | No | ISO 8601 timestamp; include only traces that started at or before this time. | |
| offset | No | Zero-based pagination offset for walking additional trace pages. | |
| service | No | Exact AgentGuard service name to filter by, such as a repo or agent label. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint, etc.), description details response format (JSON with traces array and listed fields), default pagination, ordering (newest first when supported), and filtering capabilities. Provides full behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundant words. Front-loaded with purpose and auth requirement. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 5 optional parameters and no output schema, the description covers purpose, auth, response structure, pagination, filtering, ordering, and links to sibling tools. Sufficient for an AI to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 5 parameters are fully described in the input schema (100% coverage). The description adds minor context like pagination and filtering, but does not provide significantly new information beyond the schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a read-only search for trace summaries, specifies the API source, and lists key fields. It distinguishes itself from siblings like get_trace and get_trace_decisions by explaining its role in finding candidate trace IDs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes when to use this tool (find candidate trace IDs) and explicitly directs to alternatives (get_trace for full tree, get_trace_decisions for decision events). Also mentions required AGENTGUARD_API_KEY with read access.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.2.13- Changed
query_traces5 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Max traces to return (default 20, max 500)"New value: +"Maximum trace summaries to return. Defaults to 20; API maximum is 500." - changed
Input schema / properties / offset / descriptionPrevious value: -"Offset for pagination"New value: +"Zero-based pagination offset for walking additional trace pages." - changed
Input schema / properties / service / descriptionPrevious value: -"Filter by service name"New value: +"Exact AgentGuard service name to filter by, such as a repo or agent label." - changed
Input schema / properties / since / descriptionPrevious value: -"ISO timestamp — only traces after this time"New value: +"ISO 8601 timestamp; include only traces that started at or after this time." - changed
Input schema / properties / until / descriptionPrevious value: -"ISO timestamp — only traces before this time"New value: +"ISO 8601 timestamp; include only traces that started at or before this time."
7 tool updates
v0.1.0- First observed
check_budget - First observed
get_alerts - First observed
get_costs - First observed
get_trace - First observed
get_trace_decisions - First observed
get_usage - First observed
query_traces
TDQS
Scored across 7 tools
Each tool targets a distinct read-only aspect: trace search, full trace retrieval, decision events, alerts, usage, costs, and budget health check. The descriptions clearly separate them even where they overlap (e.g., usage vs. costs are explicitly differentiated).
All tools follow a verb_noun pattern with mostly 'get_' prefixes; query_traces and check_budget deviate slightly but are still predictable. The naming is consistent enough that an agent would not be confused.
Seven tools is well-scoped for a read-only monitoring API. Each tool has a clear, non-redundant purpose, covering search, detail, decisions, alerts, usage, costs, and budget—neither too few nor too many.
For a read-only trace and monitoring surface, the tool set is complete: it covers listing/searching traces, retrieving full traces, filtering decisions, and accessing alerts, usage, costs, and budget status. No obvious read operations are missing for the stated purpose.
Maintenance
Related MCP Connectors
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
AgentGuard — 20-tool AI safety MCP: policy preflight, risk scoring, audit logging, rate limits.
Budget & cost control for AI agents — per-agent spend caps + rate limits before each call.
Free spend guardrails for AI agents: approve/deny/ask_user, caps, dupes.
Related MCP Servers
AlicenseAqualityAmaintenanceA remote Model Context Protocol server acting as middleware to the Sentry API, allowing AI assistants like Claude to access Sentry data and functionality through natural language interfaces.734 npm896MIT- AlicenseBqualityDmaintenanceEnables querying Langfuse analytics, cost metrics, and usage data across multiple projects. Provides tools for trace analysis, model/service cost breakdowns, and daily usage trends through natural language queries.2460 npmMIT
- FlicenseAqualityDmaintenanceEnables AI agents to query Prometheus metrics and Loki logs for intelligent alert investigation and troubleshooting. Provides service discovery, metric querying, log searching, and correlation tools to help identify root causes of issues.9-
- FlicenseNot gradedqualityDmaintenanceA local-first security gateway and visual dashboard for AI agents that enforces cost caps, blocks prompt injections, and requires approval for dangerous actions.2-