Agentguard47
AgentGuard
Stop runaway Python agents before they burn money.
AgentGuard47 is a zero-dependency runtime control SDK for Python agents. Add hard budget caps, loop detection, retry limits, timeouts, local traces, and incident reports without changing agent frameworks or sending data anywhere by default.
Use it when an agent can call tools, retry work, review code, or run long enough to create surprise spend.
pip install agentguard47Why AgentGuard
Most agent tooling tells you what happened after the run. AgentGuard stops the bad run while it is happening.
Problem | What AgentGuard does |
Agent loops on the same tool | Raises |
Flaky tool retries forever | Raises |
Run spends too much | Raises |
Run hangs | Raises |
Team needs proof | Writes local JSONL traces and incident reports |
Dashboard comes later |
|
Design constraints:
zero runtime dependencies
MIT licensed
local-first by default
no API key required for local proof
no network calls unless you configure
HttpSinkguards raise exceptions inside the running process
Related MCP server: Langfuse MCP Server
Real Incidents AgentGuard Prevents
PocketOS — agent deleted prod DB and backups in 9 seconds (May 2026)
A Cursor agent ran a destructive sequence against PocketOS production and wiped the live database. Backups went with it.
Reported root cause from the team's postmortem:
one API key had write + delete on both prod and backups
backups lived in the same Railway environment as prod
no confirmation step before destructive actions
the agent was given enough rope to chain the calls in one turn
Source: r/devops thread
The "AI did it" framing buries the actual lesson: the blast radius was infra, not the model. AgentGuard does not replace least-privilege creds or isolated backups. It does kill the run before a loop, retry storm, or runaway turn finishes the job.
A BudgetGuard plus LoopGuard wired around the agent loop caps how much
it can do in one session:
from agentguard import BudgetGuard, LoopGuard, RateLimitGuard, Tracer
budget = BudgetGuard(max_calls=20, max_cost_usd=1.00)
loop = LoopGuard(max_repeats=2)
rate = RateLimitGuard(max_calls_per_minute=10)
tracer = Tracer(service="cursor-agent", guards=[loop, rate])
with tracer.trace("agent.run"):
budget.consume(calls=1)
# tool call here — guards raise on overrunA 9-second sequence of destructive calls trips LoopGuard or
RateLimitGuard long before it finishes. The exception kills the run
in-process. Pair this with scoped credentials and out-of-environment
backups for the rest of the blast radius.
Local Proof In 60 Seconds
agentguard doctor
agentguard demo
agentguard quickstart --framework rawdoctor verifies the install and local trace writing.
demo proves budget, loop, and retry stops offline.
quickstart prints the smallest starter for your stack.
Installed-package proof:
agentguard demoSource-checkout proof with local incident output and hosted-compatible NDJSON:
git clone https://github.com/bmdhodl/agent47.git
cd agent47
PYTHONPATH=sdk python examples/sticky_agent_proof.py --out-dir proof/sticky-agent-proof
agentguard incident proof/sticky-agent-proof/sticky_agent_proof_traces.jsonlExpected first value moment:
BudgetGuard stops simulated spend.
LoopGuard stops repeated tool calls.
RetryGuard stops a retry storm.
No API keys. No dashboard. No network calls.Notebook version:
Copy-Paste Repo Setup
Use this when you want a coding agent or teammate to add AgentGuard safely:
pip install agentguard47
agentguard doctor
agentguard quickstart --framework raw --write
python agentguard_raw_quickstart.py
agentguard report .agentguard/traces.jsonlOptional shared local defaults, saved as .agentguard.json in the repo root:
{
"profile": "coding-agent",
"service": "my-agent",
"trace_file": ".agentguard/traces.jsonl",
"budget_usd": 5.0
}Keep the first PR local-only. Add hosted ingest later only when retained incidents, alerts, or team visibility matter.
Quickstart: Guard One Agent Run
from agentguard import BudgetGuard, JsonlFileSink, LoopGuard, Tracer
budget = BudgetGuard(max_cost_usd=5.00, max_calls=50, warn_at_pct=0.8)
loop = LoopGuard(max_repeats=3)
tracer = Tracer(
sink=JsonlFileSink(".agentguard/traces.jsonl"),
service="support-agent",
guards=[loop],
)
with tracer.trace("agent.run") as span:
budget.consume(calls=1, cost_usd=0.02)
loop.check("search", {"query": "refund policy"})
span.event("tool.call", data={"tool": "search", "query": "refund policy"})
# Call your agent or tool here.Inspect the local proof:
agentguard report .agentguard/traces.jsonl
agentguard incident .agentguard/traces.jsonlAuto-Patch Provider SDKs
If you already call OpenAI or Anthropic directly, patch once and keep using the provider normally:
from agentguard import BudgetGuard, Tracer, patch_openai
budget = BudgetGuard(max_cost_usd=5.00, warn_at_pct=0.8)
tracer = Tracer(service="support-agent")
patch_openai(tracer, budget_guard=budget)
# OpenAI chat completions are now traced and budget-enforced.When accumulated cost crosses the hard limit, BudgetExceeded is raised and
the agent stops.
Guards
Guard | Stops | Example |
| dollar, token, or call overruns |
|
| exact repeated tool calls |
|
| similar calls and A-B-A-B loops |
|
| retry storms on the same tool |
|
| long-running jobs |
|
| calls per minute |
|
| hard turns that need a stronger model |
|
Guards are static runtime checks. They do not ask another model whether a run is safe. They raise exceptions.
Examples
All examples are local-first. No API key is required unless the example says so.
Example | What it proves |
budget, loop, and retry stops | |
one CrewAI-style retry storm proof with local incident and hosted NDJSON outputs | |
review/refinement loop stopped by budget and retry guards | |
one oversized token-heavy turn can blow a run budget | |
when to escalate from a cheap model to a stronger one | |
proposal, edit, approval, and binding decision events |
Sample incident:
docs/examples/coding-agent-review-loop-incident.md
Proof gallery:
docs/examples/proof-gallery.md
Starter files:
examples/starters/
Framework Integrations
AgentGuard can wrap raw Python code or integrate with common agent stacks.
agentguard quickstart --framework raw
agentguard quickstart --framework openai
agentguard quickstart --framework anthropic
agentguard quickstart --framework langchain
agentguard quickstart --framework langgraph
agentguard quickstart --framework crewaiOptional integration extras are opt-in. The core SDK stays stdlib-only.
pip install "agentguard47[langchain]"
pip install "agentguard47[langgraph]"
pip install "agentguard47[crewai]"
pip install "agentguard47[otel]"Runtime Control vs Observability
AgentGuard is not a generic tracing platform. It is the local runtime stop layer.
Capability | AgentGuard |
In-process hard budget caps | Yes |
Kill a bad run by raising an exception | Yes |
Loop and retry-storm detection | Yes |
Local JSONL traces | Yes |
Local incident reports | Yes |
Hosted ingest | Optional |
Required dashboard | No |
Runtime dependencies | None |
Competitive notes:
Decision Traces
Capture proposal, human edit, approval, override, and binding events through the same event pipeline:
from agentguard import JsonlFileSink, Tracer, decision_flow
tracer = Tracer(sink=JsonlFileSink(".agentguard/traces.jsonl"))
with tracer.trace("agent.run") as run:
with decision_flow(
run,
workflow_id="deploy-review",
object_type="pull_request",
object_id="123",
actor_type="human",
actor_id="pat",
) as decision:
decision.proposed({"action": "merge"})
decision.approved(comment="Looks safe")
decision.bound(binding_state="merged", outcome="success")Supported event types:
decision.proposeddecision.editeddecision.overriddendecision.approveddecision.bound
Guide: docs/guides/decision-tracing.md
MCP Server
AgentGuard also ships a read-only MCP server for coding-agent workflows:
npx -y @agentguard47/mcp-serverUse the SDK to enforce local safety where the agent runs. Use MCP when a client like Codex, Claude Code, or Cursor needs read access to traces, decisions, costs, usage, and budget health.
Hosted Dashboard Boundary
The SDK is the free local proof path. The hosted dashboard is for retained history, alerts, team visibility, spend trends, hosted decision history, and dashboard-managed remote kill signals.
Use local SDK when | Use hosted dashboard when |
You are proving AgentGuard in one repo | Multiple people need the same incident history |
You need hard stops for loops, retries, timeouts, or budget burn | Runs need retained alerts and follow-up outside the terminal |
You want JSONL traces and reports without an API key | You need spend trends across traces, services, or teammates |
You are testing an agent before production | Operators need dashboard-managed remote kill signals |
Start local. Add hosted ingest when the work becomes shared, expensive, or risky enough that local files are no longer enough.
from agentguard import HttpSink, Tracer
tracer = Tracer(
sink=HttpSink(
url="https://app.agentguard47.com/api/ingest",
api_key="ag_...",
)
)HttpSink mirrors trace and decision events to the dashboard. It does not
execute remote kill signals by itself.
Dashboard contract:
docs/guides/dashboard-contract.md
Reports And CI Gates
Generate a local incident report:
agentguard incident .agentguard/traces.jsonl --format markdown
agentguard incident .agentguard/traces.jsonl --format htmlFail CI when a trace violates safety expectations:
from agentguard import EvalSuite
result = (
EvalSuite(".agentguard/traces.jsonl")
.assert_no_loops()
.assert_budget_under(tokens=50_000)
.assert_no_errors()
.run()
)
assert result.passedPackage Facts
Package:
agentguard47Python: 3.9+
License: MIT
Core runtime dependencies: zero
Trace format: JSONL
Local commands:
doctor,demo,quickstart,report,incident,evalMCP package:
@agentguard47/mcp-server
Docs
Topic | Link |
Getting started | |
Coding-agent setup | |
Safety pack | |
Dashboard contract | |
Decision traces | |
Managed sessions | |
Activation metrics design | |
Proof gallery | |
PyPI Trusted Publishing |
Architecture
agent code
|
v
Tracer
|
+-- guards raise exceptions locally
|
+-- sinks write traces locally or mirror to hosted ingestRepository layout:
sdk/ Python SDK package
mcp-server/ read-only MCP server
docs/ guides and competitive notes
examples/ runnable local examples
ops/ repo operating docs
memory/ SDK-only state and decisionsSecurity
No secrets are required for local mode.
Do not put API keys in
.agentguard.json.Hosted ingest API keys should be stored in environment variables.
Local guards remain authoritative even when hosted ingest is configured.
Report security issues through GitHub Security Advisories or by email:
pat@bmdpat.com.
Contributing
Contributions are welcome when they keep the SDK small, local-first, and zero-dependency.
Before opening a PR:
python -m pytest sdk/tests/ -v
python -m ruff check sdk/agentguard/
python scripts/sdk_release_guard.pyUseful links:
License
MIT. See LICENSE.
Available Tools
7 toolscheck_budgetARead-onlyIdempotent
Read-only hosted event-quota health check. Combines dashboard event quota and recorded cost summaries. This is not SDK BudgetGuard, not a provider invoice cap, and it does not refuse the next model call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds useful context beyond those annotations, notably that it does not refuse the next model call and that it combines two data sources. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. It front-loads the core purpose and then efficiently adds exclusions that prevent misuse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument read-only health check, the description provides enough context for correct selection and invocation. It lacks an explicit description of the return format, but this is a minor gap given the tool's simplicity and the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantic burden on the description. The schema coverage is effectively complete, and no parameter explanation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read-only health check for hosted event quota and states it combines quota and cost summary data. It is specific about what it does, though it does not explicitly distinguish itself from sibling tools like get_usage or get_costs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking hosted event quota health and includes negative guidance (not SDK BudgetGuard, not a provider invoice cap). However, it does not explicitly state when to use this tool over alternatives or name sibling tools as fallbacks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_alertsARead-onlyIdempotent
Read-only recent guard alerts (loop detection, budget exceeded) and errors from the hosted Read API. This reports stored alerts; it does not stop a running agent.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max alerts to return (default 50) | |
| since | No | ISO timestamp — only alerts after this time |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, idempotent, and non-destructive. The description adds the valuable clarification that it only reports stored alerts and never stops a running agent, which is not fully captured by the annotation hints. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundancy: the first states purpose and scope, the second clarifies a key side-effect. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with two optional, fully documented parameters and safety annotations already covering behavior, the description is largely complete. It could hint at ordering or default behavior, but nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes both limit and since with their defaults and meanings, so schema coverage is 100%. The description itself adds no parameter-level detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves recent guard alerts and errors from the hosted Read API in a read-only manner. It distinguishes the resource (alerts) from siblings like traces, usage, costs, and budget, though it does not explicitly name an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes it reports stored alerts and does not stop a running agent, giving one useful exclusion. However, it does not explicitly explain when to choose this over related tools such as check_budget or get_trace_decisions, leaving usage implications to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_costsARead-onlyIdempotent
Read-only hosted cost breakdown for the current month from the AgentGuard Read API. Estimated savings are recorded guard events, not an invoice credit.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds useful context about estimated savings being recorded guard events, not an invoice credit, which helps set expectations about the data's meaning. It doesn't mention pagination or response format, but that's acceptable for a zero-parameter tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, each earning its place: the first states the scope and source, the second clarifies an important caveat about the data's meaning. It's concise and front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with rich annotations, the description is largely complete. It clarifies the nature of the cost breakdown and the caveat about estimated savings. The only minor gap is not detailing the output structure, but since there's no output schema, it could be more explicit, but the tool is simple enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since there are 0 parameters, the schema provides full coverage at 100%, and the description doesn't need to explain any parameters. A baseline of 4 is appropriate because there is nothing to clarify, and the description correctly focuses on the output's semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (get) and resource (hosted cost breakdown for the current month), and clarifies it is read-only. It does not explicitly differentiate from siblings like get_usage or check_budget, but the resource is distinct enough for an agent to understand its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for viewing current month's costs but provides no explicit guidance on when to use this tool versus siblings like get_usage or check_budget. It mentions the source (AgentGuard Read API) but doesn't state exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_traceARead-onlyIdempotent
Read-only full event tree for one retained trace. Shows spans, tool calls, LLM calls, guard triggers, and errors. Does not enforce BudgetGuard or intercept provider calls.
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | The trace ID to look up |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, open-world, and non-destructive behavior, so the description adds value beyond them by enumerating the event types returned and explicitly negating BudgetGuard enforcement and provider-call interception. No behavioral surprises remain undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences lead with the read-only scope and core output, then list return contents and end with an important non-behavior. There is no filler or repetition of annotation metadata.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with rich annotations, the description covers what the tool returns and what side effects it avoids. The absence of an output schema is acceptable because the description enumerates the event categories the returned tree contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the only parameter, trace_id, is already described as 'The trace ID to look up'. The description reinforces that it targets a single retained trace but does not add format, source, or lookup semantics beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (get) and resource (one retained trace), and describes the output as a 'full event tree' that includes spans, tool calls, LLM calls, guard triggers, and errors. This clearly differentiates it from the sibling tools, which target listings, decisions, alerts, usage, costs, or budget checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use this for read-only inspection of a single retained trace and the full event tree. It also states exclusions ('does not enforce BudgetGuard or intercept provider calls'), helping an agent avoid expecting side effects, though it does not explicitly name sibling alternatives or when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trace_decisionsARead-onlyIdempotent
Read-only normalized decision.* events from one retained trace. Use this when a workflow includes proposal, override, approval, or binding steps. Does not approve, bind, or enforce those decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | The trace ID to inspect for decision events |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true. The description adds value by stating 'Read-only' at the beginning and explicitly saying it does not approve, bind, or enforce decisions, which reinforces the non-mutating nature beyond the annotation flags. This context about scope (retained trace, normalized events) is useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with zero fluff. It front-loads the core purpose, then gives usage guidance, then clarifies a potential misassumption. Every sentence earns its place, and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with rich annotations and no output schema, the description covers the main points: what it returns (decision.* events), the scope (one retained trace), and what it does not do (approve/bind/enforce). It is nearly complete, though it does not elaborate on what 'normalized' means or whether there are limits on event counts, which is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter trace_id already has a clear description ('The trace ID to inspect for decision events'). The tool description adds only marginal detail about the parameter ('from one retained trace'), which hints at a constraint but does not significantly expand the schema's meaning. A baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read-only normalized decision.* events from one retained trace.' This clearly distinguishes it from sibling tools like get_trace or query_traces, and the mention of proposal, override, approval, and binding steps further scopes what kind of decisions are covered.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage context: 'Use this when a workflow includes proposal, override, approval, or binding steps.' It also states what the tool does not do ('Does not approve, bind, or enforce those decisions'), which gives a when-not. However, it does not name specific alternative sibling tools, so it falls slightly short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageARead-onlyIdempotent
Read-only hosted event quota usage and plan limits from the AgentGuard Read API. This is dashboard event quota, not SDK BudgetGuard and not a provider invoice cap.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds useful context beyond annotations by clarifying the data source and scope (dashboard event quota vs SDK BudgetGuard vs provider invoice cap), which helps the agent understand exactly what this read operation covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundancy. The primary purpose is front-loaded, and the second sentence earns its place by disambiguating from related concepts. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with rich annotations, the description is complete. It states the resource, the source API, and the return scope (quota usage and plan limits), and clarifies what it is not, so an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is no parameter documentation burden on the description. The description appropriately focuses on what the tool returns rather than parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read-only'), a clear resource ('hosted event quota usage and plan limits'), and the source ('AgentGuard Read API'). It also explicitly distinguishes this from SDK BudgetGuard and provider invoice caps, which separates it from sibling tools like check_budget and get_costs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context that this tool is for dashboard event quota, not SDK BudgetGuard or provider invoice cap. It gives when-not guidance but does not explicitly name alternative sibling tools to use in those cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_tracesARead-onlyIdempotent
Read-only search for retained AgentGuard trace summaries from the AgentGuard Read API. Requires AGENTGUARD_API_KEY with read access; create keys in the AgentGuard dashboard. Returns JSON with a traces array, newest traces first when the API supports ordering; items include trace_id, service, root_name, event_count, error_count, duration_ms, started_at, API key metadata, and total_cost when available. Defaults to a small page, accepts offset pagination, exact service filtering, and ISO 8601 since/until bounds. Use this to find candidate trace_id values; use get_trace for the full event tree of one trace or get_trace_decisions for decision.* events from a known trace.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum trace summaries to return. Defaults to 20; API maximum is 500. | |
| since | No | ISO 8601 timestamp; include only traces that started at or after this time. | |
| until | No | ISO 8601 timestamp; include only traces that started at or before this time. | |
| offset | No | Zero-based pagination offset for walking additional trace pages. | |
| service | No | Exact AgentGuard service name to filter by, such as a repo or agent label. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint, etc.), description details response format (JSON with traces array and listed fields), default pagination, ordering (newest first when supported), and filtering capabilities. Provides full behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no redundant words. Front-loaded with purpose and auth requirement. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 5 optional parameters and no output schema, the description covers purpose, auth, response structure, pagination, filtering, ordering, and links to sibling tools. Sufficient for an AI to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 5 parameters are fully described in the input schema (100% coverage). The description adds minor context like pagination and filtering, but does not provide significantly new information beyond the schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a read-only search for trace summaries, specifies the API source, and lists key fields. It distinguishes itself from siblings like get_trace and get_trace_decisions by explaining its role in finding candidate trace IDs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes when to use this tool (find candidate trace IDs) and explicitly directs to alternatives (get_trace for full tree, get_trace_decisions for decision events). Also mentions required AGENTGUARD_API_KEY with read access.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.2.13- Changed
query_traces5 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Max traces to return (default 20, max 500)"New value: +"Maximum trace summaries to return. Defaults to 20; API maximum is 500." - changed
Input schema / properties / offset / descriptionPrevious value: -"Offset for pagination"New value: +"Zero-based pagination offset for walking additional trace pages." - changed
Input schema / properties / service / descriptionPrevious value: -"Filter by service name"New value: +"Exact AgentGuard service name to filter by, such as a repo or agent label." - changed
Input schema / properties / since / descriptionPrevious value: -"ISO timestamp — only traces after this time"New value: +"ISO 8601 timestamp; include only traces that started at or after this time." - changed
Input schema / properties / until / descriptionPrevious value: -"ISO timestamp — only traces before this time"New value: +"ISO 8601 timestamp; include only traces that started at or before this time."
7 tool updates
v0.1.0- First observed
check_budget - First observed
get_alerts - First observed
get_costs - First observed
get_trace - First observed
get_trace_decisions - First observed
get_usage - First observed
query_traces
TDQS
Scored across 7 tools
Each tool targets a distinct read-only aspect: trace search, full trace retrieval, decision events, alerts, usage, costs, and budget health check. The descriptions clearly separate them even where they overlap (e.g., usage vs. costs are explicitly differentiated).
All tools follow a verb_noun pattern with mostly 'get_' prefixes; query_traces and check_budget deviate slightly but are still predictable. The naming is consistent enough that an agent would not be confused.
Seven tools is well-scoped for a read-only monitoring API. Each tool has a clear, non-redundant purpose, covering search, detail, decisions, alerts, usage, costs, and budget—neither too few nor too many.
For a read-only trace and monitoring surface, the tool set is complete: it covers listing/searching traces, retrieving full traces, filtering decisions, and accessing alerts, usage, costs, and budget status. No obvious read operations are missing for the stated purpose.
Maintenance
Related MCP Connectors
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
AgentGuard — 20-tool AI safety MCP: policy preflight, risk scoring, audit logging, rate limits.
Budget & cost control for AI agents — per-agent spend caps + rate limits before each call.
Free spend guardrails for AI agents: approve/deny/ask_user, caps, dupes.
Related MCP Servers
AlicenseAqualityAmaintenanceA remote Model Context Protocol server acting as middleware to the Sentry API, allowing AI assistants like Claude to access Sentry data and functionality through natural language interfaces.734 npm896MIT- AlicenseBqualityDmaintenanceEnables querying Langfuse analytics, cost metrics, and usage data across multiple projects. Provides tools for trace analysis, model/service cost breakdowns, and daily usage trends through natural language queries.2460 npmMIT
- FlicenseAqualityDmaintenanceEnables AI agents to query Prometheus metrics and Loki logs for intelligent alert investigation and troubleshooting. Provides service discovery, metric querying, log searching, and correlation tools to help identify root causes of issues.9-
- FlicenseNot gradedqualityDmaintenanceA local-first security gateway and visual dashboard for AI agents that enforces cost caps, blocks prompt injections, and requires approval for dangerous actions.2-