Skip to main content
Glama

audit_financial

Use this before acting on a financial decision with real money at stake — a drafted trade thesis before placing the trade, a risk assessment before approving, an approve/decline call before committing, a backtest interpretation before deploying. Four frontier models critique it for tail risks, unstated assumptions, and asymmetric downside; specific weaknesses are routed back as targeted challenges; each model revises in light of specific counter-positions. Returns a structured critique with severity tags and a recommended action class. Runs ~2-5 min with no progress shown mid-call — tell the user it's working before you call. For a fast broad take use synthesize_financial; for an open-ended decision use deliberate_financial.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
assumptionsNo
constraintsNo
eval_case_idNo
relevant_dataNo
proposed_actionNo
tests_backtestsNo
continuation_tokenNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the process (four frontier models critique, weaknesses routed back, models revise), the runtime ('~2-5 min'), and the lack of progress ('no progress shown mid-call'), including an actionable instruction to tell the user it's working. However, it does not explicitly state whether the tool has side effects (e.g., writes to storage), though its read-only nature is implied. A small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the primary purpose and usage, then details the process and runtime. It is somewhat verbose but every sentence contributes value. The structure is logical, moving from use case to behavior to return format to alternatives. It could be tightened but is not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the return format at a high level ('structured critique with severity tags and a recommended action class') but lacks detail on the exact structure. More importantly, it does not explain any of the parameters, which are all optional but essential for the agent to know what to supply. The absence of parameter semantics makes the tool under-specified for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description provides no explanation of any of the 7 parameters (assumptions, constraints, eval_case_id, relevant_data, proposed_action, tests_backtests, continuation_token). The description does not mention parameter names or their intended content, leaving the agent to guess what to pass. This is a critical omission given the zero coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to audit financial decisions with real money at stake before acting. It provides concrete examples (trade thesis, risk assessment, approve/decline calls, backtest interpretation) and specifies the output (structured critique with severity tags and recommended action class). It also distinguishes from siblings by naming synthesize_financial and deliberate_financial as alternatives for different needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly defines when to use the tool ('before acting on a financial decision with real money at stake') and gives specific scenarios. It also provides clear guidance on alternatives: 'For a fast broad take use synthesize_financial; for an open-ended decision use deliberate_financial.' This gives the agent explicit routing conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.