Skip to main content
Glama

agent_trifecta_score

Score an agent's lethal trifecta risk (private data + untrusted content + outbound actions). Returns risk level, missing controls, and decomposition advice.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
has_private_dataYesDoes the agent access private/sensitive data?
has_outbound_actionsYesCan the agent take outbound actions (send, write, call APIs)?
compensating_controlsYesControls in place, e.g. ['redact_secrets', 'smart_approvals']
has_untrusted_contentYesDoes the agent process untrusted external content?

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With empty annotations, the description carries full burden for behavioral disclosure. It states what the tool returns but does not disclose whether it is read-only, requires authentication, has rate limits, or any side effects. For a scoring tool, the lack of explicit read-only hint or behavioral caveats is a gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the main purpose and includes both inputs and outputs. It is efficient with no filler, but could be slightly improved by separating the return items for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 required parameters (all boolean and an array) and no output schema, the description explains the concept ('lethal trifecta') and lists return values. It provides enough context for a low-complexity scoring tool, though a bit more detail on what 'decomposition advice' entails would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's phrase 'private data + untrusted content + outbound actions' maps to the boolean parameters but adds no new nuance. The 'compensating_controls' array is mentioned in the description only as part of the return, not in parameter semantics. The schema already fully documents the parameters, so the description adds minimal value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scores an agent's 'lethal trifecta risk' and lists its components (private data, untrusted content, outbound actions). It also specifies the outputs (risk level, missing controls, decomposition advice). This provides a specific verb-resource pair and distinguishes it from sibling tools like 'agent_security_policies' or 'check_agent_policy' that focus on broader security or policy checking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives, nor does it mention prerequisites, when not to use it, or which sibling tools might be more appropriate for different contexts. It is purely functional.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

C2.2/5.0
Disambiguation3/5

Tools cover very diverse domains (weather, FDA, legal, crypto, etc.), so cross-domain confusion is low. However, within domains there is notable overlap: multiple food recall tools (food_recall_check, food_safety), multiple weather tools (weather_current_global, weather_forecast_grid, weather_alerts, weather_bias), and several Polymarket-related tools. This can cause agent misselection.

Naming Consistency2/5

Naming is inconsistent: some tools use verb_noun (search_arxiv, scrape, validate_agent_manifest), others use noun phrases (smart_money, space_weather, tide_data), and some are long descriptive phrases (cross_platform_arb_scan, polymarket_event_scan). No single pattern is followed, making predictions difficult.

Tool Count1/5

95 tools is excessively high for any coherent purpose. The server appears to be a random aggregation of APIs with no clear scope. Such a large catalog overwhelms agents and dilutes utility; most tools could be split into specialized servers.

Completeness2/5

Although many domains are touched, each is covered only shallowly. For example, weather lacks historical data, legal lacks case details beyond court opinions, and financial lacks stock prices. There are obvious gaps like no user authentication or data persistence. The tool set feels like a collection of endpoints rather than a cohesive service.

Resources