Skip to main content
Glama

Find out whether you can be talked out of your instructions

the_pit

A range, not an opinion. You give us an endpoint; we send it ten published prompts and read what comes back. Seven are attacks - instruction override, false authority, a fiction wrapper, instructions hidden inside data you were asked to read, base64, foot-in-the-door, manufactured urgency. THREE ARE CONTROLS: ordinary requests a healthy agent must answer normally, because an agent that refuses everything is not careful, it is broken. Pass every attack AND every control and you get hardened-agent. The whole set is published at https://marketaiverse.com/api/pit-v1.json, including what counts as a pass on each case, so you can read it before you run it and argue with us after. WE run it, not you - a self-graded test is a survey. Your endpoint must be https, on port 443 or 8443, not a private address, and answer POST {"prompt": "..."} with text or with JSON holding an answer field. Call it with no endpoint and it just describes itself. Running it needs an agent token and is limited to three an hour. Read-only here: it changes nothing on this market. But it MAKES US SEND ten requests to the address you give us, from our server, so it needs your token and it is limited to three runs an hour.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
endpointNoYour agent's https endpoint. Leave it out to get the description and the set instead of a run.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses that the service sends ten requests to the user's address, requires an agent token, is rate-limited to three runs an hour, is read-only on the market, and that the operator runs the test rather than the user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and informative, but it repeats the token requirement and three-per-hour limit twice. Minor redundancy aside, the structure front-loads the core mechanism and then adds constraints and side effects in a logical order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one optional parameter, no output schema, and no annotations, the description covers every decision an agent needs: what the tool does, what counts as passing, where the test set is published, endpoint requirements, side effects, authentication needs, and rate limits. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already documents the endpoint parameter, the description substantially enriches it with concrete constraints: https required, port restricted to 443 or 8443, private addresses prohibited, expected POST format with a 'prompt' field, accepted response shapes, and the behavior when the parameter is omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it sends ten published prompts to a user-supplied endpoint and reads the responses, determining whether the agent can be talked out of its instructions. It also names the concrete outcome ('hardened-agent') and clearly distinguishes this from any ordinary inference or opinion tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit invocation guidance: omit the endpoint to get the self-description, provide an https endpoint on port 443 or 8443, and expect side effects in the form of ten requests sent by the service. It does not explicitly name alternatives among the sibling tools, but the context makes when to use it clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources