overwing-mcp
Overwing MCP lets MCP-capable agents guardrail LLM output by scoring text against rule sets and returning pass/fail/review verdicts with confidence, plus manage rule sets, retrieve evaluations, and inspect usage/plans.
evaluate: score one text against a rule set (defaultcontent-safety) and get verdict, recommended action, score, confidence, latency, and per-rule results.evaluate_batch: score up to 50 texts in one call with per-item verdicts and a summary.list_rule_sets/get_rule_set/create_rule_set: browse prebuilt rule sets, inspect rule definitions, and create custom rules (choice, score, or yes/no) with fail conditions, review thresholds, and weights.get_evaluation/list_evaluations: fetch stored evaluation results, filter by verdict or rule set, and page with a cursor.atlas_lookup: verify User-Agent claims (web bot auth, spoofable, or unattributable).atlas_agents/atlas_summary: search the AI crawler/browser agent registry and view traffic-share or spending summaries.get_usage/whoami/list_plans: check daily quota and remaining usage, identify the org/plan behind the API key, and list public pricing plans.The
overwing://guideresource provides the full plain-text API guide.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@overwing-mcpCheck this reply for PII before sending: 'Call me at 555-0142'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Overwing is an API that checks what your model said before it ships. This package exposes it to any MCP-capable agent: Claude Desktop, Claude Code, Cursor, Windsurf, VS Code, OpenAI's Agents SDK, and anything else that speaks the Model Context Protocol.
Real verdicts, not vibes. Every rule returns a typed answer, a probability, and a confidence.
failmeans a rule matched;reviewmeans it was unsure;passmeans neither.Prebuilt
content-safetyrule set: toxicity, personal data, self-harm, sexual content, severity. Or write your own rules in plain language.Built for agents. Sign up, pay, evaluate, rotate keys, and cancel, all as JSON. No CAPTCHA, no browser required. See overwing.ai/llms.txt.
Try it without installing anything: paste text into the console at overwing.ai.
Install
You need an API key. Get one at overwing.ai/login, or let your agent sign itself up:
curl -X POST https://overwing.ai/api/v1/signup \
-H "Content-Type: application/json" \
-d '{"email":"you@example.com","password":"at-least-12-chars"}'Claude Code
claude mcp add overwing -e OVERWING_API_KEY=ow_live_... -- npx -y overwing-mcpClaude Desktop, Cursor, Windsurf, VS Code (any JSON-configured client)
{
"mcpServers": {
"overwing": {
"command": "npx",
"args": ["-y", "overwing-mcp"],
"env": { "OVERWING_API_KEY": "ow_live_..." }
}
}
}Set OVERWING_BASE_URL to point at a self-hosted deployment. Requires Node 20+.
Related MCP server: Sentinel
Tools
Tool | What it does |
| Score one text against a rule set. Returns the verdict, a |
| Score up to 50 texts in one call, with a summary and per-item verdicts and recommended actions. |
| Browse the prebuilt set or define your own rules: yes/no questions, classifications, or scored scales. |
| Read stored results, filter by verdict or rule set, page with a cursor. |
| Overwing Atlas: say what a User-Agent string claims to be and whether the claim can be trusted (Web Bot Auth, spoofable string, or unattributable). 100 free a day. |
| Search the registry of 241 AI crawlers, fetchers and browser agents; get traffic shares, sector field-scan headlines, and the agent-spending summary. |
| Today's quota, the org behind the key, and the public plan catalog. |
The overwing://guide resource returns the full plain-text API guide.
Example
Ask your agent:
Check this reply before I send it: "Reach me at dana@example.com or 555-0142 to sort out the refund."
It calls evaluate and gets back:
Verdict: FAIL score=0.82 confidence=0.97 251ms
toxicity: pass (answer="safe", confidence=1)
pii_detected: fail (answer=true, confidence=0.98)
self_harm: pass (answer=false, confidence=1)
sexual_content: pass (answer="none", confidence=1)
severity: pass (answer=0.03, confidence=0.97)How verdicts work
Each rule has a fail condition, an optional review threshold, and a weight.
fail: a rule's fail condition matched. Block it, redact it, or regenerate.
review: nothing failed, but a rule's confidence was below its threshold. Route to a person or a slower model.
pass: everything else.
Every rule also carries an action (block, redact, or review), and the response rolls them up into one recommended_action: block beats redact beats review beats allow. Branch on that field. Pass a context object (recipient, channel, owns_contact_info) with the outbound-message rule set and its rules read it, so a customer's own phone number in a reply to that customer is not flagged.
aggregate_score is 0 to 1 (pass = 1, review = 0.5, fail = 0 per rule, weighted). confidence is the minimum across rules.
Pricing
Free: 250 evaluations a day. Paid plans from $29/month. Every plan includes every endpoint, custom rule sets, webhooks, and the dashboard. Full details at overwing.ai/#pricing or GET https://overwing.ai/api/v1/plans.
Links
Website and live demo: overwing.ai
API reference: overwing.ai/docs
OpenAPI: overwing.ai/api/v1/openapi.json
Guide for agents: overwing.ai/llms.txt
Support: support@overwing.ai
Development
npm install
npm run build
OVERWING_API_KEY=ow_live_... node dist/index.jsMIT © Overwing. Verdicts are produced by TypeSafe's Jev System One model; Overwing is not affiliated with TypeSafe.
Available Tools
10 toolscreate_rule_setCreate rule setB
Create a custom rule set with 1 to 25 rules. Each rule is a choice (pick one option), score (position on an ordered scale), or noul (yes/no) question with a fail condition, optional review threshold, and weight.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| slug | Yes | ||
| rules | Yes | ||
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the rule count limit (1-25) and rule structure, but doesn't disclose side effects (e.g., whether creation overwrites an existing slug, whether it requires authentication, or what happens on validation failure). For a mutation tool with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core purpose and then packs the key rule structure details. Every clause earns its place, though it could be slightly more structured with a second sentence for parameter semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with no output schema and no annotations, the description covers the rule structure well but omits important context: what the response looks like, whether slug must be unique, and any prerequisites (e.g., authentication). The sibling tools suggest a broader evaluation workflow, but the description doesn't situate this tool within that workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of 'rules' (each rule is a choice/score/noul question with fail condition, review threshold, weight) and the count constraint, which adds value beyond the raw schema. However, it doesn't explain the 'name' and 'slug' parameters' semantics (e.g., slug format, uniqueness), leaving some gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and resource ('custom rule set'), and enumerates the three question types (choice, score, noul) with their key attributes. It distinguishes itself from sibling tools like evaluate and list_rule_sets by clearly being a creation tool, though it doesn't explicitly name any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by describing what a rule set is and its constraints (1-25 rules, question types), but it doesn't explicitly state when to use this tool versus alternatives like evaluate or get_rule_set. The context is clear enough for an agent to infer this is for creating, not evaluating or listing, but no explicit when/when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluateEvaluate textA
Score any text (typically an LLM's output) against an Overwing rule set. Returns an aggregate verdict of pass, fail, or review plus per-rule answers with probability and confidence. fail means a rule's fail condition matched: block or redact. review means a rule was unsure: route to a human or slower model. The prebuilt 'content-safety' set checks toxicity, PII, self-harm, sexual content, and severity.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | The text to evaluate | |
| metadata | No | Opaque context stored with the evaluation (max 8 KB) | |
| rule_set | No | Rule set slug | content-safety |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses the aggregate verdict values, per-rule answers with probability and confidence, and the semantic meaning of fail and review. It does not explicitly mention persistence, auth, rate limits, or usage costs, but the core behavior is clearly and usefully disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose first, return shape second, verdict interpretations third. Every sentence adds meaningful information with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema, the description adequately explains return values and verdict semantics, making it callable. Minor gaps remain: it does not specify the exact JSON shape of per-rule answers or direct users to list_rule_sets for discovering available rule sets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining the rule_set concept, the default 'content-safety' set and its checks, and the intended input. This goes beyond the bare schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scores text against an Overwing rule set, with a specific verb and resource. It does not explicitly contrast itself with siblings like evaluate_batch, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it is for scoring text (commonly LLM output) against a rule set, and it explains what fail and review verdicts mean for downstream action. It does not provide explicit when-not-to-use guidance or point to alternatives such as evaluate_batch, but the context is strong enough for basic selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_batchEvaluate many textsA
Score up to 50 texts against one rule set in a single call. Each item counts as one evaluation. Returns a summary plus per-item verdicts; items can fail independently.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| rule_set | No | content-safety |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the return shape ('summary plus per-item verdicts'), failure independence ('items can fail independently'), and a business constraint ('Each item counts as one evaluation'). It does not cover authentication, rate limits, or invalid-input handling, but what it includes is genuine behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the operation, the limit/constraint, and the result shape. Front-loaded with the most important information about batching.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the description covers result shape and partial-failure semantics, which is helpful. However, it omits parameter details, usage guidance versus siblings, and any mention of the default rule set beyond its existence, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameters. It only hints at 'texts' and 'one rule set', leaving the 'items' object structure (id, metadata), the 'input' constraints, and the 'rule_set' default undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Score'), a clear resource ('up to 50 texts against one rule set'), and a distinguishing constraint ('in a single call'). It clearly separates batch evaluation from the sibling single-item 'evaluate' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies batch use via 'up to 50 texts' and 'single call', giving clear context for when this tool fits. It does not explicitly name the alternative 'evaluate' for single-text use or state exclusions, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evaluationGet evaluationA
Fetch a stored evaluation by id, including the original input and per-rule results.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it clearly indicates a read-only retrieval operation and discloses what is included in the result. It does not discuss error/not-found behavior, but for a simple get-by-id tool this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words, and the primary action and target are front-loaded. Every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter retrieval tool with no output schema, the description tells the agent what it will receive and by what identifier. It is complete enough to invoke correctly, though it could mention what happens when no evaluation matches the given id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does so by linking the 'id' parameter to the stored evaluation, adding meaning beyond the schema's bare string type. For a single obvious parameter, this is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Fetch') and resource ('stored evaluation'), and clarifies the scope ('by id') and content ('original input and per-rule results'). This clearly distinguishes it from sibling tools like list_evaluations or evaluate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Fetch a stored evaluation by id' gives clear context: use this when you have a specific evaluation ID and need the stored result. It does not explicitly name alternatives or exclusions, but the single-parameter design and retrieval wording imply the right usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rule_setGet rule setC
Fetch a rule set with its full rule definitions. Use 'content-safety' as a worked example when writing your own.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing behavior. It indicates the result includes full rule definitions, but it does not explain permissions, error behavior, output format, or whether this is a read-only operation beyond the verb 'Fetch'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The primary sentence is concise and front-loaded, but the second sentence is ambiguous and potentially distracting. It is brief but not maximally clear or purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple get-by-slug tool, the description covers the core action but omits important context: what a slug refers to, what the response contains beyond rule definitions, and any edge cases or prerequisites. No output schema or annotations exist to fill these gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and the only parameter, slug, is not explained. The phrase 'content-safety' hints at a valid slug value, but it is not clearly tied to the parameter, so the description fails to compensate adequately for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches a rule set with its full rule definitions, which is a specific verb and resource. It partially distinguishes from list_rule_sets by emphasizing full definitions, though it does not explicitly name the sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus list_rule_sets, create_rule_set, or evaluate. The second sentence about using 'content-safety' as a worked example is unclear and does not help route the agent to the correct tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageGet usageC
Daily usage, remaining quota for today, and plan limits.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states what data is returned but does not disclose whether this is a read-only operation, whether it counts against quota, whether it reflects cached or live data, or any rate-limit implications. For a usage/quota tool, this context is important.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the key data points. It earns its place, though it could add a brief note about the 'days' parameter without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and an undocumented parameter, the description is too thin. It does not explain the meaning of 'days', the time window semantics, whether the quota is per-user or per-workspace, or how the response is structured. An agent would need to guess several important details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one optional parameter, 'days', with 0% description coverage, so the schema provides no semantic meaning. The description does not mention 'days' at all, leaving the agent to infer that it controls the usage window. Since there is only one parameter and the tool name implies a time range, the gap is moderate but real.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair ('Get usage') and enumerates the specific data returned: daily usage, remaining quota for today, and plan limits. It is not a tautology and gives an agent a concrete idea of what the tool does, though it does not explicitly differentiate from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The sibling list includes list_plans and whoami, which could relate to plan/account context, but the description does not mention them or any exclusions. An agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_evaluationsList evaluationsA
List recent evaluations, newest first, with optional verdict and rule set filters. Use next_cursor to page.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| verdict | No | ||
| rule_set | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses ordering (newest first), optional filters, and pagination via next_cursor. However, it does not explicitly state that the operation is read-only or that it has no side effects, though 'list' implies it. This is adequate but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundancy. The core action and sorting are front-loaded, and the pagination note is appended. Every word serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description sufficiently conveys the purpose, ordering, filtering, and pagination. It implies the return of evaluation objects. It could mention authentication requirements, but these are likely implicit in the tool's context. Overall, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explicitly mentions verdict and rule_set as optional filters, and next_cursor for pagination, which covers three of four parameters. The limit parameter is not described, but its meaning is self-evident. The description adds value but does not fully elaborate on parameter formats or constraints beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists evaluations, sorted newest first, with optional filters. It distinguishes from siblings like get_evaluation (single item) and evaluate (create) by focusing on listing. The verb 'List' and resource 'evaluations' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It mentions optional filters and pagination, but does not explicitly state when to use this versus get_evaluation or evaluate. The context implies listing multiple evaluations, but it lacks explicit 'when not to use' guidance or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_plansList plansA
Public plan catalog: prices, daily limits, and per-minute burst limits. No API key needed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It clearly states the operation is public and requires no API key, and it enumerates the returned data categories. It does not discuss pagination or response format, but for a zero-parameter catalog this is reasonable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core purpose ('Public plan catalog') and packs the key details plus an auth note into the remainder. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, no-output-schema tool, the description covers the essential context: what the catalog contains, that it is public, and that no API key is needed. It does not describe exact response shape, but the listed contents give sufficient operational clarity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, and schema coverage is 100% by virtue of an empty schema. The description adds no parameter semantics, but the zero-parameter baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource as a public plan catalog and specifies its contents: prices, daily limits, and per-minute burst limits. This distinguishes it from sibling list tools like list_rule_sets and list_evaluations, though the description relies on the title's verb 'List' for the action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when plan information is needed and explicitly states no API key is required, which is useful access guidance. However, it does not mention when to prefer this over siblings like get_usage or list_rule_sets, nor does it provide exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_rule_setsList rule setsB
List the prebuilt and custom rule sets available to this organization.
| Name | Required | Description | Default |
|---|---|---|---|
| include_inactive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It clearly indicates this is a read/list operation scoped to the organization, but it does not disclose how the include_inactive parameter affects the returned list or whether inactive rule sets are excluded by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler words. It is concise and readable, though it could have added parameter context without much additional length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core purpose is stated, but with no annotations, no output schema, and no explanation of the only parameter, the definition is incomplete. An agent would not know the default behavior regarding inactive rule sets or how to request them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has only one parameter, include_inactive, and schema description coverage is 0%. The description does not mention this parameter at all, so an agent cannot infer whether inactive rule sets are shown, hidden, or controlled by the boolean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states the specific verb 'List' and the resource 'rule sets', and scopes it to prebuilt and custom rule sets available to the organization. This clearly differentiates it from siblings like get_rule_set, create_rule_set, and list_evaluations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided. The list phrasing and sibling names imply this is for enumerating rule sets rather than creating or evaluating, but no exclusion conditions or alternative routing are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
whoamiWho am IA
Identify the organization and plan behind the configured API key, and whether billing is set up.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of disclosing behavior. The verb 'identify' clearly communicates a read-only, non-mutating operation, and the description goes beyond the tool name by stating exactly which details are reported: organization, plan, and billing setup. This is sufficient for a simple introspection call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-formed sentence with no filler or redundancy. The key information is front-loaded and every word contributes to the tool's purpose and behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters)Skip and no output schema, the description is complete: it tells the agent what the tool does and what information it returns. Nothing essential is missing for an agent to decide whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing to explain. The description adds useful meaning by enumerating the response aspects (organization, plan, billing setup), which is the relevant semantic content for this tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Identify') and clearly names the resources it covers: the organization and plan behind the configured API key, plus billing setup status. This clearly distinguishes it from sibling tools like list_plans or get_usage, which address different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear by tying the tool to the 'configured API key' and its associated organization, plan, and billing status, which tells an agent when to call it. It does not explicitly name alternatives or exclusions, but the intent is clear enough for a no-parameter introspection tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.3.0- First observed
create_rule_set - First observed
evaluate - First observed
evaluate_batch - First observed
get_evaluation - First observed
get_rule_set - First observed
get_usage - First observed
list_evaluations - First observed
list_plans - First observed
list_rule_sets - First observed
whoami
TDQS
Scored across 10 tools
Each tool targets a distinct operation: single vs batch evaluation, rule set management, evaluation retrieval, and account/plan information. Even the plan-related tools (whoami, get_usage, list_plans) are cleanly separated by identity, usage, and catalog purposes.
Most tools follow a clear verb_noun pattern: list_rule_sets, get_rule_set, create_rule_set, get_evaluation, list_evaluations. evaluate, evaluate_batch, and whoami are minor deviations but still verb-first and predictable.
Ten tools is well-scoped for an LLM evaluation service: two evaluation paths, three rule set operations, two evaluation lookup operations, and three account/plan tools. Each tool earns its place without redundancy.
The evaluation lifecycle is well covered: evaluate, retrieve, and list evaluations, plus list/get/create rule sets. Update and delete for rule sets are missing, which is a minor gap, but agents can work around it by creating new rule sets.
Maintenance
Related MCP Connectors
Sentiment, toxicity, entity extraction, PII, translation, summary, QA, fraud scoring, safety audit.
Calibrated judgments for text: yes/no probabilities, picks from your options, or scores.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Scan text, documents, websites, and MCP metadata for prompt injection and sensitive-data risks.
Related MCP Servers
AlicenseBqualityDmaintenanceEnables MCP clients to check text content using chakoshi's guardrail API for safety and moderation, returning assessment results.13MIT- FlicenseNot gradedqualityBmaintenanceEnables AI agents to score their outgoing responses against groundedness and prompt-injection risks mid-turn, returning allow, warn, or block verdicts before the response reaches the user.11 npm-
- AlicenseAqualityCmaintenanceEnables agents to perform typed judgments—classify, score, check, match, and screen—over closed answer sets with confidence scores, without text generation.722MIT
- AlicenseAqualityAmaintenanceEnables agents to verify claims against cited evidence, screen content for prompt injection and relevance before reading it, and rank candidates by meaning, all with calibrated probability verdicts.118,716 npm415MIT