overwing-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| OVERWING_API_KEY | Yes | Your Overwing API key, e.g., ow_live_... Get one at https://overwing.ai/login | |
| OVERWING_BASE_URL | No | Optional base URL for a self-hosted Overwing deployment. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| resources | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| evaluateA | Score any text (typically an LLM's output) against an Overwing rule set. Returns an aggregate verdict of pass, fail, or review plus per-rule answers with probability and confidence. fail means a rule's fail condition matched: block or redact. review means a rule was unsure: route to a human or slower model. The prebuilt 'content-safety' set checks toxicity, PII, self-harm, sexual content, and severity. |
| evaluate_batchA | Score up to 50 texts against one rule set in a single call. Each item counts as one evaluation. Returns a summary plus per-item verdicts; items can fail independently. |
| list_rule_setsB | List the prebuilt and custom rule sets available to this organization. |
| get_rule_setC | Fetch a rule set with its full rule definitions. Use 'content-safety' as a worked example when writing your own. |
| create_rule_setB | Create a custom rule set with 1 to 25 rules. Each rule is a choice (pick one option), score (position on an ordered scale), or noul (yes/no) question with a fail condition, optional review threshold, and weight. |
| get_evaluationA | Fetch a stored evaluation by id, including the original input and per-rule results. |
| list_evaluationsA | List recent evaluations, newest first, with optional verdict and rule set filters. Use next_cursor to page. |
| get_usageC | Daily usage, remaining quota for today, and plan limits. |
| whoamiA | Identify the organization and plan behind the configured API key, and whether billing is set up. |
| list_plansA | Public plan catalog: prices, daily limits, and per-minute burst limits. No API key needed. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| guide | The plain-text API guide (llms.txt): auth, verdict semantics, limits, billing, webhooks. |
TDQS
Scored across 10 tools
Each tool targets a distinct operation: single vs batch evaluation, rule set management, evaluation retrieval, and account/plan information. Even the plan-related tools (whoami, get_usage, list_plans) are cleanly separated by identity, usage, and catalog purposes.
Most tools follow a clear verb_noun pattern: list_rule_sets, get_rule_set, create_rule_set, get_evaluation, list_evaluations. evaluate, evaluate_batch, and whoami are minor deviations but still verb-first and predictable.
Ten tools is well-scoped for an LLM evaluation service: two evaluation paths, three rule set operations, two evaluation lookup operations, and three account/plan tools. Each tool earns its place without redundancy.
The evaluation lifecycle is well covered: evaluate, retrieve, and list evaluations, plus list/get/create rule sets. Update and delete for rule sets are missing, which is a minor gap, but agents can work around it by creating new rule sets.