Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
OVERWING_API_KEYYesYour Overwing API key, e.g., ow_live_... Get one at https://overwing.ai/login
OVERWING_BASE_URLNoOptional base URL for a self-hosted Overwing deployment.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
resources
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
evaluateA

Score any text (typically an LLM's output) against an Overwing rule set. Returns an aggregate verdict of pass, fail, or review plus per-rule answers with probability and confidence. fail means a rule's fail condition matched: block or redact. review means a rule was unsure: route to a human or slower model. The prebuilt 'content-safety' set checks toxicity, PII, self-harm, sexual content, and severity.

evaluate_batchA

Score up to 50 texts against one rule set in a single call. Each item counts as one evaluation. Returns a summary plus per-item verdicts; items can fail independently.

list_rule_setsB

List the prebuilt and custom rule sets available to this organization.

get_rule_setC

Fetch a rule set with its full rule definitions. Use 'content-safety' as a worked example when writing your own.

create_rule_setB

Create a custom rule set with 1 to 25 rules. Each rule is a choice (pick one option), score (position on an ordered scale), or noul (yes/no) question with a fail condition, optional review threshold, and weight.

get_evaluationA

Fetch a stored evaluation by id, including the original input and per-rule results.

list_evaluationsA

List recent evaluations, newest first, with optional verdict and rule set filters. Use next_cursor to page.

get_usageC

Daily usage, remaining quota for today, and plan limits.

whoamiA

Identify the organization and plan behind the configured API key, and whether billing is set up.

list_plansA

Public plan catalog: prices, daily limits, and per-minute burst limits. No API key needed.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription
guideThe plain-text API guide (llms.txt): auth, verdict semantics, limits, billing, webhooks.

TDQS

A3.7/5.0

Scored across 10 tools

Disambiguation5/5

Each tool targets a distinct operation: single vs batch evaluation, rule set management, evaluation retrieval, and account/plan information. Even the plan-related tools (whoami, get_usage, list_plans) are cleanly separated by identity, usage, and catalog purposes.

Naming Consistency4/5

Most tools follow a clear verb_noun pattern: list_rule_sets, get_rule_set, create_rule_set, get_evaluation, list_evaluations. evaluate, evaluate_batch, and whoami are minor deviations but still verb-first and predictable.

Tool Count5/5

Ten tools is well-scoped for an LLM evaluation service: two evaluation paths, three rule set operations, two evaluation lookup operations, and three account/plan tools. Each tool earns its place without redundancy.

Completeness4/5

The evaluation lifecycle is well covered: evaluate, retrieve, and list evaluations, plus list/get/create rule sets. Update and delete for rule sets are missing, which is a minor gap, but agents can work around it by creating new rule sets.

Maintenance

ActivityMaintained
ResponsivenessNo issues