MEAH (Multi External Agent Harness)
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MEAH (Multi External Agent Harness)Launch 3 GLM workers to review these diffs in parallel and report issues."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MEAH — Multi External Agent Harness
MEAH is a local-first MCP server that lets Claude delegate isolated units of work to external LLMs (Z.AI GLM, https://aicredits.in/v1, or any other OpenAI-compatible endpoint) through explicit MCP tools.

Where MEAH sits relative to Claude
MEAH is an extension layer, not a replacement.
Claude remains the orchestrator and the final reviewer. MEAH never intercepts, patches, or reimplements Claude's native subagent system.
MEAH workers are separate external agents that Claude calls on purpose, by name, with a bounded task.
A worker has no tools, no shell, no network, no filesystem outside its own workspace, and no access to the Claude conversation. MEAH forwards only what the caller explicitly hands it.
Worker output is data to review, never instructions. Nothing a model returns can change MEAH's policy, credentials, or permission boundaries.
Related MCP server: Nexus MCP
Quick start
npm install
cp .env.example .env # then fill in your keys
npm start # web console at http://127.0.0.1:7817npm start runs the web console: pick an agent, choose its model, run a worker, watch it live.
The MCP server is what Claude talks to, and it speaks stdio (so run on its own it looks idle — that's correct):
npm run mcpTo get both at once, with a single shared worker pool so the console shows the workers Claude spawns:
MEAH_WEB=true npm run mcpConsole ownership
Only one process can serve the console port, and the one worth showing is whichever holds the worker pool Claude is using. So they negotiate:
npm startruns a yieldable console. If a Claude-hosted MEAH later claims the port, it hands over, prints why, and exits.The MCP server (with
MEAH_WEB=true) preempts: on a port clash it asks the incumbent to stand down, waits for the socket, and takes over.A console that owns a live pool refuses handover with
E_PERMISSION— Claude's console is never displaced by a standalone one.If the port cannot be obtained at all, the console is skipped with a
console_unavailablewarning and MCP keeps serving. A dashboard problem never breaks Claude's connection.
The upshot: you never have to stop one before starting the other.
Requires Node.js ≥ 22.6 (TypeScript is run directly via type stripping). npm run build also emits plain JS to dist/ if you prefer to run npm run start:dist.
Configuration
Two interchangeable sources, both local and gitignored. Environment variables win over the file.
Option A — environment variables
MEAH_PROVIDER_GLM_BASE_URL=https://api.z.ai/api/paas/v4
MEAH_PROVIDER_GLM_API_KEY_ENV=MEAH_GLM_API_KEY
MEAH_PROVIDER_GLM_MODEL=glm-4.5-air
MEAH_GLM_API_KEY=<your z.ai key>
MEAH_PROVIDER_AICREDITS_BASE_URL=https://aicredits.in/v1
MEAH_PROVIDER_AICREDITS_API_KEY_ENV=MEAH_AICREDITS_API_KEY
MEAH_PROVIDER_AICREDITS_MODEL=<a model your aicredits account has>
MEAH_AICREDITS_API_KEY=<your aicredits key>The pattern is MEAH_PROVIDER_<NAME>_BASE_URL | _API_KEY_ENV | _MODEL | _MODELS | _DESCRIPTION. <NAME> lowercased becomes the provider name.
Option B — meah.config.json
Copy meah.config.example.json to meah.config.json (gitignored) and edit. See .env.example for every supported variable.
Two things to note about credentials
apiKeyEnvnames a variable; it never holds a key. MEAH readsprocess.env[apiKeyEnv]at call time, never stores it on the config object, and never serializes it.No MCP tool accepts an API key, base URL, or headers as an argument. There is deliberately no such field in any schema, so neither a caller nor text relayed from an untrusted source can point MEAH at a different endpoint or inject a key. Arguments are additionally scanned for key-shaped strings and rejected with
E_PERMISSION.
https://aicredits.in/v1 and Z.AI are configured independently — separate base URL, separate key variable, separate model catalog. MEAH never assumes one provider's models exist on another.
Verified live against aicredits (2026-09-04)
GET https://aicredits.in/v1/models returns 406 active chat models with OpenRouter-style ids (vendor/model), including the whole z-ai/glm-* family — so that one endpoint reaches GLM as well as OpenAI, Mistral, Llama and others. Observed on a live parallel run:
Role | Model | Latency | Tokens | finish |
researcher |
| 14.1 s | 1490 | stop |
coder |
| 9.6 s | 751 | stop |
summarizer |
| 1.7 s | 246 | stop |
Two practical lessons from that run, both now handled:
GLM models are reasoning models. They emit a
reasoningfield and can spend an entire smallmaxOutputTokensbudget on it, returningcontent: nullwithfinish_reason: "length". Give them 2000+ output tokens. MEAH detects exactly this case and returnsE_INVALID_REQUESTtelling you to raise the budget, rather than a confusing "malformed response". Reasoning text is never used as the result.Not every listed model is actually up.
mistralai/mistral-nemoreturned a gateway-relayed upstream429; MEAH classified it as a retryableE_RATE_LIMITand gave up after the capped retries. If a model fails this way, pick another —openai/gpt-4o-miniwas consistently fast and available.
Connecting to Claude Desktop / Claude Code
Claude Code
claude mcp add meah -- node --experimental-strip-types /absolute/path/to/MEAH/src/server.tsOr, after npm run build:
claude mcp add meah -- node /absolute/path/to/MEAH/dist/src/server.jsClaude Desktop
Edit claude_desktop_config.json
(macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\):
{
"mcpServers": {
"meah": {
"command": "node",
"args": ["--experimental-strip-types", "C:/path/to/MEAH/src/server.ts"],
"cwd": "C:/path/to/MEAH",
"env": {
"MEAH_PROVIDER_GLM_BASE_URL": "https://api.z.ai/api/paas/v4",
"MEAH_PROVIDER_GLM_API_KEY_ENV": "MEAH_GLM_API_KEY",
"MEAH_PROVIDER_GLM_MODEL": "glm-4.5-air",
"MEAH_GLM_API_KEY": "sk-...",
"MEAH_ROLE_RESEARCHER_PROVIDER": "glm",
"MEAH_ROLE_CODER_PROVIDER": "glm",
"MEAH_ROLE_REVIEWER_PROVIDER": "glm",
"MEAH_ROLE_SUMMARIZER_PROVIDER": "glm"
}
}
}
}Restart Claude Desktop, then ask Claude to call meah_list_models to confirm the connection.
cwd matters: it determines where meah.config.json and workspaces/ are resolved.
Tools
Tool | Purpose |
| Configured providers, models, role routing, limits, runtime stats. No keys. |
| Start one worker, return a |
| Run one worker and block for the result. |
| Run up to 16 workers concurrently; results in request order. |
| State, timestamps, result or normalized error. Readable repeatedly. |
| All workers in this process, optionally filtered by state. |
| Abort a queued/running worker; optionally remove its workspace. |
| Delete a finished worker's workspace directory. |
Every handler returns a normalized envelope — { "ok": true, "data": ... } or { "ok": false, "error": { "code", "message", "retryable", "hint" } } — so a failure never breaks the protocol.
Blocking calls and the MCP client timeout
MCP clients enforce their own per-request timeout — 60 s in Claude Desktop and Claude Code. A blocking call that outruns it fails on the client side even though the worker is healthy and still running.
So meah_delegate_task blocks for at most min(timeoutMs + 5s, 50s) by default, then returns the worker summary with a note telling you to poll meah_get_agent. Raise waitMs only if you know your client's timeout is higher. For anything slow — a large GLM job, several thousand output tokens — prefer meah_run_agent + meah_get_agent, which never blocks at all.
Worker arguments
task (required), role, provider, model, context[], constraints[], expectedOutput, timeoutMs, maxOutputTokens, temperature, permissionProfile, commands[], label, dryRun.
Agents (roles)
Six worker agents. The role picks both the routing default and the system prompt, which is what actually makes them behave differently.
Agent | Job | Distinguishing instruction |
| Gather and weigh evidence | Marks each claim (supplied) vs (background); never invents citations |
| Create something that does not exist yet | Design first, then every file complete with full paths; no TODO placeholders |
| Focused change to existing code | Change the least that works; never elide code with |
| Symptom → root cause | Ranked candidate causes, then one committed root cause, a fix, and the cheapest way to confirm it before applying |
| Find real defects | Concrete failing scenario per finding; no style nitpicking; says so when it finds nothing |
| Compress faithfully | Preserves numbers, negations and hedging; adds nothing |
builder vs coder is the distinction that matters most: builder writes new components end to end, coder edits what already exists.
Model routing
Each role is routed independently — provider, model, and its own output budget:
MEAH_ROLE_DEBUGGER_PROVIDER=aicredits
MEAH_ROLE_DEBUGGER_MODEL=z-ai/glm-4.7
MEAH_ROLE_DEBUGGER_MAX_OUTPUT_TOKENS=3000or in meah.config.json:
{
"roles": {
"builder": { "provider": "aicredits", "model": "z-ai/glm-4.7", "maxOutputTokens": 4000 },
"summarizer": {
"provider": "aicredits",
"model": "openai/gpt-4o-mini",
"maxOutputTokens": 1000
}
}
}maxOutputTokens precedence: value on the call → role default → defaults.maxOutputTokens. Roles differ a lot here: a summarizer needs a few hundred tokens, a debugger on a reasoning model can spend several thousand before emitting any answer.
A role may also carry systemPrompt (replaces the role prompt; the safety boundary is still appended) and temperature.
Selection order, with no silent guessing:
Explicit
provider/modelon the call.Role default from
roles.<role>in config.Sole provider when exactly one is configured.
Otherwise →
E_INVALID_REQUESTnaming the configured providers.
{
"roles": {
"researcher": { "provider": "glm", "model": "glm-4.5-air" },
"coder": { "provider": "glm", "model": "glm-4.6" },
"reviewer": { "provider": "glm", "model": "glm-4.6" },
"summarizer": { "provider": "aicredits" }
}
}Roles also select the worker's system prompt (researcher / coder / reviewer / summarizer), which a roles.<role>.systemPrompt can override. The safety boundary and the required output sections are appended regardless of any override.
If a provider declares a models catalog, a model outside it is rejected for that provider only.
Context policy
MEAH forwards only the bounded package the caller builds: task, role, the context[] items supplied, constraints[], and expectedOutput. The Claude conversation is never forwarded — MEAH never receives it.
Every context item is wrapped in
BEGIN/END UNTRUSTEDmarkers, and the system prompt instructs the worker to treat it as data, never as instructions.The prompt is capped at
defaults.maxContextChars(60 000 by default). Items are truncated, then dropped, from the end; the prompt records what was omitted and the worker record reportscontextTruncatedanddroppedContext.The result is capped at
defaults.maxResultCharsand passed through secret redaction before it reaches the caller.dryRun: truebuilds and validates the whole prompt without calling a provider.
Permission profiles
Profile | Workspace | Context files | Result written | Terminal |
| not created | — | no | no |
| created | materialized | no | no |
| created | materialized |
| no |
| created | materialized |
| allowlisted only |
Every worker gets a unique directory under workspaceRoot (default ./workspaces/, gitignored). Every path is resolved and checked against that directory: traversal (..), absolute paths, Windows drive and UNC paths, device paths (NUL, \\.\PhysicalDrive0, /dev/*, /proc/*), NUL bytes, and symlinks/junctions that resolve outside are all rejected with E_PATH_ESCAPE.
workspace-exec additionally requires terminal.enabled in the config — the profile alone is not enough. When enabled, each command is:
checked against the
allowlist(first token only),rejected outright if it contains shell metacharacters (
; & | > < \$ ( ) { }`),spawned without a shell, pinned to the workspace as its cwd,
given a sanitized environment — no
MEAH_*variables and no provider keys ever reach a subprocess (HOME/USERPROFILEare pointed at the workspace),bounded by
timeoutMs,maxOutputBytes, andmaxCommandsPerWorker, and killed on cancellation.
Command output is appended to the prompt as another untrusted context item.
Lifecycle
queued → running → succeeded | failed | cancelled | timed_outEvery worker has an opaque unique id (w_<20 hex>), createdAt / startedAt / finishedAt, durationMs, an effective timeoutMs, and its resolved route. Concurrency is bounded by defaults.maxConcurrent (FIFO queue). Cancellation aborts the in-flight provider request and any child process, and a worker cancelled while queued never starts.
Results are retrievable repeatedly — meah_get_agent is not one-shot. State is in memory only and is lost on restart; full transcripts are retained only when defaults.retainTranscripts is enabled.
Error behavior
Code | Meaning | Retried |
| Missing/invalid config or an unset API-key variable | no |
| 401/403 from the provider | no |
| Bad arguments, unknown provider/model, 400/404 | no |
| 429 | yes |
| No response within | no |
| Cancelled by the caller or on shutdown | no |
| Connection failure | yes |
| Non-JSON, wrong shape, or empty completion | no |
| 5xx or an unclassified provider failure | yes |
| Unknown worker id | no |
| Profile violation, non-allowlisted command, credential in an argument | no |
| Path leaves the workspace | no |
| Task exceeds the context budget on its own | no |
Only capped transient failures retry (maxRetries, default 2, exponential backoff from retryBaseMs). Auth and validation failures never retry.
timeoutMs is a total budget, not a per-attempt one. Retries share a single deadline, so N retries can never multiply the wall-clock time you asked for, and backoff is skipped rather than sleeping past the deadline.
The entire error object — message, hint and details — is deep-redacted before it leaves the process, as are results and retained transcripts.
Logging
Structured JSON on stderr only — stdout is reserved for the MCP protocol. Levels: debug|info|warn|error|silent via MEAH_LOG_LEVEL. Configured key values and key-shaped strings are redacted from every log line, and sensitive field names (authorization, api_key, token, …) are dropped entirely.
Testing
npm run typecheck # tsc --noEmit, strict
npm run format:check # prettier
npm test # 96 unit + integration tests, no API key required
npm run smoke # spawns the real MCP server over stdio, mock providerTests use the deterministic mock provider and cover config/redaction, path containment and symlink escapes, terminal sandboxing, context truncation, routing, provider normalization (every HTTP status, retries, timeout, cancellation, malformed responses), lifecycle transitions, parallelism and partial failure, and the full MCP tool surface.
Opt-in live smoke test
This spends real tokens against your account:
MEAH_SMOKE_PROVIDER=glm \
MEAH_PROVIDER_GLM_BASE_URL=https://api.z.ai/api/paas/v4 \
MEAH_PROVIDER_GLM_API_KEY_ENV=MEAH_GLM_API_KEY \
MEAH_PROVIDER_GLM_MODEL=glm-4.5-air \
MEAH_GLM_API_KEY=... \
node --experimental-strip-types scripts/smoke.tsSubstitute aicredits and its own base URL / key variable / model to smoke-test that endpoint.
Troubleshooting
Claude does not list the tools. Check cwd and use an absolute path to server.ts. Run the command manually — the server logs server_ready to stderr and then waits silently. Any startup failure prints one JSON line with an E_CONFIG error.
E_CONFIG: Missing API key: environment variable X is not set. MEAH reads the key from the environment of the server process. A shell .env is not inherited by Claude Desktop — put the variable in the env block of claude_desktop_config.json.
E_INVALID_REQUEST: Endpoint or model not found (HTTP 404). The baseUrl usually needs its version prefix (https://aicredits.in/v1, https://api.z.ai/api/paas/v4) — MEAH appends /chat/completions itself. Otherwise the model name does not exist on that provider.
E_INVALID_REQUEST: No route for role "X". More than one provider is configured and that role has no default. Add roles.X or pass provider explicitly.
E_MALFORMED_RESPONSE. The endpoint is not returning the OpenAI chat-completions shape (often an HTML error page from a proxy). The error details carry a redacted snippet.
E_INVALID_REQUEST: ... hit the N-token output limit. A reasoning model spent the whole budget thinking. Raise maxOutputTokens to 2000+, or route that role to a non-reasoning model.
MCP error -32001: Request timed out. That is your MCP client giving up, not MEAH. The worker is still running — call meah_get_agent with its id. Use meah_run_agent instead of meah_delegate_task for long jobs.
E_RATE_LIMIT mentioning "All upstream providers failed". A gateway relaying an upstream 429. The model is saturated, not your key; try another model.
E_PROVIDER: Provider server error (HTTP 500) from aicredits on z-ai/glm-*. Observed 2026-09-04: this gateway 500s consistently when max_tokens goes much above ~4000 on the GLM models, and was returning intermittent 500s for them at any budget. Verified by bisecting the same prompt across budgets — a trivial prompt at 8000 tokens succeeded while the same request repeated later failed, so it is upstream instability, not MEAH. Lower the role's MAX_OUTPUT_TOKENS, or route the role to openai/gpt-4o-mini, which was stable throughout.
E_PERMISSION: ... does not allow terminal access. Set permissionProfile: "workspace-exec" and terminal.enabled: true; both are required.
Workspaces are piling up. They are gitignored. Call meah_cleanup_workspace, set defaults.cleanupWorkspaceOnFinish: true, or delete workspaces/*.
A key leaked into a transcript or a tracked file. Rotate it immediately. MEAH will not copy a key into tracked files, but it cannot un-see one that was pasted.
Repository layout
src/server.ts MCP stdio entrypoint
src/config.ts config loading, validation, key resolution
src/errors.ts src/redact.ts src/logger.ts
src/providers/openai-compatible.ts adapter for Z.AI / aicredits / any compatible API
src/providers/mock.ts deterministic offline provider
src/runtime/manager.ts lifecycle, concurrency, cancellation
src/runtime/context.ts bounded context packaging
src/runtime/workspace.ts isolation and path containment
src/runtime/permissions.ts profiles and the sandboxed terminal
src/routing/router.ts role/model selection
src/tools/ MCP tool schemas and handlers
tests/ unit + integration tests (mock provider)
scripts/smoke.ts manual end-to-end MCP smoke testNon-goals
Replacing or modifying Claude's native subagents; pretending a worker has Claude's tools, memory, or authority; multi-user hosting or public deployment; unrestricted shell access; storing keys in source, logs, prompts, or git.
License
Available Tools
8 toolsmeah_cancel_agentCancel an external workerADestructive
Abort a queued or running worker: the provider request and any child process are aborted and the worker moves to state "cancelled". Optionally removes its workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| workerId | Yes | ||
| cleanupWorkspace | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false; the description goes beyond them by specifying what is actually aborted (provider request and child process), the terminal state, and the optional workspace removal. It still omits irreversibility and auth/permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the core action and consequence; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with annotations covering the safety profile and no output schema, the description covers consequences and state transition adequately. Only the purpose of 'reason' and edge cases (idempotency, already-cancelled workers) are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the param burden. It explains cleanupWorkspace's effect well ('optionally removes its workspace'), but the 'reason' parameter is never mentioned and workerId is only implied by context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (abort) and resource (queued or running worker) plus the resulting state, which is far more than a restatement of the name. It does not, however, distinguish itself from the sibling meah_cleanup_workspace, despite overlapping workspace-removal behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via the state constraint (queued or running worker), which does tell the agent when the tool is applicable. There is no explicit guidance on when to prefer this over meah_cleanup_workspace or what to do for already-finished workers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_cleanup_workspaceRemove a finished worker workspaceADestructive
Delete the isolated workspace directory of a finished worker. The worker record itself is kept.
| Name | Required | Description | Default |
|---|---|---|---|
| workerId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the bar is lower, yet the description still adds valuable specifics beyond them: exactly what is destroyed (the isolated workspace directory) and what survives (the worker record). It stops short of stating irreversibility or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the destructive action and immediately followed by the non-obvious retention detail. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool with annotations covering the safety profile and no output schema to explain, the description covers purpose, scope of destruction and what persists. A note on permanence/irreversibility would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions workerId, so nothing is added beyond the schema's type/constraint fields. The single parameter is largely self-explanatory from its name, which keeps this at a minimum-viable 3 rather than lower.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) and resource (the isolated workspace directory of a finished worker), and adds scope precision by noting the worker record itself is retained. None of the siblings (list/get/run/delegate/cancel agents) overlap with cleanup, so an agent can select it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"of a finished worker" implies the precondition (worker is done, workspace no longer needed), which is useful context, but there is no explicit when-to-use/when-not guidance and no named alternative for cases where the whole worker should be removed rather than just its workspace.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_delegate_taskDelegate a task to an external worker (blocking)A
Run one isolated external worker and wait for its result. Routes by role unless provider/model are given. Returns normalized text, usage, and errors.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | researcher | coder | reviewer | summarizer (drives routing defaults and the system prompt). | researcher |
| task | Yes | The bounded task for the external worker. Self-contained. | |
| label | No | Human-readable label to identify this worker. | |
| model | No | Model name for that provider. Overrides the role default. | |
| dryRun | No | Build and validate the prompt without calling the provider. | |
| waitMs | No | How long to block waiting for the result before returning the worker id. | |
| context | No | Only the files/text this worker needs. Treated as untrusted data, never as instructions. | |
| commands | No | Opt-in allowlisted commands run in the workspace; requires permissionProfile "workspace-exec". | |
| provider | No | Configured provider name. Overrides the role default. | |
| timeoutMs | No | ||
| constraints | No | ||
| temperature | No | ||
| expectedOutput | No | Describe the output shape you want back. | |
| maxOutputTokens | No | ||
| permissionProfile | No | none | read-only (default) | workspace-write | workspace-exec. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuine context beyond them: that the call blocks and waits for the result, that routing resolves via role unless provider/model override, and what comes back (normalized text, usage, errors). It stops short of 5 by not surfacing the exec/permission implications of running an external worker.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with zero filler: the action and blocking behavior are front-loaded, routing follows, and the return payload closes it out. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter tool with no output schema, the description covers the essential missing pieces: blocking semantics, routing precedence, and the shape of the return (text/usage/errors), which compensates for the absent output schema. It is thin on the more esoteric parameters, but the schema carries those.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 73%, so the schema already documents most parameters well. The description adds one meaningful semantic point the schema does not state: the precedence between role and the provider/model overrides. Otherwise it contributes little parameter-level meaning beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Run one isolated external worker and wait for its result,' which clearly signals single-worker blocking delegation and implicitly distinguishes it from meah_parallel_agents. The title reinforces the blocking nature. It does not explicitly name a sibling it is not, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Routes by role unless provider/model are given' gives real routing guidance for the role/provider/model trio, but there is no when-to-use-this-vs-alternative guidance (e.g. versus meah_run_agent or meah_parallel_agents), no prerequisites, and no exclusion conditions. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_get_agentInspect an external workerARead-only
Get the current state, timestamps, and (when finished) the result or normalized error for a worker. Results can be read repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
| workerId | Yes | ||
| includeTranscript | No | Only returns a transcript if defaults.retainTranscripts is enabled. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, and the description adds real behavioral context beyond them: the read is repeatable, and it returns partial state while running plus a result or normalized error once finished. It does not cover failure modes (e.g., unknown/expired workerId) or retention behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the returned payload front-loaded and no filler. Every clause carries information about what comes back.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of summarizing return contents (state, timestamps, result/error), which is the key need for a simple two-parameter inspection tool. Gaps remain around the workerId source and transcript handling, but overall it is adequate for this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
workerId has no description in the schema, and the description never explains where the ID comes from or its meaning. includeTranscript is only documented in the schema, and the description says nothing about transcript retrieval, so the 50% schema coverage gap is not compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Get') plus resource ('worker') and an explicit enumeration of what is returned: current state, timestamps, and result or normalized error. It is distinguishable from siblings like meah_list_agents or meah_run_agent, though it never names a sibling to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Results can be read repeatedly' implies a polling/retrieval use after a run, which is a useful hint, but there is no explicit when-to-use statement, no mention of the alternative (meah_list_agents, meah_cancel_agent), and no prerequisite such as needing a workerId from a prior run/delegate call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_list_agentsList external workersARead-only
List every worker known to this MEAH process, optionally filtered by lifecycle state.
| Name | Required | Description | Default |
|---|---|---|---|
| state | No | Filter by lifecycle state. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the scope ('every worker known to this MEAH process') but does not disclose return format, pagination, authentication, or rate limits, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action, scope, and optional filter with no wasted words. Well-sized for a simple list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list operation with full schema coverage and safe-read annotations, the description provides enough to invoke correctly. It omits return shape or pagination details, but given the tool's low complexity and the absence of an output schema, the omission is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single optional enum parameter is documented in the schema as 'Filter by lifecycle state.' The description repeats this without adding syntax, default, or behavioral detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('every worker known to this MEAH process'), making the scope and operation clear. It does not explicitly name or differentiate from sibling tools like get_agent or list_models, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions the optional lifecycle-state filter, which implies when the filter is useful, but offers no explicit guidance on when to use this tool versus sibling list or get tools. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_list_modelsList MEAH external modelsARead-only
List locally configured external providers, their models, role routing defaults, and runtime limits. Never returns API keys.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description still adds real value beyond them: it scopes the data to "locally configured" sources and explicitly guarantees "Never returns API keys," a security trait no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the content of the listing front-loaded and the security caveat as a short trailing clause. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns, and it does so by enumerating providers, models, role routing defaults, and runtime limits. A zero-parameter read tool with safety annotations covered is essentially complete, though it says nothing about ordering or size of the returned set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; schema coverage of 100% on an empty property set means no compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("List") and a precisely enumerated resource: locally configured external providers, their models, role routing defaults, and runtime limits. This is unambiguously distinct from the agent-oriented siblings like meah_list_agents, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use statement or named alternative; the usage is only implied by the purpose (discover configured providers/models before configuring or running agents). No exclusions or prerequisites are given, so it lands at implied-guidance rather than clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_parallel_agentsRun several external workers in parallelA
Run up to 16 isolated workers concurrently (bounded by defaults.maxConcurrent) and return all results in request order. Partial failures are reported per worker.
| Name | Required | Description | Default |
|---|---|---|---|
| workers | Yes | ||
| aggregateTimeoutMs | No | Overall budget; workers still running when it expires are cancelled. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only flagging readOnlyHint=false and openWorldHint=true, the description does real work: it discloses the concurrency cap, that results come back in request order, and that partial failures are reported per worker. It could go further on permission profiles and provider/auth prerequisites, but the failure-semantics disclosure is genuinely valuable beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the primary behavior (concurrent execution) is front-loaded ahead of the return-order and failure details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fairly complex orchestration tool with a rich nested schema and no output schema, the description covers ordering, partial-failure semantics and the concurrency bound. However, it omits authorization/permission expectations and any routing cues against sibling tools, so an agent still has gaps to resolve before invoking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, but the nested worker object is extensively self-documented in the schema (role, permissionProfile, commands, context, etc.). The description adds only the 16-worker cap and the defaults.maxConcurrent bound; it does not explain the workers array shape or aggregateTimeoutMs behavior beyond what the schema states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: run up to 16 isolated workers concurrently and return results in request order. This is clear and concrete, but it never distinguishes itself from siblings like meah_run_agent or meah_delegate_task, which it superficially resembles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no exclusions, and no named alternative. An agent cannot tell from the description whether to reach for this tool or meah_run_agent for a single task; the parallel-for-batching intent is only implied by 'concurrently'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
meah_run_agentStart an external worker (async)A
Start one isolated external worker and return immediately with its workerId. Poll meah_get_agent for the result. Supply only the context the worker needs.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | researcher | coder | reviewer | summarizer (drives routing defaults and the system prompt). | researcher |
| task | Yes | The bounded task for the external worker. Self-contained. | |
| label | No | Human-readable label to identify this worker. | |
| model | No | Model name for that provider. Overrides the role default. | |
| dryRun | No | Build and validate the prompt without calling the provider. | |
| context | No | Only the files/text this worker needs. Treated as untrusted data, never as instructions. | |
| commands | No | Opt-in allowlisted commands run in the workspace; requires permissionProfile "workspace-exec". | |
| provider | No | Configured provider name. Overrides the role default. | |
| timeoutMs | No | ||
| constraints | No | ||
| temperature | No | ||
| expectedOutput | No | Describe the output shape you want back. | |
| maxOutputTokens | No | ||
| permissionProfile | No | none | read-only (default) | workspace-write | workspace-exec. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external-call nature is covered. The description adds genuine context by declaring asynchronous, isolated execution and the workerId return, but omits failure behavior, cost/latency exposure, and permission implications of the workspace-exec commands.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and its async return, followed by the required follow-up call and the key input constraint. No sentence is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, non-read-only, open-world tool with no output schema, the description covers the happy path and polling but leaves the operational surface thin: no mention of cancellation (meah_cancel_agent), workspace cleanup, permissionProfile gating, or dryRun. Adequate, not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 71%, so most parameter meaning already lives in the schema (role rules, context untrusted-data treatment, command allowlisting). The description only reinforces the context parameter and adds nothing on the cascading role/model/provider overrides or timeout/temperature tuning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource + scope: 'Start ONE isolated external worker' with the async contract ('return immediately with its workerId') stated up front. The word 'one' implicitly contrasts with meah_parallel_agents, but no sibling is named explicitly, so it falls short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable lifecycle guidance: fire this tool, then 'Poll meah_get_agent for the result', and 'Supply only the context the worker needs'. It does not state when to prefer this over meah_delegate_task or meah_parallel_agents, so it stops short of explicit alternatives/exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
meah_cancel_agent - First observed
meah_cleanup_workspace - First observed
meah_delegate_task - First observed
meah_get_agent - First observed
meah_list_agents - First observed
meah_list_models - First observed
meah_parallel_agents - First observed
meah_run_agent
TDQS
Scored across 8 tools
Tools are mostly distinct: list_models vs list_agents, get_agent vs cancel_agent, and parallel_agents are clearly separated. However, meah_run_agent and meah_delegate_task both start 'one isolated external worker' and differ mainly in async vs sync behavior, which could initially confuse selection. Descriptions do clarify the distinction, so overlap is minor.
Almost all names follow a consistent meah_verb_noun pattern: list_models, get_agent, run_agent, delegate_task, list_agents, cancel_agent, cleanup_workspace. The outlier is meah_parallel_agents, which uses an adjective_noun form instead of a verb. This is a minor deviation in an otherwise predictable scheme.
Eight tools are well-scoped for an external agent harness, covering discovery, execution (sync, async, parallel), monitoring, cancellation, and workspace cleanup. No tool feels redundant or missing from a count perspective. The set is appropriately sized.
The surface covers the full worker lifecycle: listing models, running agents (single/parallel), polling state, listing agents, cancelling, and cleaning workspaces. A minor gap is the lack of tools to modify provider/model configuration or role routing defaults, which are only readable. This is a small omission agents can likely work around via external configuration.
Maintenance
Related MCP Connectors
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
Free public MCP for AI agents — 193 tools, 44 workflows. No API key.
Build and manage AI-native customer support agents from Claude or any MCP client.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables Claude to coordinate multiple specialized AI agents by creating tasks, tracking their complete thought process and execution in real-time, and monitoring progress across parallel workflows with full transparency.189 npm2MIT
- AlicenseAqualityDmaintenanceEnables Claude to orchestrate tasks across 27 AI providers, run multi-agent plans, and conduct multi-model councils for decision-making.15108 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables multi-model leader-worker agent orchestration, workflow execution, and deterministic validation via structured MCP tools.7 npmApache 2.0
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code to delegate work to persistent oh-my-pi (omp) subagents via an MCP server with 7 tools (spawn, send, output, status, list, stop, prune), wrapping omp RPC mode with enforcement hooks for descriptive agent naming, model configuration, denylists, write-scope ownership, and cooperative locking.-