Skip to main content
Glama
phamviet86

codex-hermes-a2a-bridge

by phamviet86

Codex Hermes A2A Bridge

A local bridge so Codex can be the “front desk”: Codex calls MCP tools over stdio, the bridge turns requests into A2A v1.0/JSON-RPC to the Hermes profile default, then keeps conversation/task mapping in SQLite. Hermes remains the “brain” that runs the agent loop, memory, skills, tools, and internal orchestration.

Current version: v0.1.1. Only loopback endpoints are bound/called; there are no tools for model switching, plugins, configuration, updates, a shell, or controlling the Hermes service.

Independent project: this is independent community software, not an official product, not sponsored, and not representative of Nous Research/Hermes Agent or OpenAI/Codex. Brand names are used only to describe interoperability.

Architecture

Codex client --MCP stdio--> MCP server --> bridge core --> Hermes A2A :9900
                                      \--> SQLite context/task mapping
  • Python 3.11 and a dedicated venv, not Hermes' venv.

  • Official Python MCP SDK, httpx async, Pydantic, and SQLite stdlib.

  • Each conversation_key opens a map entry to a Hermes contextId; subsequent turns reuse that mapping.

  • Original prompts are not persisted; the bridge stores fingerprint, route, state, result, and minimal errors.

Related MCP server: hermes-mcp-bridge

Requirements and quick setup

  • Python 3.11.

  • Hermes Agent 0.20.5 with the A2A gateway running on loopback.

  • A Codex client with MCP stdio support.

cd /absolute/path/to/codex-hermes-a2a-bridge
python3.11 -m venv .venv
.venv/bin/python -m pip install -e .
.venv/bin/codex-hermes-a2a-bridge doctor

Contributors can install additional testing tools with python -m pip install -e '.[dev]'. See .env.example for overrides; do not commit a real .env file.

Safe default configuration:

Environment variable

Default

Meaning

HERMES_A2A_ENDPOINT

http://127.0.0.1:9900

A2A root resource; only loopback URLs are accepted.

HERMES_A2A_TOKEN

Empty

Bearer token read from the environment, never through tool arguments.

HERMES_BRIDGE_STATE_PATH

~/.local/state/codex-hermes-a2a-bridge/state.sqlite3

SQLite file mode 0600.

HERMES_BRIDGE_DEFAULT_TIMEOUT

60

Default timeout, clamped to 300 seconds maximum.

HERMES_BRIDGE_AUTO_WAIT

15

How long auto waits before returning the task handle.

HERMES_BRIDGE_SYNC_WAIT

30

Inline wait limit for sync; after that it returns a handle and correlation continues.

HERMES_BRIDGE_CORRELATION_TIMEOUT

300

Lifetime of the SSE worker to keep the A2A task ID/result after the initial timeout.

HERMES_A2A_CONVERSATION_DIR

~/.hermes/a2a_conversations

Read-only fallback when the in‑memory TaskStore is gone.

HERMES_BRIDGE_MAX_TURNS

5

Turn budget/context to avoid agent loops.

HERMES_BRIDGE_MAX_CONCURRENCY

4

Number of concurrent outbound calls.

Enabling Hermes A2A and registering Codex

On initially installed Hermes 0.20.5:

hermes plugins enable a2a-platform --no-allow-tool-override
hermes config set gateway.platforms.a2a.enabled true
hermes gateway run --no-supervise

When the foreground process, a user service can be installed (no sudo):

hermes gateway install --start-now --start-on-login

Register the bridge in Codex shared MCP config:

codex mcp add codex-hermes-a2a-bridge -- \
  /absolute/path/to/codex-hermes-a2a-bridge/.venv/bin/codex-hermes-a2a-bridge serve
codex mcp get codex-hermes-a2a-bridge

A new Codex client must be opened/restarted to pick up the fresh entry. MCP stdio writes only protocol frames to stdout; diagnostics go to stderr.

Seven MCP tools in v0.1

Tool

Purpose

hermes_status

Health, Agent Card summary, DB counts, and connection.

herms_chat

Create/continue brings conversation; auto, sync or async; profile default.

herms_task_get

Reconcile state, result, error, or input_required.

herms_task_list

List durable bridge tasks by conversation / state.

herms_task_wait

Wwait on an active stream, subsub SSE into SSE, then poll fallback.

herms_task_cancel

Sent best-effort cancel; does not claim computation stopped.

herms_contexts

List/inspect/close mappings; close does not delete Hermes data.

The set of four MCP operations originally described in the study (discover, send, get, continue) is not full A2A. v0.1 folds these into seven high-level tools for conversation/task work; lower-level A2A operations such as push-notification CRUD and Hermes administration are not expё.

Example workflow

  1. Codex calls OM hermes_status.

  2. Codex calls hermes_chat(message=..., conversation_key=<stable>, mode="auto").

  3. If the task is still running, use hermes_task_wait or hermes_task_get; do not blindly resend after an ambiguous timeout.

  4. If neds_input=true, ask the user, then call hermes_chat with your conversation_key / context_id.

  5. The next conversation line continues with the same mapp; hermes_contexts(action="close") only closes the bridge mapping.

For tasks with side effects, provide an idempotency_key. Hermes 0.20.5 does not have wire-level idempotency, so the bridge will not retry mutation sends when the result transport is unclear.

From v0.1.1, all three modes use SendStreamingMessage to get the A2A task ID from the very first event. sync still waits inline up to 30 seconds (or timeout whenever set lower); the stream remains alive until correlation timeout. For older records in outcome_unknown that lack an A2A ID, hermes_task_get / hermes_task_wait first try ListTasks(contextId) and then read Hermes' official conversation persistence. Recovery attaches a result only when there is exactly one local unresolved task and exactly one remote/disk candidate; ambiguous cases remain unchanged, with no resdend and no guessing. The disk fallback has no A2A state, so it issues a warning and treats an already persisted agent reply as completed.

Testing and operations

.venv/bin/pytest --cov=codex_hermes_a2a_bridge --cov-report=term-missing
.venv/bin/codex-hermes-a2a-bridge doctor
.venv/bin/codex-hermes-a2a-bridge smoke \
  'Reply with exactly MY_MARKER and nothing else.' \
  --conversation-key manual-smoke
.venv/bin/python scripts/live_check.py manual-smoke

Clock pytest uses a fake A2A server on an ephemeral loopback port and does not require a real Hermes. doctor and live_check.py are read-only. A smoke command sends a real task; run it only actively with g harmless content.

Security and privacy

  • v0.1.1 rejects endpoints and Agent Card URLs that are not loopback, does not follow redirects, and does not accept tokens through MCP tool arguments.

  • SQLite is stored by default outside the source tree with mode 0600; it holds mapping, fingerprint, state, result/artifact and minimal errors. Results can contain sensitive data, so apply an appropriate retention/change &backup policy.

  • Original prompts are not persisted by bridge; Hermes may still write its own conversation / audit log. Fallback recovery only reads from the configured Hermes conversation store.

  • The MCP server must be run by a trusted user; the seven tools can trigger Hermes, and Hermes may use skills / tools with side effects. Use an idempotency_key and do not resnd blindly when in outcome_unknown.

  • Report vulnerabilities via [SECURITY.MD]. Do not post tokens, transcripts, or SQLite in issues.

Guarantees and upstream guarantees

The bridge guarantees a loopback policy, a durable local mapping, no retried sent send after ambiguity, and honest cancel semantics. The bridge does not guarantee that HProc Hermes will stop computation, that token‑level streaming is implemented, that wire‑level idempotency exists, or that tasks survive a Hermes restart.

Hermes 0.20.5 uses an in‑memory TaskStore, lifecycle SSE and protocol cancel does not abort the current turn. The bridge's conversation‑store recovery is a read‑only fallback, not a substitute for the upstream's durable task store. Verified details are in the Hermes A2A reference.

Troubleshooting

  • a2a_unreachable: run hermes gateway status, check the card at http://127.0.1:9900/.well-known/agent-card.json.

  • A2A plugin enabled but no port: check hermes config get gateway.platforms.a2a.enabled, then restart the gateway.

  • Codex doesn't Pages or lists not present: run codex mcp get codex-herMes-a2a-bridge, then use a new Codex process/client.

  • outcome_unknown: call hermes_task_get/`hermes; if still ambiguous, do not resend a task with side effects; ask user.

  • turn_budget_exeeded: close the map and create a new conversation; do not reset budget/changes just to let the agent loop on forever.

  • Hermes 0.20.5 loses the A2A TaskStore on restart; bridge keeps local task/result but remote refresh may say the task no longer exists.

  • On macOS currently theif launchctl bootstrap returns exit 5, Hermes uses detached fallback: it works but does not auto-start/auto-restart. Use hermes gateway status to confirm.

Rollback

See scripts/rollback.sh. The script defaults to only printing the plan. scripts/rollback.sh --apply strips the exact MCP entry and the A2A configuration/plugin, but keeps the gateway service because the service may serve another platform. Only add --include-gateway-service if a gateway is installed solely for this rollout. Source, .venv, SQLite, and Hermes transcripts remain in place.

Scoped backups are created next to the config file and suffixed with .pre-codex-hermes-a2a-bridge-v0.1.bak; no full automatic restore is done, because it could overwrite fresh user changes.

Documentation

Canonical sources: OpenAI Codex MCP, Hermes A2A guide,, NousResearch/hermes-agent. Where they differ, the local Hermes 0.20.5 commit d5a... is authoritative.

Available Tools

7 tools
hermes_chatB
Destructive

Start or continue a Hermes conversation; returns a durable bridge task and A2A context mapping.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoauto waits briefly, sync waits, async returns earlyauto
originNo
messageYesUser request for Hermes
profileNoHermes profile; v0.1 supports default onlydefault
task_idNo
timeoutNoAbsolute task/stream timeout in seconds
context_idNoExisting A2A contextId; normally reuse the returned value
idempotency_keyNoClient key used to deduplicate exactly matching submissions
conversation_keyNoStable Codex conversation identifier

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety profile is covered externally. The description usefully adds that a durable task and A2A context mapping are returned (relevant for continuation), but it never explains the non-idempotent/destructive posture or how mode affects blocking behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the action and followed by the return-value clause; nothing is padded. It is dense with domain jargon ('A2A context mapping', 'bridge task') that is never unpacked, which slightly undercuts clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich 9-parameter schema, full annotations and an output schema, the description does not need to carry everything. Still, for a conversation-continuation tool it omits the practical guidance an agent most needs: how the returned task/context IDs feed back into subsequent calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, so the schema already documents mode, profile, timeout, context_id, idempotency_key and conversation_key. The description's phrase 'A2A context mapping' loosely echoes context_id but adds no syntax, defaults, or reuse rules beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start or continue a Hermes conversation') and adds a distinctive output promise ('durable bridge task and A2A context mapping'). It does not, however, name or contrast any sibling (e.g. hermes_task_wait, hermes_contexts), so an agent must infer where this sits in the family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start or continue' implies a lifecycle but gives no explicit when-to-use guidance: nothing says when to reuse task_id/context_id versus starting fresh, nor when to prefer sync over async. No alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hermes_contextsA
Idempotent

List, inspect, or close bridge-owned conversation/context mappings; close never deletes Hermes data.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows/tasks
actionNoMapping operationlist
context_idNoSelect a mapping by A2A contextId
conversation_keyNoSelect a mapping by Codex conversation

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=true and destructiveHint=false. The description adds a specific behavioral guarantee that 'close never deletes Hermes data,' which goes beyond the annotations and clarifies safety. No contradictions with annotations, and the tool is low-risk, so this level of disclosure is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core actions (list, inspect, close) and adds a crucial caveat about data safety. Every word earns its place, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with all parameters optional and documented, an output schema present, and annotations covering idempotency and destructiveness. The description adequately covers the actions and a behavioral guarantee. It does not explicitly address parameter-action pairing, but the schema descriptions already convey that, so the overall context is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter having a clear description (e.g., 'Mapping operation', 'Select a mapping by A2A contextId'). The tool description does not add additional parameter-level meaning, so the baseline of 3 for full schema coverage is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: listing, inspecting, or closing bridge-owned conversation/context mappings. It names the specific resource and actions, making the purpose unambiguous. While it does not explicitly name sibling tools for differentiation, the resource is distinct enough that the purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its usage for managing context mappings but provides no explicit guidance on when to choose this over alternatives or when not to use it. Siblings are clearly different in scope, so the decision is straightforward, but the lack of explicit routing or exclusion prevents a higher score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hermes_statusA
Read-onlyIdempotent

Check bridge persistence, Hermes health, connectivity, and a concise Agent Card summary.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds the specific scopes of the check (bridge persistence, health, connectivity, Agent Card summary), which is useful context, but it does not detail return structure, latency, or other behavioral nuances. This is consistent with the annotations, so no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no filler. It front-loads the primary action ('Check') and then lists the elements, making it immediately scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple status-check tool with no parameters, full read-only annotations, and an output schema present, this description provides all necessary context. It covers the full scope of the tool's function and leaves no ambiguity about what it reports.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description does not need to explain parameter semantics. With an empty schema, the baseline of 4 applies, and the description's mention of what is checked is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Check' and enumerates the exact resources checked: bridge persistence, Hermes health, connectivity, and Agent Card summary. This clearly distinguishes it from sibling tools like hermes_chat or hermes_task_get, which perform other functions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage is implied: the description states it checks various status aspects, making it evident this is for status queries. However, it does not explicitly mention when to use it instead of alternatives or any exclusions, lacking the direct guidance seen in stronger examples.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hermes_task_cancelA

Request task cancellation; response is explicit that Hermes may continue underlying computation.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYesbridge_task_id or known A2A task id
timeoutNoCancel request timeout in seconds

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the annotations (all false) by explicitly warning that cancellation is only a request and that Hermes may continue underlying computation. This is a critical behavioral disclosure that prevents the agent from assuming the task will be stopped, and it surfaces a non-obvious execution semantic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence that front-loads the core action ('Request task cancellation') and immediately follows with the most important caveat. Every word earns its place; there is no redundancy or extraneous detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with only two parameters and has an output schema, so the description only needs to cover the critical behavioral uncertainty, which it does. It doesn't discuss edge cases (e.g., cancelling a completed task), but given the presence of an output schema and the straightforward nature of the operation, this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for both parameters (task_id and timeout) with meaningful descriptions. The tool description adds no additional information about parameter usage or syntax, so it relies on the schema, which is the baseline case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Request task cancellation' clearly identifies the action (request cancel) and the target (a task), and the 'request' caveat immediately distinguishes it from guarantee-style operations. This separates it cleanly from sibling tools like hermes_task_get, hermes_tasks_list, and hermes_task_wait.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the name and description, but there is no explicit guidance on when to choose cancel over wait or get, nor any mention of conditions or exclusions. It doesn't tell an agent when cancellation is appropriate or when it might be too late to attempt.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hermes_task_getB

Get a task/result; optionally acknowledge a consumed result_id with expected_origin verification.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoRefresh a nonterminal task from Hermes when possible
task_idYesbridge_task_id or known A2A task id
expected_originNo
acknowledge_result_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and idempotentHint=false, so the agent already knows this is not a pure read. The description adds that acknowledgement is optional and that origin is verified, which explains the mutating profile, but it doesn't disclose what acknowledging actually does (marks consumed, irreversible?) or why the call is non-idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the core action and attaches the optional behaviour at the end. No wasted words, though the clause is packed tightly enough to be slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a tool that mixes a read with an optional state-changing acknowledgement, the definition is adequate but leaves the acknowledge side-effect and sibling selection under-explained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: task_id and refresh are documented in the schema, while expected_origin and acknowledge_result_id are not. The description partially compensates by naming both undocumented parameters and their roles (consumption, origin verification), which is the baseline for this coverage level, but it gives no format or matching semantics for expected_origin.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: 'Get a task/result' names both the action and the object, and the acknowledgement clause signals the consumption semantics. It does not, however, distinguish itself from siblings like hermes_task_wait, hermes_tasks_list, or hermes_task_cancel, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description hints at a retrieve-then-acknowledge flow ('optionally acknowledge a consumed result_id'), which implies one usage scenario, but it never states when to use this tool versus hermes_task_wait or hermes_tasks_list, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hermes_tasks_listA
Read-onlyIdempotent

List durable bridge tasks, optionally filtered by conversation and bridge state.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum tasks
statusNoOptional bridge state such as working or completed
conversation_keyNoOptional Codex conversation identifier

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is known. The description adds the 'durable' characteristic and filter behavior, which is useful, but it does not disclose ordering, pagination behavior, or how status values map to concrete bridge states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It communicates the core operation and the optional filters efficiently, and every word contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple optional-parameter listing tool, the description combined with fully documented schema, strong annotations, and an output schema is nearly complete. It could be improved by explicitly directing agents to sibling tools for single-task retrieval, but no critical invocation details are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters limit, status, and conversation_key are already documented. The description only loosely echoes the filtering parameters without adding new format constraints, allowed values, or behavioral details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a specific resource ('durable bridge tasks'), and the optional filtering dimensions ('conversation and bridge state'). It clearly distinguishes this tool from siblings like hermes_task_get, hermes_task_wait, and hermes_task_cancel by signaling a listing operation rather than a single-task or mutation operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: when you need to list bridge tasks, optionally filtered by conversation or status. However, it provides no explicit guidance about when not to use it or when a sibling such as hermes_task_get or hermes_task_wait would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hermes_task_waitA
Read-onlyIdempotent

Wait for task progress/result using the active stream, A2A subscribe, then polling fallback.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYesbridge_task_id or known A2A task id
timeoutNoMaximum wait in seconds

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the operational mechanism (active stream, A2A subscribe, polling fallback), which adds value beyond the annotations. Since annotations already declare readOnlyHint=true and idempotentHint=true, the description's detail about stream/subscribe/polling provides useful context about how the wait is implemented without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core action ('Wait for task progress/result') before detailing the fallback mechanism. No redundant words or filler; every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the primary purpose and mechanism but omits explicit usage scenarios versus alternatives, timeout behavior (e.g., what happens on timeout), and error handling. While the output schema and annotations provide some coverage, the description alone is insufficient for an agent to fully understand when and how to use this tool in a broader workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes both parameters (task_id with 'bridge_task_id or known A2A task id' and timeout with 'Maximum wait in seconds'), so the description adds no extra parameter meaning. With 100% schema coverage, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with 'Wait for task progress/result', which clearly states the action (wait) and resource (task). It differentiates from siblings like hermes_task_get (which likely fetches status without blocking) and hermes_task_cancel (which cancels). The mechanism detail further clarifies intent, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for blocking until a task progresses or completes, but it does not explicitly state when to prefer it over hermes_task_get or hermes_status. No alternatives are named and no 'when not to use' guidance is given, leaving the agent to infer from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.5.0
    • Changedhermes_chat2 fields changed
      • addedInput schema / properties / origin
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": {
        +        "type": "string"
        +      },
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Origin"
        +}
      • addedInput schema / properties / task_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Task Id"
        +}
    • Changedhermes_task_get2 fields changed
      • addedInput schema / properties / acknowledge_result_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Acknowledge Result Id"
        +}
      • addedInput schema / properties / expected_origin
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": {
        +        "type": "string"
        +      },
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Expected Origin"
        +}
  2. 7 tool updatesv0.1.1
    • First observedhermes_chat
    • First observedhermes_contexts
    • First observedhermes_status
    • First observedhermes_task_cancel
    • First observedhermes_task_get
    • First observedhermes_task_wait
    • First observedhermes_tasks_list

TDQS

A3.8/5.0

Scored across 7 tools

Disambiguation4/5

Most tools have clearly distinct purposes: chat/context management, task retrieval, cancellation, listing, waiting, and bridge status. The main possible confusion is between hermes_task_get and hermes_task_wait, since both can surface task results, but their descriptions distinguish immediate retrieval from waiting behavior.

Naming Consistency4/5

All tools use the hermes_ prefix and snake_case, which makes the set predictable overall. Minor deviations exist because some names are noun-only (hermes_chat, hermes_contexts, hermes_status) and task/tasks singular-plural usage varies.

Tool Count5/5

Seven tools is well-scoped for a bridge server focused on Hermes conversation/context and A2A task lifecycle. Each tool has a distinct operational role, and the set avoids unnecessary surface area.

Completeness5/5

The surface covers conversation start/continue, context list/inspect/close, task list/get/wait/cancel, and bridge health/status. Result acknowledgment and origin verification are included in hermes_task_get, so the core A2A bridge lifecycle appears complete.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers