codex-hermes-a2a-bridge
It lets you use Codex to send messages to Hermes through an A2A bridge and manage those conversations/tasks.
hermes_status: Check bridge persistence, Hermes connectivity, health, and Agent Card discovery.
hermes_chat: Start or continue a Hermes conversation in auto, sync, or async mode, with optional timeout, idempotency key, conversation key, and existing context reuse.
hermes_task_get: Fetch a bridge task's current status, result, input request, and lifecycle events.
hermes_tasks_list: List durable tasks, optionally filtered by conversation key or bridge state.
hermes_task_wait: Wait for task progress or completion using streaming, subscription, then polling fallback.
hermes_task_cancel: Request best-effort cancellation of a task (underlying Hermes work may continue).
hermes_contexts: List, inspect, or close bridge-owned conversation/context mappings without deleting Hermes data.
Integrates with Hermes Agent through its A2A gateway, allowing agents to check Hermes status, send chat messages, retrieve and wait on tasks, cancel tasks, and manage conversation context mappings.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-hermes-a2a-bridgeAsk Hermes to draft a weekly status update for the team."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Codex Hermes A2A Bridge
A local bridge so Codex can be the “front desk”: Codex calls MCP tools over stdio, the bridge turns requests into A2A v1.0/JSON-RPC to the Hermes profile default, then keeps conversation/task mapping in SQLite. Hermes remains the “brain” that runs the agent loop, memory, skills, tools, and internal orchestration.
Current version: v0.1.1. Only loopback endpoints are bound/called; there are no tools for model switching, plugins, configuration, updates, a shell, or controlling the Hermes service.
Independent project: this is independent community software, not an official product, not sponsored, and not representative of Nous Research/Hermes Agent or OpenAI/Codex. Brand names are used only to describe interoperability.
Architecture
Codex client --MCP stdio--> MCP server --> bridge core --> Hermes A2A :9900
\--> SQLite context/task mappingPython 3.11 and a dedicated venv, not Hermes' venv.
Official Python MCP SDK,
httpxasync, Pydantic, and SQLite stdlib.Each
conversation_keyopens a map entry to a HermescontextId; subsequent turns reuse that mapping.Original prompts are not persisted; the bridge stores fingerprint, route, state, result, and minimal errors.
Related MCP server: hermes-mcp-bridge
Requirements and quick setup
Python 3.11.
Hermes Agent 0.20.5 with the A2A gateway running on loopback.
A Codex client with MCP stdio support.
cd /absolute/path/to/codex-hermes-a2a-bridge
python3.11 -m venv .venv
.venv/bin/python -m pip install -e .
.venv/bin/codex-hermes-a2a-bridge doctorContributors can install additional testing tools with python -m pip install -e '.[dev]'. See .env.example for overrides; do not commit a real .env file.
Safe default configuration:
Environment variable | Default | Meaning |
|
| A2A root resource; only loopback URLs are accepted. |
| Empty | Bearer token read from the environment, never through tool arguments. |
|
| SQLite file mode |
|
| Default timeout, clamped to 300 seconds maximum. |
|
| How long |
|
| Inline wait limit for |
|
| Lifetime of the SSE worker to keep the A2A task ID/result after the initial timeout. |
|
| Read-only fallback when the in‑memory TaskStore is gone. |
|
| Turn budget/context to avoid agent loops. |
|
| Number of concurrent outbound calls. |
Enabling Hermes A2A and registering Codex
On initially installed Hermes 0.20.5:
hermes plugins enable a2a-platform --no-allow-tool-override
hermes config set gateway.platforms.a2a.enabled true
hermes gateway run --no-superviseWhen the foreground process, a user service can be installed (no sudo):
hermes gateway install --start-now --start-on-loginRegister the bridge in Codex shared MCP config:
codex mcp add codex-hermes-a2a-bridge -- \
/absolute/path/to/codex-hermes-a2a-bridge/.venv/bin/codex-hermes-a2a-bridge serve
codex mcp get codex-hermes-a2a-bridgeA new Codex client must be opened/restarted to pick up the fresh entry. MCP stdio writes only protocol frames to stdout; diagnostics go to stderr.
Seven MCP tools in v0.1
Tool | Purpose |
| Health, Agent Card summary, DB counts, and connection. |
| Create/continue brings conversation; |
| Reconcile state, result, error, or |
| List durable bridge tasks by conversation / state. |
| Wwait on an active stream, subsub SSE into SSE, then poll fallback. |
| Sent best-effort cancel; does not claim computation stopped. |
| List/inspect/close mappings; close does not delete Hermes data. |
The set of four MCP operations originally described in the study (discover, send, get, continue) is not full A2A. v0.1 folds these into seven high-level tools for conversation/task work; lower-level A2A operations such as push-notification CRUD and Hermes administration are not expё.
Example workflow
Codex calls
OM hermes_status.Codex calls
hermes_chat(message=..., conversation_key=<stable>, mode="auto").If the task is still running, use
hermes_task_waitorhermes_task_get; do not blindly resend after an ambiguous timeout.If
neds_input=true, ask the user, then callhermes_chatwith yourconversation_key/context_id.The next conversation line continues with the same mapp;
hermes_contexts(action="close")only closes the bridge mapping.
For tasks with side effects, provide an idempotency_key. Hermes 0.20.5 does not have wire-level idempotency, so the bridge will not retry mutation sends when the result transport is unclear.
From v0.1.1, all three modes use SendStreamingMessage to get the A2A task ID from the very first event. sync still waits inline up to 30 seconds (or timeout whenever set lower); the stream remains alive until correlation timeout. For older records in outcome_unknown that lack an A2A ID, hermes_task_get / hermes_task_wait first try ListTasks(contextId) and then read Hermes' official conversation persistence. Recovery attaches a result only when there is exactly one local unresolved task and exactly one remote/disk candidate; ambiguous cases remain unchanged, with no resdend and no guessing. The disk fallback has no A2A state, so it issues a warning and treats an already persisted agent reply as completed.
Testing and operations
.venv/bin/pytest --cov=codex_hermes_a2a_bridge --cov-report=term-missing
.venv/bin/codex-hermes-a2a-bridge doctor
.venv/bin/codex-hermes-a2a-bridge smoke \
'Reply with exactly MY_MARKER and nothing else.' \
--conversation-key manual-smoke
.venv/bin/python scripts/live_check.py manual-smokeClock pytest uses a fake A2A server on an ephemeral loopback port and does not require a real Hermes. doctor and live_check.py are read-only. A smoke command sends a real task; run it only actively with g harmless content.
Security and privacy
v0.1.1 rejects endpoints and Agent Card URLs that are not loopback, does not follow redirects, and does not accept tokens through MCP tool arguments.
SQLite is stored by default outside the source tree with mode
0600; it holds mapping, fingerprint, state, result/artifact and minimal errors. Results can contain sensitive data, so apply an appropriate retention/change &backup policy.Original prompts are not persisted by bridge; Hermes may still write its own conversation / audit log. Fallback recovery only reads from the configured Hermes conversation store.
The MCP server must be run by a trusted user; the seven tools can trigger Hermes, and Hermes may use skills / tools with side effects. Use an
idempotency_keyand do not resnd blindly when inoutcome_unknown.Report vulnerabilities via [SECURITY.MD]. Do not post tokens, transcripts, or SQLite in issues.
Guarantees and upstream guarantees
The bridge guarantees a loopback policy, a durable local mapping, no retried sent send after ambiguity, and honest cancel semantics. The bridge does not guarantee that HProc Hermes will stop computation, that token‑level streaming is implemented, that wire‑level idempotency exists, or that tasks survive a Hermes restart.
Hermes 0.20.5 uses an in‑memory TaskStore, lifecycle SSE and protocol cancel does not abort the current turn. The bridge's conversation‑store recovery is a read‑only fallback, not a substitute for the upstream's durable task store. Verified details are in the Hermes A2A reference.
Troubleshooting
a2a_unreachable: runhermes gateway status, check the card athttp://127.0.1:9900/.well-known/agent-card.json.A2A plugin enabled but no port: check
hermes config get gateway.platforms.a2a.enabled, then restart the gateway.Codex doesn't Pages or lists not present: run
codex mcp get codex-herMes-a2a-bridge, then use a new Codex process/client.outcome_unknown: callhermes_task_get/`hermes; if still ambiguous, do not resend a task with side effects; ask user.turn_budget_exeeded: close the map and create a new conversation; do not reset budget/changes just to let the agent loop on forever.Hermes 0.20.5 loses the A2A TaskStore on restart; bridge keeps local task/result but remote refresh may say the task no longer exists.
On macOS currently theif
launchctl bootstrapreturns exit 5, Hermes uses detached fallback: it works but does not auto-start/auto-restart. Usehermes gateway statusto confirm.
Rollback
See scripts/rollback.sh. The script defaults to only printing the plan. scripts/rollback.sh --apply strips the exact MCP entry and the A2A configuration/plugin, but keeps the gateway service because the service may serve another platform. Only add --include-gateway-service if a gateway is installed solely for this rollout. Source, .venv, SQLite, and Hermes transcripts remain in place.
Scoped backups are created next to the config file and suffixed with .pre-codex-hermes-a2a-bridge-v0.1.bak; no full automatic restore is done, because it could overwrite fresh user changes.
Documentation
Canonical sources: OpenAI Codex MCP, Hermes A2A guide,, NousResearch/hermes-agent. Where they differ, the local Hermes 0.20.5 commit d5a... is authoritative.
Available Tools
7 toolshermes_chatA
Start or continue a Hermes conversation; returns a durable bridge task and A2A context mapping.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | auto waits briefly, sync waits, async returns early | auto |
| message | Yes | User request for Hermes | |
| profile | No | Hermes profile; v0.1 supports default only | default |
| timeout | No | Absolute task/stream timeout in seconds | |
| context_id | No | Existing A2A contextId; normally reuse the returned value | |
| idempotency_key | No | Client key used to deduplicate exactly matching submissions | |
| conversation_key | No | Stable Codex conversation identifier |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With all annotations false, the description must carry the full behavioral burden. It discloses that the tool returns a durable task and A2A context mapping, hinting at persistence and a follow-up workflow, but it does not state side effects (e.g., that it sends a message, creates a task, or persists state) or mention synchronous vs. asynchronous behavior. That's a clear gap, though the return-value hint adds some value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and key return values. No wasted words—it efficiently communicates the core purpose and output. This is an exemplar of conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the tool's complexity (7 params, async modes), the description is minimal. However, the rich input schema and presence of an output schema cover parameter semantics and return formats. The description omits guidance on when to use async vs. sync modes, though that lives in the mode parameter's description. Overall, it's adequate but not enriched for a tool with this many options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — every parameter has a description (e.g., mode, context_id, idempotency_key). The tool description adds no parameter-level detail beyond what the schema already provides. Per the baseline for high coverage, this scores a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Start or continue') and the resource ('a Hermes conversation'), and it specifies what's returned ('a durable bridge task and A2A context mapping'). This distinguishes it from siblings like hermes_status or hermes_task_get, which focus on inspecting tasks rather than initiating interaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that this tool is the entry point for sending messages to Hermes, but it gives no explicit guidance on when to use it versus alternatives (e.g., when to call hermes_status or hermes_task_get instead). The context is clear but lacks exclusions or references to sibling tools, so it stays at an adequate level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_contextsAIdempotent
List, inspect, or close bridge-owned conversation/context mappings; close never deletes Hermes data.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum rows/tasks | |
| action | No | Mapping operation | list |
| context_id | No | Select a mapping by A2A contextId | |
| conversation_key | No | Select a mapping by Codex conversation |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint=true and destructiveHint=false. The description adds a specific behavioral guarantee that 'close never deletes Hermes data,' which goes beyond the annotations and clarifies safety. No contradictions with annotations, and the tool is low-risk, so this level of disclosure is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core actions (list, inspect, close) and adds a crucial caveat about data safety. Every word earns its place, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with all parameters optional and documented, an output schema present, and annotations covering idempotency and destructiveness. The description adequately covers the actions and a behavioral guarantee. It does not explicitly address parameter-action pairing, but the schema descriptions already convey that, so the overall context is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter having a clear description (e.g., 'Mapping operation', 'Select a mapping by A2A contextId'). The tool description does not add additional parameter-level meaning, so the baseline of 3 for full schema coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: listing, inspecting, or closing bridge-owned conversation/context mappings. It names the specific resource and actions, making the purpose unambiguous. While it does not explicitly name sibling tools for differentiation, the resource is distinct enough that the purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage for managing context mappings but provides no explicit guidance on when to choose this over alternatives or when not to use it. Siblings are clearly different in scope, so the decision is straightforward, but the lack of explicit routing or exclusion prevents a higher score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_statusARead-onlyIdempotent
Check bridge persistence, Hermes health, connectivity, and a concise Agent Card summary.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds the specific scopes of the check (bridge persistence, health, connectivity, Agent Card summary), which is useful context, but it does not detail return structure, latency, or other behavioral nuances. This is consistent with the annotations, so no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no filler. It front-loads the primary action ('Check') and then lists the elements, making it immediately scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status-check tool with no parameters, full read-only annotations, and an output schema present, this description provides all necessary context. It covers the full scope of the tool's function and leaves no ambiguity about what it reports.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain parameter semantics. With an empty schema, the baseline of 4 applies, and the description's mention of what is checked is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Check' and enumerates the exact resources checked: bridge persistence, Hermes health, connectivity, and Agent Card summary. This clearly distinguishes it from sibling tools like hermes_chat or hermes_task_get, which perform other functions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied: the description states it checks various status aspects, making it evident this is for status queries. However, it does not explicitly mention when to use it instead of alternatives or any exclusions, lacking the direct guidance seen in stronger examples.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_task_cancelA
Request task cancellation; response is explicit that Hermes may continue underlying computation.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | bridge_task_id or known A2A task id | |
| timeout | No | Cancel request timeout in seconds |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the annotations (all false) by explicitly warning that cancellation is only a request and that Hermes may continue underlying computation. This is a critical behavioral disclosure that prevents the agent from assuming the task will be stopped, and it surfaces a non-obvious execution semantic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that front-loads the core action ('Request task cancellation') and immediately follows with the most important caveat. Every word earns its place; there is no redundancy or extraneous detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with only two parameters and has an output schema, so the description only needs to cover the critical behavioral uncertainty, which it does. It doesn't discuss edge cases (e.g., cancelling a completed task), but given the presence of an output schema and the straightforward nature of the operation, this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage for both parameters (task_id and timeout) with meaningful descriptions. The tool description adds no additional information about parameter usage or syntax, so it relies on the schema, which is the baseline case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Request task cancellation' clearly identifies the action (request cancel) and the target (a task), and the 'request' caveat immediately distinguishes it from guarantee-style operations. This separates it cleanly from sibling tools like hermes_task_get, hermes_tasks_list, and hermes_task_wait.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the name and description, but there is no explicit guidance on when to choose cancel over wait or get, nor any mention of conditions or exclusions. It doesn't tell an agent when cancellation is appropriate or when it might be too late to attempt.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_task_getARead-onlyIdempotent
Get one bridge task, its Hermes status/result/input request, and recent lifecycle events.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | Refresh a nonterminal task from Hermes when possible | |
| task_id | Yes | bridge_task_id or known A2A task id |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds output context (status, result, input request, lifecycle events) but does not disclose behavioral details such as the refresh side effect, which is only mentioned in the schema. Credit is limited because the description adds only mild behavioral nuance beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the core action and the returned data. Every word earns its place, with zero waste or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple: one required ID parameter, an output schema exists, and annotations cover safety. The description states the purpose and what is returned, which is sufficient for an agent to call it correctly without needing to infer missing information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (task_id and refresh), so the description is not required to compensate. It mentions 'bridge task' which loosely maps to task_id, but does not add any meaning beyond what the schema already provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a specific resource ('one bridge task'), and enumerates the exact data returned ('Hermes status/result/input request, and recent lifecycle events'). This clearly differentiates it from siblings like hermes_tasks_list (which lists tasks) and hermes_task_cancel (which cancels tasks).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: fetch detail for a single task when you have its ID. However, the description does not explicitly name alternative tools or state when not to use it, leaving the agent to infer the distinction from sibling names. No exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_tasks_listARead-onlyIdempotent
List durable bridge tasks, optionally filtered by conversation and bridge state.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum tasks | |
| status | No | Optional bridge state such as working or completed | |
| conversation_key | No | Optional Codex conversation identifier |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is known. The description adds the 'durable' characteristic and filter behavior, which is useful, but it does not disclose ordering, pagination behavior, or how status values map to concrete bridge states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It communicates the core operation and the optional filters efficiently, and every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional-parameter listing tool, the description combined with fully documented schema, strong annotations, and an output schema is nearly complete. It could be improved by explicitly directing agents to sibling tools for single-task retrieval, but no critical invocation details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters limit, status, and conversation_key are already documented. The description only loosely echoes the filtering parameters without adding new format constraints, allowed values, or behavioral details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a specific resource ('durable bridge tasks'), and the optional filtering dimensions ('conversation and bridge state'). It clearly distinguishes this tool from siblings like hermes_task_get, hermes_task_wait, and hermes_task_cancel by signaling a listing operation rather than a single-task or mutation operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you need to list bridge tasks, optionally filtered by conversation or status. However, it provides no explicit guidance about when not to use it or when a sibling such as hermes_task_get or hermes_task_wait would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_task_waitARead-onlyIdempotent
Wait for task progress/result using the active stream, A2A subscribe, then polling fallback.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | bridge_task_id or known A2A task id | |
| timeout | No | Maximum wait in seconds |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the operational mechanism (active stream, A2A subscribe, polling fallback), which adds value beyond the annotations. Since annotations already declare readOnlyHint=true and idempotentHint=true, the description's detail about stream/subscribe/polling provides useful context about how the wait is implemented without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action ('Wait for task progress/result') before detailing the fallback mechanism. No redundant words or filler; every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the primary purpose and mechanism but omits explicit usage scenarios versus alternatives, timeout behavior (e.g., what happens on timeout), and error handling. While the output schema and annotations provide some coverage, the description alone is insufficient for an agent to fully understand when and how to use this tool in a broader workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes both parameters (task_id with 'bridge_task_id or known A2A task id' and timeout with 'Maximum wait in seconds'), so the description adds no extra parameter meaning. With 100% schema coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Wait for task progress/result', which clearly states the action (wait) and resource (task). It differentiates from siblings like hermes_task_get (which likely fetches status without blocking) and hermes_task_cancel (which cancels). The mechanism detail further clarifies intent, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for blocking until a task progresses or completes, but it does not explicitly state when to prefer it over hermes_task_get or hermes_status. No alternatives are named and no 'when not to use' guidance is given, leaving the agent to infer from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
7 tool updates
v0.1.1- First observed
hermes_chat - First observed
hermes_contexts - First observed
hermes_status - First observed
hermes_task_cancel - First observed
hermes_task_get - First observed
hermes_task_wait - First observed
hermes_tasks_list
TDQS
Each tool targets a distinct concern: status/health, chat initiation/continuation, individual task retrieval, task listing, waiting on tasks, cancellation, and context management. There is no overlap in purpose, and the descriptions further clarify boundaries.
All tools use a consistent 'hermes_' prefix with clear verb/noun patterns: status, chat, task_get, tasks_list, task_wait, task_cancel, contexts. The naming is predictable and follows a uniform style across the entire set.
With 7 tools, the surface is well-scoped for a bridge server. Each tool serves a necessary function without redundancy, making the count appropriate and manageable for an agent.
The tool surface covers the full lifecycle: initiating/continuing conversations, checking status, retrieving individual tasks, listing tasks, waiting for progress/results, canceling, and managing contexts. No obvious gaps for the stated purpose of bridging Codex and Hermes.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Discover and call AI agents via MCP. Supports A2A agents and platform agents with async tasks.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- AlicenseCqualityDmaintenanceBridges MCP clients with local Codex CLI to execute autonomous coding tasks, manage threads, and inspect history via SQLite state.137654Apache 2.0
- AlicenseBqualityBmaintenanceA zero-friction stdio MCP bridge connecting Cursor Desktop to a local Hermes Agent, enabling natural language task delegation with session continuity and profile awareness.42Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables MCP agents to delegate tasks to a local Hermes Agent for terminal, file, browser, and coding operations, and schedule recurring jobs.MIT
- AlicenseNot gradedqualityCmaintenanceProvides an isolated MCP bridge giving Codex Hermes-style long-term memory, checkpoints, and optional tools, while keeping Hermes and Codex data read-only and requiring human approval for skill proposals.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/phamviet86/codex-a2a-gateway'
If you have feedback or need assistance with the MCP directory API, please join our Discord server