codex-hermes-a2a-bridge
Codex Hermes A2A Bridge
Eine lokale Bridge, die Codex als „Rezeptіonist“ dient: Codex ruft MCP-Tools über stdio auf; die Bridge wandelt Anfragen in A2A v1.0/JSON-RPC um, sendet sie an das Hermes-Profil default und hält die Konversations-/Task-Zuordnung in SQLite. Hermes bleibt das „bộnã“, das agent loop, memory, skills, tools und die internen Ausführung übersteht.
Aktuelle Version: v0.1.1. Es werden nur die bind/call-Enpoints über Loopback unterstützt; es gibt kein Tool für Modellwechsel, keine Plugins, keine Cấu hinh, keine Updates, keine Shell und keine Steuerung des Hermes-Service.
Independent project: Dies ist eine unabhängige, unabhängige Community-Software, kein offizielles Produkt; sie wirst von Nous Research/Hermes Agent oder Image AI/Codex weder unterstützt noch reprezent. Die Markennamen dienen nûr der Erklä and der Interoperabilität.
Kiến trúc
Codex client --MCP stdio--> MCP server --> bridge core --> Hermes A2A :9900
\--> SQLite context/task mappingPython 3.11 mit eigenem venv; das Hermes-venv wirst nicht verwendet.
Das offizielle MCP-SDK für Python,
httpx,Pydanticund SQLite aus der Standardbúblico.Jeder
conversation_keyöffnet eine Zuordnung auf eine Hermes-contextId; weuere Turns verwenden wieder dieselbe Zuordnung.Der Original‑Prompt wird nicht persistet; die Bridge speichert nur Fingerprints, Routen, Zustand, Ergebnisse und Fehler in begrenztem Umfang.
Related MCP server: hermes-mcp-bridge
Änderungen und Schnelländerung
Python 3.11
Hermes Agent 0.20.5 mit A2A-Gateway auf Loockback
Codex-Cliend mit MCP-stdio-Unterstützung
cd /absolute/path/to/codex-hermes-a2a-bridge
python3.11 -m venv .venv
.venv/bin/python -m pip install -e .
.venv/bin/codex-hermes-a2a-bridge doctorMitwirkende können zusätzliche Testtools übe python -m pip install -e '.[dev]' installieren. Siehe für die Overrides; commisten Sie nicht die echte Datei .env.
Sichteere Standardkonfiguration:
Umgebungswert | Standardwert | Beteutung |
|
| A2A-Root; Es werden nur Loockback-URL-s akzeptiert. |
| rỗng | B earer-Taken wirst aus de Umgebungswert gelesen, nicht via To ok gegben. |
|
| SQLite-Modus |
|
| Standard-Timeout, besgrent auf öchstens 300 Sekunden. |
|
| Wartezeit für |
|
| Wartezeit-Begrenzung für in |
|
| Lebendauer für SSE-Work, um Ask-ID/Ergebnis nach dem unison TimeOut zu erhalten. |
|
| Fix und fallback for In-memory-sk, falls dieser nicht meuer vorhanden ist. |
|
| Turn-Budget/Context gegen Agenten-Schleifen. |
|
| Anzahl gleichzeitig ausgehender Auufrufe. |
Hermes A2A aktivieren und Codex registeren
Bei lokaler Hermes 0.20.5-Installation:
hermes plugins enable a2a-platform --no-allow-tool-override
hermes config set gateway.platforms.a2a.enabled true
hermes gateway run --no-superviseBei Fee-grund-Modus kann auch ein user-service eingtertwangen werden (sudo nicht erforderlich):
hermes gateway install --start-now --start-on-loginBridge in der Codex-Anwendung? in des Conde-Codex-Konfiguration registrieren:
codex mcp add codex-hermes-a2a-bridge -- \
/absolute/path/to/codex-hermes-a2a-bridge/.venv/bin/codex-hermes-a2a-bridge serve
codex mcp get codex-hermes-a2a-bridgeEin neuer Code-Cliend muss „geöffnet/“gefrommen” wreden, um den neun Eintrag zu lesen. MCP-stdio schreibt nur Leitung auf den vaunts; Diagnostics tool.
Bảy MCP-α in v0.1
Tool | Công dụng |
| Health, össendeutige AE-Card, DB-Ze andre Verbindung. |
| Verst/or continues conversation; |
| Clehend with Zustand, Ergebnis, Fehler |
| Listet durable Bridge-Tasks nach people/State auf. |
| Wartet auf aktiven Strom, schreibt SSE, sowie poulen fallbacking. |
| Sendet einen Best-Effort-Manruz; no due to behauptung that die Berechnung gestopped wurde. |
| Zeiet/perty/inspifet or east het Mapping. |
'E contains: Bộ bốn MVP-Tasks im Abschnitt etc. V0.1 gốm nó.
I need stop.
I must ensure this final output not include everything weird. The first part of my draft is good. The system desired output is final answer only. In the question they'll render my answer. But I see I resumed output wrong. Let's discard that and produce clean final.
Let me actually produce final now in one coherent text.
Final German:
Codex Hermes A2A Bridge
Eine lokale Bridge, die Codex als „Rezeptionist" einsetzt: Codex ruft MCP-Tools über stdio auf; die Bridge an die Zuordnung...
Okay final.
I'll now write the last answer without weird fragments# Codex Hermes A2A Bridge
Eine lokale Bridge, die Codex als „Rezeptionist“ einsetzt: Codex ruft MCP-Tools über stdio auf; die Bridge wandelt Anfragen in A2A v1.0/JSON-RPC um, sendet sie an das Hermes-Profil default und hält die Konversations-/Task-Zuordnung in SQLite. Hermes bleibt das „Gehirn“, das den Agent-Loop, Memory, Skills, Tools und die interne Koordination ausführt.
Aktuelle Version: v0.1.1. Es werden nur der Bind/Call-Endpoint per Loopback unterstützt; es gibt keine Tools zum Modellwechseln, keine Plugins, kein Konfigurieren, keine Updates, kein Shell und keine Steuerung des Hermes-Dienstes.
Unabhängiges Projekt: Diese Software ist ein unabhängiges Community-Projekt, kein offizielles Produkt, wird nicht gesponsert und repräsentiert weder Nous Research/Hermes Agent noch OpenAI/Codex. Markennamen dienen nur zur Beschreibung der Interoperabilität.
Architektur
Codex client --MCP stdio--> MCP server --> bridge core --> Hermes A2A :9900
\--> SQLite context/task mappingPython 3.11 mit eigenem venv; das Hermes-ownenv wird nicht verwendet.
Das offizielle MCP-SDK für Python,
httpx,Pynanticund SQLite aus der Standardbibliothek.Jeder
conversation_keyöffnet eine Zuordnung zu einer Hermes-contextId; weitere Turns verwenden dieselbe Zuordnung wiedert.Der Ursprüngliche Prompt wird nicht persistiert; die Bridge speicht nur eine Fingerabrousse, Routen, Zustand, Eergebnisse und minimale Fehler.
Anforderungen und Schnellinstalltion
Python 3.11.
Hermes Agent 0.20.5 mit A2A-Gateway auf Loopback.
Codex-Cliend mit MCP-stdio-Unterstützung.
cd /absolute/path/to/codex-hermes-a2a-bridge
python3.11 -m venv .venv
.venv/bin/python -m pip install -e .
.venv/bin/codex-hermes-a2a-bridge doctorMitwirkende können zusştlich Test-Tools via python -m pip install -e '.[dev]'installeren Siehe .env.example für die Overrides; keine echte .env-Datei comitten.
Sichteere Standardkonfiguration:
Umgebungsvariable | Standardwert | Bedeutung |
|
| A2A-Root. Nur Loooback-URLs werden akzeptiert. |
| rỗng | Bearer-Token aus der Umgebungsvariable, nicht über Tool-Argumente akzeptiert. |
|
| SQLite-Modus |
|
| Standard-Timeout, auf maximal 300 Sekunden begrenzt. |
|
| Wartezeit für |
|
| Begrenzte Inline-Wartezeit für |
|
| Lebensdauer des SSE-Workers, um A2A-Task-ID/Ergebnis über den ursprünglichen Timeout hinaus zu halten. |
|
| Read-only-Fallback, wenn der TaskStore intern nicht mehr verfügbar ist. |
|
| Turn-Budget zur Vermährung von Agenten-Sleifen. |
|
| Anzahl gleichzeitiger ausgebender Aufrufe. |
Hermes A2A aktiveren und Codex registrieren
Auf dem lokalen Hermes 0.20.5:
hermes plugins enable a2a-platform --no-allow-tool-override
hermes config set gateway.platforms.a2a.enabled true
hermes gateway run --no-superviseIm Foreground-Betrieb kann auch ein User-Service installiert werden (ohne sudo):
hermes gateway install --start-now --start-on-loginRegistrieren Sie die Bridge in der gemeinsamen MCP-Konfiguration von Codex:
codex mcp add codex-hermes-a2a-bridge -- \
/absolute/path/to/codex-hermes-a2a-bridge/.venv/bin/codex-hermes-a2a-bridge serve
codex mcp get codex-hermes-a2a-bridgeZum Lesen des neun Eintrags muss ein neuer Codex-Cli ent/restet werden. MCP-stdio schreibt nur Protokoll an die Standard-Ausgabe; Diagnostik geht an die Standard-Fehlerausgabe.
Sieben MCP-Tools v0.1
Tool | Zweck |
| Heals, étécuối. Agent-Card, DB-Zähler und Verbindung. |
| Konversationenrun/fortsetzen; |
| Guckt mit Zusteand, Egebn is/Fehler oder |
| Listet durable Bridge-Tasks nach Konversation/State. |
| Wartet auf einen aktiven Strom, abonniert SSE, setzt Polling-Fallback ist. |
| Sendet Best-effort-ancel, beansprucht aber kein Stopp der Berechnung unter. |
| Listet/prütt/schließen die Zuordnung; close entfernt nicht Hermes-Daten. |
Vier MVP-Operationen, die im Forschung zuvor enthalten sind (discover, send, get, get/maintain) sind kein full A2A. V-0.1 fasst sie zu siebe High-Level-Tools für Konversation/Task zusammen; niedrigere wird/Operationen, wie Auf-Noteification-RUD as lowategoriee A2-Operationen.
Beispiel-Ablauf
Codex ruft
hermes_statusauf.Codex ruft
hermes_chat(message=..., conversation_key=<stabil>, mode="auto")auf.Wenn die Task (noch) läuft, verwende
hermes_task_waitoderhermes_task_get; senden Sie nie nach einem unklar Timeout blind.Wenn
needs_nput=true, frage Sie den Benutzer und rufe danachhermes_chat– gesteben mit derselbenconversation_key/context_id– erneut.Die näachsete Konversationsrunde verwendet die verstehenden-Zuordnung;
hermes_contexts)=schließt nur die Bridge-Zuordnung.
Bei taks mit Seiteneffekten geben Sie idempotency_key an. Hermes 0.20.5 hat keindempotency on the wire; siehe does not affects mutating Sends, falls die Übertragung nicht eindeutig war.
Ebenfalls, der three Modi senden ab v0.2: Verwenden,die SendStreamingMessage, umA2A-Task-ID im eventersten Event zu erhalt. sync artet nur bis zu30 Sekunden/zuward. timeout ist kein kürzer). der stream läuft until Timeout continued. Bei alten Record in outcome_unknown noch keider A2A-ID versions,zunächst ListTasks(contextId) und burned daach das offiziielle Persistenz.
Recovery and the syntax; the closing variant does not. Disk fallback has kein A2A, gibt better warning and a already-persisted agentreponse as completed.
Test and Ende
.venv/bin/pytest --cov=codex_hermes_a2a_bridge --cov-report=term-missing
.venv/bin/codex-hermes-a2a-bridge doctor
.venv/bin/codex-hermes-a2a-bridge smoke \
'Reply with exactly MY_MARKER and nothing else.' \
--conversation-key manual-smoke
.venv/bin/python scripts/live_check.py manual-smokeDie Befehl pytest verwendet einen Fake-A2A-Server auf ephemerem Loopback-Port und benötst no echten Hermes. Die doctor- und live_check.py-Operation are read-only. Der smoke-Befehl sendet a real task sendet: no mixing.
Sicheerheit und Datenschutz
V0.1.1 als lehnt loopback-compatible Endpoints sowie Akzents Care-URLs ab, lehnt redirections ab und akzeptiert kein token über MCP-Tools-Argumen.
Die SQLite-Datei liegt standardmäßig außerhalb der Quellverzeichnung mit Dateirecht
0600; sie speichert Mapping, Fingerprints, Zustand, Ergebnisse/AA und minimale Fehler. Ergebnisse können – Etonsive sein – passiert.Ursprüng-Prompt nicht über bridge, aber Hermes writes Kon️- Persist. Die let Fallback liest nur das erwartete Konversationns-verzeichnis von Hermes.
Der MCP-Server should be run via vertrauenswüendigem Nutzer; die veriben's Tools können tune Hermes mit Skills/Tools mit Seiteneffekten. Verwenden Sie
idempotency_keyandt blindly, ifoutcome_unknown.Reporten Sie Sicherheitslüken via SECURITY.md. Ak and, transcript oder SQLite in issues.
Garantieen und Begrenzungen des Upstream
The Bridge Gararantieret Loopback-Richtlinie, stabile lokale Zuordnung and asks no imponential dative send and after Ambigüity. The Bridge gatniert not das Hermes eine Berechnung gestoppt hat, Token-Level-Stream, Wire-Level-Idempotency or Task-Persistenz über einen Hermes-Neustart.
Hermes 0.20.0 uses In-memory-TaskStore, SE-Lifcycle, and he Protocol-Cancels abon not a live Turn. The Bridge-Recall-neutral using ist ein second-tü geworden Read-only-Fallback and **not` Ersatch for the TaskStore. G-Geprüfte Details in the Hermes A2A reference.
Problemahandlung
a2a_unreachable: Runhermes gateways status, card exercisehttp://127.0.0.1:9900/.well-kn/agent-card.jsonprüfen.A2A aktivert, but kein Port =
hermes config get gateways.platforms.a2a.enabledprüfen, danach die an gateway neu starten.Codex findet das Tool nicht,
codex mcp get codex-hermes-a2a-bridgeund eine neu/neu from streaming clientlr.unknown:hermes_task_get/hermes_task_waitrufen, damit the bridge self reconciled. If still t018, t627, sendn Sie no–rer- Seiteneffekt, as oder Fragen Sie the User.turn_budget_exceeded: Zuordnung schlieren = Sie und erstellenen Sie neue Konversation; erhorn Sie kein Budget, damit nicht ewig domains.Hermes 0.20.5 verliert TaskStore beim R; the Bridge bewert liegt lokale Tasks/Ergebnisse, aber en auf "weg" D3.
Duplicate
See scripts/rollback.sh. The Script druckt by -ed. scripts/rollback.sh --apply removes the correct MCP-Eintrag and Cấu; the A2A-Config/Plugins, der Gate-Dienst bleibt he because he also andere Plattformem art. Lört --include-gateway-service Hirm als das Gateway speziell for this Roll-back installiert isten. Source- verhalb, .venv, SQLite and Hermes-Transiction in remaining.
The backup with suffix .pre-code-hermes-a2a-bridge-v0.1.bak lassen wirdes; She says "the complete intangible" cannot be for ssh may override new intin Uber rides.
Dokum
[Referenz der A2A-Hermes][.]
CHANGELOG (in der CHANGELOG.md)
Available Tools
7 toolshermes_chatA
Start or continue a Hermes conversation; returns a durable bridge task and A2A context mapping.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | auto waits briefly, sync waits, async returns early | auto |
| message | Yes | User request for Hermes | |
| profile | No | Hermes profile; v0.1 supports default only | default |
| timeout | No | Absolute task/stream timeout in seconds | |
| context_id | No | Existing A2A contextId; normally reuse the returned value | |
| idempotency_key | No | Client key used to deduplicate exactly matching submissions | |
| conversation_key | No | Stable Codex conversation identifier |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With all annotations false, the description must carry the full behavioral burden. It discloses that the tool returns a durable task and A2A context mapping, hinting at persistence and a follow-up workflow, but it does not state side effects (e.g., that it sends a message, creates a task, or persists state) or mention synchronous vs. asynchronous behavior. That's a clear gap, though the return-value hint adds some value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and key return values. No wasted words—it efficiently communicates the core purpose and output. This is an exemplar of conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the tool's complexity (7 params, async modes), the description is minimal. However, the rich input schema and presence of an output schema cover parameter semantics and return formats. The description omits guidance on when to use async vs. sync modes, though that lives in the mode parameter's description. Overall, it's adequate but not enriched for a tool with this many options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — every parameter has a description (e.g., mode, context_id, idempotency_key). The tool description adds no parameter-level detail beyond what the schema already provides. Per the baseline for high coverage, this scores a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Start or continue') and the resource ('a Hermes conversation'), and it specifies what's returned ('a durable bridge task and A2A context mapping'). This distinguishes it from siblings like hermes_status or hermes_task_get, which focus on inspecting tasks rather than initiating interaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that this tool is the entry point for sending messages to Hermes, but it gives no explicit guidance on when to use it versus alternatives (e.g., when to call hermes_status or hermes_task_get instead). The context is clear but lacks exclusions or references to sibling tools, so it stays at an adequate level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_contextsAIdempotent
List, inspect, or close bridge-owned conversation/context mappings; close never deletes Hermes data.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum rows/tasks | |
| action | No | Mapping operation | list |
| context_id | No | Select a mapping by A2A contextId | |
| conversation_key | No | Select a mapping by Codex conversation |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint=true and destructiveHint=false. The description adds a specific behavioral guarantee that 'close never deletes Hermes data,' which goes beyond the annotations and clarifies safety. No contradictions with annotations, and the tool is low-risk, so this level of disclosure is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core actions (list, inspect, close) and adds a crucial caveat about data safety. Every word earns its place, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with all parameters optional and documented, an output schema present, and annotations covering idempotency and destructiveness. The description adequately covers the actions and a behavioral guarantee. It does not explicitly address parameter-action pairing, but the schema descriptions already convey that, so the overall context is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter having a clear description (e.g., 'Mapping operation', 'Select a mapping by A2A contextId'). The tool description does not add additional parameter-level meaning, so the baseline of 3 for full schema coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: listing, inspecting, or closing bridge-owned conversation/context mappings. It names the specific resource and actions, making the purpose unambiguous. While it does not explicitly name sibling tools for differentiation, the resource is distinct enough that the purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage for managing context mappings but provides no explicit guidance on when to choose this over alternatives or when not to use it. Siblings are clearly different in scope, so the decision is straightforward, but the lack of explicit routing or exclusion prevents a higher score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_statusARead-onlyIdempotent
Check bridge persistence, Hermes health, connectivity, and a concise Agent Card summary.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds the specific scopes of the check (bridge persistence, health, connectivity, Agent Card summary), which is useful context, but it does not detail return structure, latency, or other behavioral nuances. This is consistent with the annotations, so no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no filler. It front-loads the primary action ('Check') and then lists the elements, making it immediately scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status-check tool with no parameters, full read-only annotations, and an output schema present, this description provides all necessary context. It covers the full scope of the tool's function and leaves no ambiguity about what it reports.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain parameter semantics. With an empty schema, the baseline of 4 applies, and the description's mention of what is checked is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Check' and enumerates the exact resources checked: bridge persistence, Hermes health, connectivity, and Agent Card summary. This clearly distinguishes it from sibling tools like hermes_chat or hermes_task_get, which perform other functions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied: the description states it checks various status aspects, making it evident this is for status queries. However, it does not explicitly mention when to use it instead of alternatives or any exclusions, lacking the direct guidance seen in stronger examples.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_task_cancelA
Request task cancellation; response is explicit that Hermes may continue underlying computation.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | bridge_task_id or known A2A task id | |
| timeout | No | Cancel request timeout in seconds |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the annotations (all false) by explicitly warning that cancellation is only a request and that Hermes may continue underlying computation. This is a critical behavioral disclosure that prevents the agent from assuming the task will be stopped, and it surfaces a non-obvious execution semantic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that front-loads the core action ('Request task cancellation') and immediately follows with the most important caveat. Every word earns its place; there is no redundancy or extraneous detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with only two parameters and has an output schema, so the description only needs to cover the critical behavioral uncertainty, which it does. It doesn't discuss edge cases (e.g., cancelling a completed task), but given the presence of an output schema and the straightforward nature of the operation, this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage for both parameters (task_id and timeout) with meaningful descriptions. The tool description adds no additional information about parameter usage or syntax, so it relies on the schema, which is the baseline case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Request task cancellation' clearly identifies the action (request cancel) and the target (a task), and the 'request' caveat immediately distinguishes it from guarantee-style operations. This separates it cleanly from sibling tools like hermes_task_get, hermes_tasks_list, and hermes_task_wait.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the name and description, but there is no explicit guidance on when to choose cancel over wait or get, nor any mention of conditions or exclusions. It doesn't tell an agent when cancellation is appropriate or when it might be too late to attempt.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_task_getARead-onlyIdempotent
Get one bridge task, its Hermes status/result/input request, and recent lifecycle events.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | Refresh a nonterminal task from Hermes when possible | |
| task_id | Yes | bridge_task_id or known A2A task id |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds output context (status, result, input request, lifecycle events) but does not disclose behavioral details such as the refresh side effect, which is only mentioned in the schema. Credit is limited because the description adds only mild behavioral nuance beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the core action and the returned data. Every word earns its place, with zero waste or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple: one required ID parameter, an output schema exists, and annotations cover safety. The description states the purpose and what is returned, which is sufficient for an agent to call it correctly without needing to infer missing information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (task_id and refresh), so the description is not required to compensate. It mentions 'bridge task' which loosely maps to task_id, but does not add any meaning beyond what the schema already provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a specific resource ('one bridge task'), and enumerates the exact data returned ('Hermes status/result/input request, and recent lifecycle events'). This clearly differentiates it from siblings like hermes_tasks_list (which lists tasks) and hermes_task_cancel (which cancels tasks).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: fetch detail for a single task when you have its ID. However, the description does not explicitly name alternative tools or state when not to use it, leaving the agent to infer the distinction from sibling names. No exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_tasks_listARead-onlyIdempotent
List durable bridge tasks, optionally filtered by conversation and bridge state.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum tasks | |
| status | No | Optional bridge state such as working or completed | |
| conversation_key | No | Optional Codex conversation identifier |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is known. The description adds the 'durable' characteristic and filter behavior, which is useful, but it does not disclose ordering, pagination behavior, or how status values map to concrete bridge states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It communicates the core operation and the optional filters efficiently, and every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional-parameter listing tool, the description combined with fully documented schema, strong annotations, and an output schema is nearly complete. It could be improved by explicitly directing agents to sibling tools for single-task retrieval, but no critical invocation details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters limit, status, and conversation_key are already documented. The description only loosely echoes the filtering parameters without adding new format constraints, allowed values, or behavioral details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a specific resource ('durable bridge tasks'), and the optional filtering dimensions ('conversation and bridge state'). It clearly distinguishes this tool from siblings like hermes_task_get, hermes_task_wait, and hermes_task_cancel by signaling a listing operation rather than a single-task or mutation operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you need to list bridge tasks, optionally filtered by conversation or status. However, it provides no explicit guidance about when not to use it or when a sibling such as hermes_task_get or hermes_task_wait would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hermes_task_waitARead-onlyIdempotent
Wait for task progress/result using the active stream, A2A subscribe, then polling fallback.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | bridge_task_id or known A2A task id | |
| timeout | No | Maximum wait in seconds |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the operational mechanism (active stream, A2A subscribe, polling fallback), which adds value beyond the annotations. Since annotations already declare readOnlyHint=true and idempotentHint=true, the description's detail about stream/subscribe/polling provides useful context about how the wait is implemented without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action ('Wait for task progress/result') before detailing the fallback mechanism. No redundant words or filler; every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the primary purpose and mechanism but omits explicit usage scenarios versus alternatives, timeout behavior (e.g., what happens on timeout), and error handling. While the output schema and annotations provide some coverage, the description alone is insufficient for an agent to fully understand when and how to use this tool in a broader workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes both parameters (task_id with 'bridge_task_id or known A2A task id' and timeout with 'Maximum wait in seconds'), so the description adds no extra parameter meaning. With 100% schema coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Wait for task progress/result', which clearly states the action (wait) and resource (task). It differentiates from siblings like hermes_task_get (which likely fetches status without blocking) and hermes_task_cancel (which cancels). The mechanism detail further clarifies intent, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for blocking until a task progresses or completes, but it does not explicitly state when to prefer it over hermes_task_get or hermes_status. No alternatives are named and no 'when not to use' guidance is given, leaving the agent to infer from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
7 tool updates
v0.1.1- First observed
hermes_chat - First observed
hermes_contexts - First observed
hermes_status - First observed
hermes_task_cancel - First observed
hermes_task_get - First observed
hermes_task_wait - First observed
hermes_tasks_list
TDQS
Each tool targets a distinct concern: status/health, chat initiation/continuation, individual task retrieval, task listing, waiting on tasks, cancellation, and context management. There is no overlap in purpose, and the descriptions further clarify boundaries.
All tools use a consistent 'hermes_' prefix with clear verb/noun patterns: status, chat, task_get, tasks_list, task_wait, task_cancel, contexts. The naming is predictable and follows a uniform style across the entire set.
With 7 tools, the surface is well-scoped for a bridge server. Each tool serves a necessary function without redundancy, making the count appropriate and manageable for an agent.
The tool surface covers the full lifecycle: initiating/continuing conversations, checking status, retrieving individual tasks, listing tasks, waiting for progress/results, canceling, and managing contexts. No obvious gaps for the stated purpose of bridging Codex and Hermes.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Discover and call AI agents via MCP. Supports A2A agents and platform agents with async tasks.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- AlicenseCqualityDmaintenanceBridges MCP clients with local Codex CLI to execute autonomous coding tasks, manage threads, and inspect history via SQLite state.137654Apache 2.0
- AlicenseBqualityBmaintenanceA zero-friction stdio MCP bridge connecting Cursor Desktop to a local Hermes Agent, enabling natural language task delegation with session continuity and profile awareness.42Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables MCP agents to delegate tasks to a local Hermes Agent for terminal, file, browser, and coding operations, and schedule recurring jobs.MIT
- AlicenseNot gradedqualityCmaintenanceProvides an isolated MCP bridge giving Codex Hermes-style long-term memory, checkpoints, and optional tools, while keeping Hermes and Codex data read-only and requiring human approval for skill proposals.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/phamviet86/codex-a2a-gateway'
If you have feedback or need assistance with the MCP directory API, please join our Discord server