okto-nexus
Okto Nexus
Local-first coordination for teams of AI agents.
Okto Nexus is an MCP server and operator hub for agents working in the same repository. It gives them durable identities, presence, messages, inboxes, handoffs, artifacts, an event log, governance controls, and a live dashboard without requiring a cloud broker.
okto-nexus serve exposes the complete hub on one port:
/mcp— MCP over streamable HTTP, authenticated as an agent;/api/v1— operator REST APIs, read-only monitor endpoints, and SSE;/— the bundled React dashboard.
The classic stdio transport remains supported. Coordination records live in one
SQLite database in WAL mode. Optional metrics, referenced workspace files, and
the derived shared.md view live outside that database.
Release fact | Value |
Package |
|
Python |
|
MCP surface | 43 tools by default; 46 with memory enabled |
MCP resources | 12 versioned reference resources |
MCP prompts | 0 |
Surface revision | 33 |
Database schema | 28 migrations, 34 tables |
Storage | local SQLite/WAL catalog + adapter-backed artifact payloads |
Contents
Related MCP server: Nexus Memory
Why Nexus
One coordination space per project. An absolute
project_rootis canonicalized and hashed into a deterministicworkspace_id. Every client that resolves the same real path joins the same workspace.Durable delivery. Messages fan out into per-recipient inbox lanes with leases, redelivery, acknowledgements, delivery status, and optional read receipts.
Single-winner work dispatch. Handoffs support atomic claim, leases, rejection, cancellation, optional verification, and optional dependency graphs.
Explicit identity and presence. Operators create identities and API keys; agents open sessions and heartbeat to remain present.
Targeted routing. Direct, capability, role, tag, broadcast, mixed, and direct-with-fallback strategies share one validated grammar.
Governed communication. Permissions, communication scopes, versioned policies, quotas, guardrails, groups, and human approval can restrict writes without exposing the control plane to agents.
Observable by design. Monotonic event IDs, cursor reads, long-poll, replay export, SSE, health aggregates, and a live dashboard expose what the team is doing.
Local-first and fail-closed. Configuration, target grammars, catalogs, workspace paths, API keys, and state transitions are validated before writes.
Token-aware MCP docs. First-use guidance remains resident; deeper reference material is available through versioned MCP resources on demand.
Install
The recommended install includes the HTTP hub, dashboard, local embedding provider, and tokenizer:
uv tool install "okto-nexus[serve]"
okto-nexus serveEquivalent with pipx:
pipx install "okto-nexus[serve]"Available extras:
Extra | Includes | Use it when |
none | stdio MCP core | You only need a lightweight local stdio server |
| FastAPI, Uvicorn, dashboard | You need HTTP without Torch/model dependencies |
| sentence-transformers | You want the local embedding provider separately |
| HTTP stack, embeddings, tokenizer | You want the complete supported hub |
| pytest, FastAPI, Uvicorn, httpx | You are developing or testing Nexus |
The published wheel and sdist contain the compiled dashboard. Node.js is only needed when rebuilding the frontend from source.
From a checkout:
git clone https://github.com/OktoLabsAI/okto-nexus.git
cd okto-nexus
uv sync --extra dev
# Complete HTTP build:
uv sync --extra serve --extra devStart the hub
okto-nexus serveDefaults:
dashboard:
http://127.0.0.1:8202/;MCP:
http://127.0.0.1:8202/mcp;data directory:
~/.okto_nexus;database:
~/.okto_nexus/nexus.db;initial workspace context: the current directory. The dashboard keeps a saved selection when present and otherwise may open the all-workspaces view.
Useful variants:
okto-nexus serve --project-root /absolute/path/to/project
okto-nexus serve --host 0.0.0.0 --port 8202
okto-nexus serve --trust-mode strict
okto-nexus serve --embedding-mode localOn first use, open Agents → New agent in the dashboard. Create one identity
per participant and copy its nxs_... key immediately: Nexus stores only the
hash and shows plaintext only at creation or regeneration. Capability and tag
values must first exist in Registry before an identity can use them.
Connect an MCP client
The dashboard generates snippets for Claude Code, Claude Desktop, Codex, Cursor, VS Code, Windsurf, and Cline. The generic streamable-HTTP URL is:
http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_MECredential extraction order is query api_key, x-api-key header, then
Authorization: Bearer. Treat client configuration containing a query key as
a secret.
Examples:
claude mcp add -t http okto-nexus \
"http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME"
codex mcp add okto-nexus \
--url "http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME"Generic JSON:
{
"mcpServers": {
"okto-nexus": {
"url": "http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME"
}
}
}Stdio
Run okto-nexus without a subcommand for stdio:
{
"mcpServers": {
"okto-nexus": {
"command": "okto-nexus",
"args": [],
"env": {
"OKTO_NEXUS_HOME": "/absolute/path/to/nexus-home"
}
}
}
}Without OKTO_NEXUS_API_KEY, stdio preserves the cooperative anonymous model.
Set that variable to an active nxs_... key to bind the process to the same
authenticated identity rules as HTTP. An invalid configured key fails closed.
Agent pre-flight
Every authenticated agent should do this on its first turn:
Call
agent_whoami(), use the returnedagent_idconsistently, and treat itsroleandcommunication.contentas the default operating contract.Call
workspace_resolve(project_root=<absolute cwd>).Call
session_open(agent_id=<you>, workspace_id=<resolved id>)and retain the returnedsession_idand one-timesession_secret.Check
inbox_count(agent_id=<you>); pull and acknowledge backlog.Anchor monitoring with
event_cursor(project_root=..., agent_id=<you>, stream="workspace").
The full procedure is available at
okto-nexus://reference/preflight.
The role guides responsibilities, operating perspective, and decision boundaries. The communication block guides tone, format, language, verbosity, structure, and agent-to-agent as well as user-facing communication. Follow both unless the user explicitly directs otherwise for the current task or interaction. That task-scoped override does not modify the Nexus profile or bypass permissions, policies, guardrails, approvals, communication scope, safety rules, or higher-priority host instructions. Capabilities are routing claims, not authorization or persona.
Direct message
{
"project_root": "/absolute/path/to/project",
"from_agent_id": "researcher",
"subject": "API findings",
"body": "The endpoint is idempotent; details are attached.",
"target": {
"strategy": "direct",
"agent_id": "implementer"
},
"from_session_id": "ses_...",
"session_secret": "..."
}The response names the resolved recipients and delivery count. The recipient uses:
inbox_count(agent_id="implementer")
inbox_pull(agent_id="implementer", session_id="ses_...", session_secret="...")
inbox_ack(agent_id="implementer", message_ids=[...],
session_id="ses_...", session_secret="...")Messages are delivered through the inbox. event_get and event_wait are
observability tools, not delivery.
Long-poll
event_wait is a snapshot when timeout_seconds is omitted, null, or 0.
Long-poll is explicit:
event_wait(
project_root="/absolute/path/to/project",
agent_id="researcher",
stream="workspace",
cursor=123,
timeout_seconds=25,
profile="summary"
)Always continue from next_cursor. The waiter uses SQLite
PRAGMA data_version plus bounded sleep polling; HTTP runs the blocking wait
in a worker thread so it does not block the shared event loop.
Architecture
MCP stdio MCP HTTP REST / SSE / SPA CLI
\ | | /
+---------------- inbound adapters and transport auth ----------------+
|
application services
identity · messages · inbox · handoffs · events · artifacts
permissions · policies · approvals · guardrails · memory · health
|
domain models and pure rules
|
+------------------------- outbound ports ---------------------------+
| SQLite repositories | files/shared.md | waiter | telemetry | embed |
+--------------------------------------------------------------------+bootstrap() resolves configuration, creates the store, applies migrations,
wires repositories, telemetry, embeddings, and approval execution, then seeds
the reserved operator, backfills the capability catalog, and creates built-in
permission/communication presets. create_server() lazily imports FastMCP and
registers the effective tools and resources. serve wraps the same composition
in FastAPI/Uvicorn; tail and admin are separate CLI adapters.
Important boundaries:
domain code contains state machines, routing, IDs, and invariants;
application services own the core coordination use cases; operator CRUD/maintenance routes may drive repositories and units of work directly;
inbound adapters translate MCP, HTTP, SSE, and CLI calls;
outbound adapters implement SQLite, files, telemetry, tokenization, embeddings, and waiting;
coordination truth is durable in SQLite; bounded process caches are implementation details, not authoritative state.
HTTP surfaces and authentication
Surface | Loopback bind | Non-loopback bind |
SPA shell/assets, | Public | Public |
REST data/control plane | Keyless operator trust | Active |
MCP | Active | Active |
EPT monitor endpoints | Scoped | Scoped |
MCP-over-HTTP connections always represent an agent; stdio may use the cooperative anonymous mode. The dashboard/REST loopback trust path represents the local operator. Browser-origin checks protect mutating operator routes, and binding beyond loopback removes keyless REST trust.
On a non-loopback bind, use the reserved operator identity's key for the
dashboard/control plane. Participant keys authenticate requests but
operator-only routes return PERMISSION_DENIED. When a store has no keys at
all, startup creates the operator key and prints its plaintext once.
Permanent agent keys can authenticate REST and MCP, but helper monitors should
receive only a short-lived ephemeral poll token (nxsept_...). EPTs are bound
to the issuing session, agent, and workspace and are accepted only as
Authorization: Bearer nxsept_...; query api_key and x-api-key are rejected
for this token type. The bearer is valid only on:
GET /api/v1/eventsandGET /api/v1/events/cursor;GET /api/v1/inbox/countandGET /api/v1/inbox/peek.
They cannot call MCP or mutate state.
Dashboard
The bundled dashboard provides:
Graph — toggle between detailed agent cards and compact activity-sized circles, with profile colours, live presence status badges, recent message flow, unread traffic, open handoffs, and claimed relationships;
Messages — inbox lanes, peer conversations, undelivered targeting outcomes, receipts, and optional semantic search;
Handoffs — a six-column Kanban including
VERIFYING, claim details, dependency state, verification, cancellation, and results;Meta-harness — an OpenAI-style chat for private or broadcast messages and handoffs, with agent filtering, aggregated acknowledgement flags, and replies/results in one timeline;
Artifacts — paginated deliverables with list/Explorer-style grid views, type-aware icons, rich previews, metadata, and managed-payload downloads;
Events — filtered event history, trace navigation, and live SSE updates;
Memory — durable memory browse/search/curation when
feature_memoryis enabled;Workspaces — sessions, analytics, and coordination health;
Agents — identities, keys, activation, roles, capabilities, metadata, colors, permissions, tags, inbound/outbound audiences, communication style, and steering;
Registry — operator-managed capability and tag vocabularies;
Policies — versioned policies and per-agent bindings;
Guardrails — groups, versioned content rules, assignments, and scrubbed denial audit;
Communication — versioned communication presets and bindings;
Approvals — pending and decided human-in-the-loop actions;
Settings — runtime-manageable settings, feature flags, retention, and database maintenance; metrics use their own header-menu panel.
Semantic search requires embedding_mode=stub or local. off returns
EMBEDDINGS_UNAVAILABLE on the REST search endpoint. The local provider
needs the embeddings extra; stub is deterministic and is intended for tests
or demonstrations.
Coordination model
Workspaces, agents, and sessions
Agents are global identities; workspaces represent canonical project roots.
Most coordination tools accept
project_root.session_openandshared_md_renderconsume a resolvedworkspace_id.An authenticated agent can update only its own profile with
agent_register, subject toidentity.update_profileandidentity.update_capabilities. Operators create identities.Authenticated discovery is reachability-scoped.
agent_listandagent_gethide unreachable peers;capability_listreturns the complete catalog but filters owner identities.workspace_listis permission-gated; absolute paths require a separate permission.Cross-workspace errors are intentionally operation-specific:
WORKSPACE_MISMATCHfor ownership guards,NOT_FOUNDfor hidden artifact or memory reads, andDEPENDENCY_NOT_FOUNDfor dependency creation.
Presence is explicit. A session is considered present while its heartbeat is
within presence_ttl_seconds. A trust-sensitive write advances the heartbeat
only when it authenticates with that session's credentials; in
trust_mode=open, a credential-free write advances no session. Read-only tools
do not heartbeat. Call session_heartbeat during long read-only or idle periods
and session_close when finished.
Messages and inboxes
message_create persists one message and resolves recipients at send time.
Each recipient gets a durable delivery row in its global inbox. The lanes are:
Lane | Meaning |
| Available to pull |
| Pulled and protected by an in-flight lease |
| Acknowledged |
| Dead-lettered after exhausting delivery claims; not redelivered automatically |
An expired delivered lease becomes pullable again. Delivery is therefore
at-least-once until acknowledgement. message_status lets the sender inspect
each recipient's lane. Every pull/redelivery emits message.delivered; ack
emits message.read. By default, ack also sends one synthetic read-receipt
message to the sender; receipts do not recursively create receipts.
Channels are organizational labels, not ACLs or delivery mechanisms. Access still intersects with permissions, policy, guardrails, and communication reachability. Message retention can remove aged messages and their deliveries, including unread or in-flight rows, so durability is bounded by configured retention and explicit database reset.
Communication intent
Choose the coordination mechanism by intended outcome: if another agent is expected to perform work or produce a deliverable, create a handoff. If the recipient only needs to know something or reply, use a message.
Handoff — executable work. Any request to execute, investigate, change, build, test, review, validate, or otherwise produce a deliverable must use
handoff_create; a direct message must not be its sole record. Target the intended assignee directly when known, or use a capability, role, tag, mixed, broadcast, or direct-with-fallback target when the first eligible claimant should own the work. The handoff is the canonical, operator-visible record of ownership, lifecycle, result, and verification. Include the objective, context, scope, constraints, and expected deliverable; add acceptance criteria, dependencies, and a verifier when applicable and the correspondingfeature_verification/feature_dagflag is enabled. Those fields are rejected while their feature is off.Broadcast message — shared alignment. Use it for shared context, decisions, announcements, discoveries, risk or blocker alerts, and general alignment across every selected reachable recipient. It informs recipients but assigns no owner and creates no task lifecycle. If anyone is expected to act, create one or more handoffs.
Direct message — conversation and informal coordination. Use it for status checks, questions, clarifications, acknowledgements, focused context exchange, and informal coordination. It may discuss an existing handoff, but new executable work requires a handoff; reference that handoff in the conversation. Use
handoff_getfor canonical lifecycle state and direct messages for contextual updates or blocker explanations.
A broadcast message is informational fan-out to many recipients. A handoff
with a broadcast target is a claim pool for one executor. If several agents
must produce independent results, create separate handoffs.
Routing
There are seven strategies:
Strategy | Descriptor | Notes |
|
| One named identity |
|
| One or any of several registered capabilities |
|
| Exact role match |
|
| Registered tag selector |
|
| Present workspace agents for messages; globally registered eligible agents for handoffs |
|
| Non-empty union of non-broadcast rules |
| direct plus | Handoffs only |
Messages support direct, capability, role, tag, broadcast, and mixed. Omitting a message target means broadcast. Handoffs require an explicit target and additionally support direct-with-fallback.
Tag selectors use AND across keys and OR across values. Rich
In/NotIn/Exists/DoesNotExist expressions are also supported.
Capability and tag names fail closed against operator-managed catalogs.
Target resolution is intersected with presence where applicable and with the caller's effective communication reach/audience. Separate enforcement layers then allow, deny, limit, or intercept the write:
per-agent permissions and recipient/rate limits;
versioned policy action rules and quotas;
content guardrails;
optional HITL approval interception.
A channel itself adds no ACL. Visibility controls who may see an item; eligibility controls who may claim it.
Event log
Events are immutable after insertion during normal operation. event_id is
globally monotonic and never reused. Retention may delete old rows, so retained
history can contain gaps.
Streams are workspace, agent, and handoff. Supported filters are
type, agent_id, task_id, handoff_id, and trace_id. Authenticated event
reads require events.read and omit actors outside the caller's communication
reach, except the caller's own and system events.
Response profiles:
Profile | Behavior |
| Safe trim: preserves contextual fields while removing empty or duplicated data; oversized event payloads can yield a follow-up hint |
| Aggressive projection: omits heavy bodies/payloads and returns follow-up hints |
| Raw debugging escape hatch |
event_get is non-blocking. event_cursor returns the current end in O(1).
event_wait long-polls only when timeout_seconds > 0. Dashboard SSE already
provides operator UI updates; agent monitoring remains MCP polling/long-poll or
the EPT read-only REST plane.
Handoffs
OPEN -> CLAIMED
OPEN -> REJECTED | CANCELLED
CLAIMED -> COMPLETED
CLAIMED -> VERIFYING when acceptance criteria exist
CLAIMED -> REJECTED claimant rejects
CLAIMED -> OPEN lease expires
VERIFYING -> COMPLETED verifier passes
VERIFYING -> CLAIMED verifier fails; executor reworksOnly one agent wins handoff_claim. A claim returns the confidential payload.
Available-list responses omit payload; handoff_get includes it only for the
claimant. Events and synthetic notifications never carry the payload.
With feature_verification=true, creation may include
acceptance_criteria and verify_by. The executor cannot verify its own
result. A failed verdict persists feedback, returns the handoff to CLAIMED,
and renews the lease.
With feature_dag=true, creation may include depends_on. The persisted
status remains OPEN, but blockedness is derived. Blocked items are excluded
from handoff_list_available and claim returns DEPENDENCY_NOT_MET until every
dependency is COMPLETED.
Artifacts and shared.md
Artifacts are published as text/json/markdown/html content or imported from a
workspace-contained file. Path containment is checked after canonical
resolution; escapes return PATH_OUTSIDE_WORKSPACE. Text content is bounded by
max_inline_bytes.
Payload bytes and free-form metadata do not live in SQLite. The database keeps
only the minimal searchable and authorization catalog. The default local
ArtifactStore adapter writes each payload and its manifest.json below
~/.okto_nexus/artifacts/<workspace>/<agent>/<artifact-id>/. Imported files
are copied into that managed location, so the artifact survives later changes
to the original workspace file. The dashboard's Artifacts screen pages
through the managed payload catalog, filters by one or more producing agents
and a production date interval, and previews or downloads the selected payload.
The publisher's effective outbound audience is frozen on the artifact. The
same audience controls artifact.created visibility and artifact_get.
Unauthorized, cross-workspace, and missing reads all return NOT_FOUND.
shared_md_render atomically rewrites a derived workspace shared.md with four
fixed sections: relevant agents/sessions, open tasks, open handoffs, and recent
events. It never becomes the source of truth and never renders handoff payloads.
Authenticated callers need shared_md.render.
Permissions, policies, guardrails, and approvals
The operator control plane manages:
permission presets and per-agent effective permissions;
capability/tag catalogs and communication scopes;
versioned attachable policies with deny-overrides and quota windows;
versioned communication guidance returned privately by
agent_whoami;agent groups and versioned guardrails assigned by scope and priority;
one-shot HITL approval decisions and operator steering.
Guardrail denials persist scrubbed metadata, not rejected raw content. When
feature_hitl=true and a policy requires approval, message_create or
handoff_create may return status: "pending_approval". Do not resend the
action; watch the returned approval ID.
Opt-in features
All seven flags default to off:
Flag | Effect |
| Accept and project |
| Enable new |
| Enable handoff acceptance criteria and verdicts |
| Enable handoff dependencies |
| Register three memory MCP tools and show Memory UI; restart required |
| Enable |
| Enable REST replay export |
Nuances:
approval history and decisions remain available when interception is off;
disabling verification or DAG blocks new contracts/dependencies, while already persisted workflow state remains enforceable and decidable;
workspace health REST remains available when the MCP health flag is off;
CLI replay export is operator-shell access and is not gated by
feature_replay;memory REST supports operator curation independently, but MCP memory tools are registered only when the flag is on at startup.
MCP surface
The default server exposes 43 tools: 42 across tool modules plus
nexus_info. Enabling feature_memory at startup adds three tools, for 46.
Both transports expose the same effective tool/resource surface for the same
configuration.
Area | Tools |
Metadata |
|
Identity/workspace |
|
Events |
|
Messages/channels |
|
Inbox |
|
Handoffs |
|
Artifacts |
|
Derived view |
|
Health |
|
Catalog |
|
Ephemeral monitor tokens |
|
Optional memory |
|
message_get, message_list, and message_wait are intentional migration
shims. They return MIGRATED with replacements in the inbox/event surface
instead of failing as unknown tools.
coordination_health stays registered but returns VALIDATION_ERROR while
feature_health is off. memory_put/get/search are absent until
feature_memory is enabled and the server restarts.
Session credentials are trust-sensitive on:
message_create;handoff_claim/complete/verify/reject/cancel;inbox_pull/ack/extend;memory_putwhen published.
In trust_mode=open they are optional but validated if supplied. In
trust_mode=strict they are required. poll_token_issue/renew/revoke always
require a valid session_id and session_secret in both modes.
Versioned MCP resources
URI | Version |
| 4 |
| 3 |
| 5 |
| 5 |
| 3 |
| 2 |
| 2 |
| 4 |
| 5 |
| 4 |
| 2 |
| 2 |
nexus_info reports package/schema/surface versions, the URI-to-version map,
and effective feature flags. Use that live metadata instead of assuming a
cached surface.
Response envelope
Successful tools return:
{
"ok": true,
"data": {}
}Failures return:
{
"ok": false,
"error": {
"code": "VALIDATION_ERROR",
"message": "Human-readable explanation"
}
}Transient SQLite lock/busy failures use DB_ERROR with
details.retryable=true.
Resident token footprint
For the 0.1.10 default surface:
Component | Characters |
Server instructions | 3,992 |
Tool docstrings | 6,563 |
Parameter schemas/descriptions | 15,720 |
Cuttable resident surface | 26,275 (~6,568 tokens) |
Total measured surface | 37,345 (~9,336 tokens) |
Deep explanations live in resources so clients load them only when needed. The current measured cuttable reduction against the frozen baseline is about 44.3%.
Configuration
For serve settings managed by the runtime catalog, effective precedence is:
CLI flag > environment variable > stored dashboard override > defaultWithin the core/stdio bootstrap there is no stored layer, so precedence is
CLI > env > default. Unknown flags, missing values, invalid enums, and
out-of-range numbers fail closed with CONFIG_ERROR. Boolean CLI flags take
an explicit value such as --feature-trace true.
Core runtime
Environment | CLI | Default | Notes |
|
|
| Runtime data directory |
|
|
| SQLite database |
|
| 5000 | Minimum 0 |
|
| 200 | Waiter interval; minimum 1 |
|
| 30 | Server wait ceiling; minimum 0 |
|
| 300 | Minimum 1 |
|
| 65536 | Inline artifact/content limit |
|
| 300 | Minimum 1 |
|
| 60 | Derived stale threshold |
|
| 1800 | Broadcast/tag presence window |
|
| 86400 | Opportunistic stale close |
|
| 1000 | Render ceiling |
|
| 1000 | Event page ceiling |
|
| 3600 | Minimum 60 |
|
|
|
|
|
|
|
|
|
|
| Sender inbox receipts |
|
|
|
|
|
|
| Operator REST/dashboard path disclosure |
|
|
| One bounded best-effort startup pass |
Retention
Environment | CLI | Default | Minimum |
|
| 30 | 0 |
|
| 14 | 0 |
|
| 7 | 0 |
|
| 30 | 7 |
Pruning removes aged events, read deliveries, closed sessions, and messages
older than the message window. Message retention is pure-age: it can remove
unread, in-flight, or parked messages and cascades to their deliveries and
embeddings. Handoffs and other non-message live rows are not age-pruned.
Metrics
Environment | CLI | Default | Notes |
|
|
|
|
|
|
| Local telemetry JSONL/state |
|
|
| Used only in beacon mode |
|
| 30 | Minimum 0 |
|
| 3600 | Minimum 60 |
Metrics are opt-in. Local mode stores bounded per-event metadata; beacon mode
publishes aggregate hourly counts only. Message bodies, prompts, workspace file
paths, coordination IDs, keys, tokens, URLs, and stack traces are excluded.
Local telemetry JSONL is not currently pruned automatically; operators must
manage those files even though metrics_retention_days is validated and
exposed in configuration.
Feature flags
Environment | CLI | Default |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
feature_memory changes tool registration and requires restart/reconnect.
The other flags gate live behavior.
Serve-only and transport-specific settings
Environment | CLI | Default | Scope |
|
| 8202 |
|
|
|
|
|
|
|
|
|
— |
|
| Initial dashboard workspace |
| — | unset | Suppress serve banner |
| — | unset | Optional stdio authenticated identity |
Operations
CLI commands
Command | Purpose |
| Start MCP HTTP, REST, SSE, and dashboard |
| Start MCP over stdio |
| Operator NDJSON follower over the event service |
| Enforce retention, optionally vacuum |
| Add keys to legacy keyless identities |
| Export a workspace replay stream as NDJSON |
Use --help on every command for the full argument grammar.
Tail
okto-nexus tail \
--project-root /absolute/path/to/project \
--agent-id observer \
--stream workspace \
--from latesttail applies per-agent event visibility. An optional --cursor-file belongs
to that consumer only; corrupted checkpoints fail closed.
Retention
Start with a dry run:
okto-nexus admin prune \
--project-root /absolute/path/to/project \
--dry-runThen execute:
okto-nexus admin prune \
--project-root /absolute/path/to/project \
--messages-keep-days 30 \
--vacuumRetention spans the whole shared store even though --project-root is
validated as the command anchor. --vacuum is the only option that compacts
freed pages on disk. There is no always-running coordination reaper;
auto_prune_on_start is a bounded opportunistic pass.
Issue legacy keys
okto-nexus admin issue-keys \
--project-root /absolute/path/to/projectThis is additive and idempotent. Existing keys are never rotated. Newly issued plaintext keys are printed once.
Replay export
okto-nexus admin export \
--project-root /absolute/path/to/project \
--trace-id trc_... \
--output nexus-events.ndjsonThe first line is a manifest; subsequent lines are raw events ordered by
event_id. CLI export is operator-shell access and remains available even
when the REST replay flag is off.
Ephemeral monitor token
An authenticated agent can call poll_token_issue, give only the returned
nxsept_... token and base URL to a read-only helper, renew it before expiry,
and revoke it on teardown. The raw token is returned only on issue/renew.
Data model and migrations
The current schema contains 34 tables:
Area | Tables |
Core coordination |
|
Settings/security/catalogs |
|
Search and memory |
|
Governance/workflows |
|
Not every table is workspace-scoped: agents and catalogs are global, inbox deliveries are keyed by recipient identity, and bindings/control-plane records have their own ownership rules.
Migrations are embedded in the package and applied in order:
001–008: core schema, close metadata, handoff payload/result, presence, durable inbox deliveries, leases, and session secrets;
009–015: API keys, settings, permissions, embeddings, tags/scopes, capability catalog, and trace IDs;
016–021: governance, approvals, verification, dependencies, memory, and health/event indexes;
022–026: versioned attachable policies, communication presets, display colors, groups/guardrails, and ephemeral poll tokens.
Nexus refuses to run against an unsupported newer schema. Runtime SQLite databases and their WAL/SHM/journal sidecars are ignored and must not be committed.
Errors
The domain/MCP contract has a closed catalog of 29 canonical codes:
Area | Codes |
Workspace |
|
Validation/identity |
|
Catalog/control plane |
|
Governance |
|
Dependencies/state |
|
Handoffs |
|
Content/path |
|
Infrastructure |
|
Compatibility/internal |
|
REST adapters also use transport-specific codes such as AUTH_FAILED,
CROSS_ORIGIN_BLOCKED, INVALID_PARAM, INVALID_WINDOW,
INVALID_SETTING, EMBEDDINGS_UNAVAILABLE, and INTERNAL.
Development
uv sync --extra dev
uv run pytest -qFor the complete HTTP/embedding environment:
uv sync --extra serve --extra dev
uv run pytest -qFrontend:
cd frontend
npm ci
npm run buildThe build writes packaged static assets under
src/okto_nexus/adapters/inbound/http/static/.
Release checks:
uv lock --check
uv build --out-dir dist/release-0.1.10
uvx twine check \
dist/release-0.1.10/okto_nexus-0.1.10-py3-none-any.whl \
dist/release-0.1.10/okto_nexus-0.1.10.tar.gzPublish only explicitly named current-version artifacts. The top-level
dist/ may contain older builds.
Project layout
src/okto_nexus/
adapters/
inbound/
cli/ serve, tail, admin
http/ FastAPI, REST, SSE, packaged SPA
mcp/ server, resources, projections, 43/46 tools
outbound/
sqlite/ repositories and migrations adapter
embedding/ optional semantic provider
file/ sharedmd/ artifact and derived-view I/O
telemetry/ tokenizer/ metrics support
application/ use cases and ports
domain/ entities, routing, state machines, policies
migrations/ 001 through 027
testing/ reusable test/replay harnesses
frontend/ React dashboard source
tests/ unit, contract, integration, replay tests
docs/design/ architecture and design records
pyproject.toml package metadata and extras
uv.lock reproducible dependency lockContributing
See CONTRIBUTING.md for the supported Python and frontend development environments, validation commands, and pull-request workflow.
Troubleshooting
event_wait returns immediately
Pass timeout_seconds > 0. Omitted, null, and zero are snapshots by design.
An agent misses broadcasts
Check that it opened a session in the correct resolved workspace and continues to heartbeat. Read-only event/inbox checks do not advance presence.
A direct peer is missing from discovery
Authenticated discovery is filtered by communication reachability. Inspect the caller's outbound and the peer's inbound communication scopes/tags in the dashboard.
Capability or tag targeting fails
Create the capability/tag value in Registry first. Catalog validation is fail-closed.
Memory tools are absent
Set feature_memory=true and restart/reconnect. Unlike live behavior flags,
this flag changes MCP tool registration.
Semantic search is unavailable
Use embedding_mode=stub or install the embeddings/serve extra and use
embedding_mode=local. off intentionally disables search.
DB_ERROR reports a lock
If details.retryable=true, retry the same call after the competing writer
commits. Avoid opening the runtime SQLite file with tools that hold long write
transactions.
The hub says another server already owns the home
serve holds {home}/nexus.serve.lock and refuses a second server using that
same home. Stop the other hub or choose a different --home; changing only
--db-path does not change the lock scope.
Workspace path is rejected
Pass an existing absolute path. Nexus canonicalizes the real path before deriving the workspace ID and validating artifact containment.
A monitor cannot mutate state
That is expected for nxsept_ tokens. They are intentionally read-only and
accepted only on the four monitor endpoints.
Cached docs appear stale
Call nexus_info and compare surface_revision and resource_versions before
reusing cached MCP reference content.
Security and limitations
Report vulnerabilities privately according to SECURITY.md. Do not include vulnerability details, credentials, tokens, or private workspace data in a public issue.
Nexus is designed for local or controlled single-tenant coordination, not as a public multi-tenant broker.
It does not terminate TLS. Put an authenticated TLS reverse proxy in front of a remote bind.
The dashboard shell and public health/info/license assets remain public; data/control REST requires authentication outside loopback.
API keys are hash-only at rest and shown once. Regeneration invalidates the old key immediately.
Session secrets are stored in plaintext in the local SQLite database; anyone who can read that file is inside the session trust boundary.
Channels are labels, not security boundaries.
SQLite is the only built-in coordination store. There is no Redis, PostgreSQL, or cloud broker adapter.
There is no always-running coordination scheduler/reaper. Expiry is checked opportunistically and retention runs manually or at startup when enabled.
The HTTP server does use worker threads/tasks for operational needs such as blocking waits and telemetry; “no scheduler” does not mean “no threads.”
Message durability is bounded by message retention and explicit reset.
Memory is experimental and changes the MCP surface at startup.
Artifact file references are confined to the canonical workspace.
Avoid committing runtime databases, sidecars, metrics output, and local secrets to version control.
Future direction
Likely extension points are additional durable-store adapters, a push-backed waiter that removes internal sleep polling, stronger remote deployment packaging, and further generated documentation from the live MCP schemas. The delivered surface already includes permissions, catalogs, communication scopes, policies, guardrails, HITL, trace correlation, verified/DAG handoffs, memory, health, replay, ephemeral poll tokens, embeddings, metrics, REST, SSE, and the dashboard.
Release notes
0.1.10 — current
Validation-hardening release: lands the external PR backlog and closes the open validation bugs (#26–#30). The MCP contract remains at surface revision 33 and the latest database migration remains 028.
Bumped the package and distribution metadata from 0.1.9 to 0.1.10.
Capped message artifact references at 20 (
MAX_ARTIFACTS) with{count, max}diagnostics on both MCP and REST; the dashboard Meta-harness route now defers to the domain cap (the legacy 10-attachment fence was removed) andVALIDATION_ERRORdetails survive the REST envelope.Rejected exact-duplicate artifact references; the error names the offending index and the duplicated reference.
Capped tag selector value lists at 20 values per key with set-based de-duplication, removing an O(n²) scan reachable from
message_create/handoff_createtag targets.Added length caps to the routing target grammar identifier fields (
agent_id/role/capability, 256 chars) and to eachdepends_onid (128 chars).Resynced the dashboard ColorPicker draft when the
valueprop changes while the component stays mounted.Consolidated the duplicated bounded-list count check into the shared domain helper
check_list_size.
0.1.9
Dashboard presentation release. The MCP contract remains at surface revision 33 and the latest database migration remains 028.
Bumped the package and distribution metadata from 0.1.8 to 0.1.9.
Added list/grid switching to the Artifacts catalog, including an Explorer-style grid with type-aware file icons, names, and payload sizes.
Collapsed Meta-harness read-receipt notifications into acknowledgement flags on their original messages by default. Grey flags identify incomplete target sets; green flags mean every target acknowledged, and their detail modal identifies pending recipients and delivery/read timestamps.
Added the
meta_harness_receipt_displayinterface setting so operators can restore separate receipt messages with thetimelinemode.
0.1.8
Artifact storage and catalog release. The MCP contract is at surface revision 33 and the latest database migration is 028.
Moved artifact payloads and free-form metadata out of SQLite into an adapter-backed store, with a local filesystem adapter by default.
Imported path-based artifacts into managed storage so they remain available independently of the original workspace file.
Added the Artifacts catalog with paginated producer/date filtering, detail previews, expanded Raw/Rich rendering, and managed-payload downloads.
Added rendered metadata fields plus safe Rich previews for Markdown, JSON, and HTML artifacts.
0.1.7
Dashboard usability and operator-interaction release. The MCP surface remains unchanged; the latest database migration is 027.
Added the Meta-harness chat with independent Message/Handoff and Private/Broadcast controls.
Combined outgoing turns, incoming agent messages, and handoff completion or rejection outcomes in a live, agent-filterable timeline.
Rendered structured chat content as readable labels and lists instead of raw JSON.
Reworked Guardrails group composition and rule authoring, including agent auto-completion, capability-scoped assignments, regex assistance, and stricter validation.
Added completed-agent responses and rejection reasons to Handoff cards and details.
Removed nested page scrolling from the dashboard shell.
0.1.6
Startup compatibility maintenance release. The MCP contract and database schema are unchanged: surface revision remains 32 and the latest migration remains 026.
Rebuilt the MCP v1 FastMCP settings model after import, eliminating the unresolved
lifespanforward-reference warning introduced bypydantic-settings2.15.Applied the compatibility path to both stdio and streamable-HTTP server construction.
Constrained the MCP SDK dependency to
mcp>=1.0,<2so adopting the breaking v2 API requires an explicit migration.Added a regression test that promotes the startup warning to an error and verifies the settings model is complete.
Verified the installed executable with local MiniLM warm-up, a live health request, and the complete test suite.
0.1.5
Live MCP smoke-test maintenance release. The MCP contract and database schema are unchanged: surface revision remains 32 and the latest migration remains 026.
Restored the real stdio smoke test on clean stores after capability registration became fail-closed.
Documented when and how to run the isolated two-agent smoke flow on Unix and Windows, including its
LIVE E2E RESULT: PASScompletion signal.Added semantic cards for structured
kindmessages in the Graph conversation drawer, with distinct read-receipt and handoff lifecycle treatments instead of raw JSON.Synchronized
pyproject.toml,uv.lock, and release commands at version 0.1.5.
0.1.4
Dashboard, observability, and coordination-guidance release. The MCP guidance contract advances to surface revision 32; the database schema is unchanged and the latest migration remains 026.
Added detailed and compact activity-based agent graph modes, profile colours, live status badges, richer relationship context, and graph conversations.
Added an event timeline, expanded event and handoff filters, message detail hydration, workspace display names, and catalog import/export workflows.
Improved dashboard views for agents, approvals, communication, events, guardrails, handoffs, messages, policies, and workspaces.
Extended observability APIs and repositories with message lookup, filtered handoff/event queries, and bucketed event timeline data.
Completed PyPI project URLs, keywords, and supported-Python classifiers, with a focused metadata regression test.
Added contributor and security policies, structured GitHub issue forms, and a dedicated documentation-assets location for product screenshots.
Clarified that agents follow the role and communication profile returned by
agent_whoamiunless the user supplies a task-scoped override, without bypassing platform governance.Defined handoffs as the required, traceable mechanism for executable work, broadcasts as shared alignment or information fan-out, and direct messages as status, clarification, and informal coordination channels.
Rewrote this README against the current CLI, transport/authentication model, dashboard, 43/46-tool surface, seven routing strategies, message retention, governance features, verification/DAG workflows, 34-table schema, 29-code error catalog, operations, testing, and limitations.
Synchronized
pyproject.tomland the root project entry inuv.lockat version 0.1.4, and aligned the package/dashboard license label with the addendum that is actually included inLICENSE.Removed tracked runtime SQLite databases and added ignore coverage for database files and journal/WAL/SHM sidecars.
Verified release archives contain the dashboard and migrations without runtime SQLite databases or sidecars.
0.1.2
Documentation-accuracy sweep at surface revision 31.
Corrected pre-flight, monitoring, heartbeat, inbox, trace, policy, HITL, artifact-audience, and target-grammar documentation.
Kept 12 deep reference resources versioned and reduced the resident cuttable surface to about 26.3k characters.
Corrected the token-reduction gate to include experimental-surface growth.
0.1.1
Hardened authenticated self-only identity/session rules and permission checks.
Added permission-gated workspace paths, shared view, health, and memory.
Added configurable aggregate metrics telemetry.
0.1.0
Added ephemeral remote-monitor tokens and read-only monitor endpoints.
Added attachable policies, communication presets, guardrail/group administration, enriched agent graph cards, and loopback trust hardening.
Made memory a registration-time experimental MCP surface.
0.0.x
Built the MCP reference-resource system, HTTP hub, live dashboard, inbox receipts, semantic search, monitoring guidance, and dashboard observability waves.
License
Copyright 2026 Okto Labs.
Okto Nexus is distributed under the Elastic License 2.0 together with the project's SaaS, competing-service, internal-use, and branding addendum. It permits internal and qualifying single-tenant use and prohibits the specified multi-tenant, white-label/OEM, competing, and large-scale internal-platform uses; applicable notices and attribution remain required.
Read the complete
LICENSE before use
or redistribution. It is also included in the source distribution and served
by the running hub at GET /api/v1/license.
Available Tools
43 toolsagent_getA
Return one agent's details incl. last_seen_at. Scoped by reachability: an agent outside your comm scope reads as NOT_FOUND, indistinguishable from a non-existent agent_id.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | The agent_id to look up. REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose a genuinely important trait: out-of-scope agents return NOT_FOUND, indistinguishable from a missing id. That prevents an agent from misreading a permission failure as non-existence. It does not state read-only status or return shape, but the output schema covers the latter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the core action leads and the surprising scoping caveat follows. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with an output schema, the description covers the one non-obvious failure mode an agent must know. Return values are handled by the output schema, so nothing critical is missing, though read-only confirmation is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single agent_id parameter is fully documented in the schema (100% coverage, including 'REQUIRED'), so the baseline is 3. The description adds no format, naming, or resolution semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return one agent's details') and names a key returned field (last_seen_at), which is crisper than the bare name. It never explicitly contrasts with agent_list or agent_whoami, so sibling differentiation is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this rather than agent_list, agent_whoami, or agent_register. The only qualifying information is a behavioral scoping caveat, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_listA
List registered agents (global), each with role/capabilities and last_seen_at. Authenticated callers see only agents their comm scope can reach (plus themselves); anonymous callers see all.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses a caller-dependent visibility rule (authenticated callers see only agents within their comm scope plus themselves; anonymous callers see all). It omits pagination/ordering behavior, but the output schema exists to cover the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the scope ('global') and the visibility rule are front-loaded rather than buried. Every clause carries information the caller needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless list tool with an output schema and no annotations, the description supplies the one thing structured data cannot: who sees which agents. Only pagination/results-size behavior is left unstated, a minor gap given the schema covers the return shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is nothing for the description to disambiguate. Baseline 4 applies for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List registered agents'), names the scope ('global'), and previews the returned fields (role/capabilities, last_seen_at). It does not explicitly distinguish itself from agent_get or agent_whoami, but 'List ... (global)' implies enumeration vs single-record retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to reach for this tool versus agent_get, agent_whoami, or agent_register, nor does it state any preconditions for calling it. Usage context is only weakly implied by the word 'List'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_registerA
Update YOUR OWN profile (role/capabilities/metadata); SELF-ONLY (else PERMISSION_DENIED). Capabilities are fail-closed against the central catalog. Docs: okto-nexus://reference/tool-docs/identity.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | Logical role, e.g. validator, worker (optional); matched exactly/case-sensitively by role-strategy targets. | |
| agent_id | Yes | The logical agent identity (stable, opaque string); agents are GLOBAL, not per-workspace. REQUIRED. | |
| metadata | No | Free-form JSON object of extra attributes stored with the agent (optional). | |
| capabilities | No | What this agent can do - used by capability routing + capability_list discovery (optional). Accepts a flag-map ({"ocr":true}), a list (["ocr","pdf"]), or a single name string. Blank names dropped. FAIL-CLOSED: every name must already exist in the central capability catalog (discover with capability_list). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does meaningful work: it discloses the authorization model (self-only, PERMISSION_DENIED on violation) and the fail-closed capability policy against the central catalog. It omits whether unspecified fields are preserved or cleared, and whether updates are idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: the action, the self-only restriction, and the capability constraint each appear in their own short clause, with the doc link last. No filler, though the clipped style borders on telegraphic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, the description leaves two real gaps for a mutation tool: whether this creates an agent that doesn't yet exist (the name implies registration) and why a required agent_id is needed if the operation is strictly self-only. For a write tool with no annotations, these should be resolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already explains role matching, the global agent_id, free-form metadata, and the three capability input shapes with the fail-closed rule. The description restates the field list and the catalog constraint without adding syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: update the caller's own agent profile, with the affected fields (role/capabilities/metadata) enumerated. It is clearly distinguishable from read-oriented siblings like agent_get, agent_list, and agent_whoami. The tool name 'agent_register' suggests creation rather than update, which slightly muddies the read of the purpose, but the prose is unambiguous about the action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition (SELF-ONLY, otherwise PERMISSION_DENIED) and routes capability discovery to capability_list. It does not, however, state when an agent should call this versus agent_whoami (read) or when a first-time registration versus a subsequent edit applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_whoamiA
Return YOUR OWN profile: agent_id, role, capabilities, metadata, permissions, effective_policies + governance, plus communication style when set. Docs: okto-nexus://reference/tool-docs/identity.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose the return surface (agent_id, role, capabilities, metadata, permissions, effective_policies + governance) and notes communication style is conditional ('when set'). It stops short of stating auth requirements or that the call is strictly non-mutating, but the self-read framing makes that clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence plus a docs pointer; no waste. The field list is dense but each item earns its place by telling the agent what it will receive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless self-profile read backed by an output schema and a docs link, this is essentially complete. The only soft gap is the absent explicit contrast with agent_get/agent_list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4; there is nothing for the description to disambiguate. The enumerated fields are extra value, though the output schema already documents returns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and a precisely scoped resource ('YOUR OWN profile'), and enumerates the fields returned. The 'YOUR OWN' framing cleanly distinguishes it from sibling agent_get/agent_list, which fetch other agents' records.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the self-scoped wording rather than stated. It never explicitly says to use agent_get for other agents or when this is preferable, so the agent must infer the routing from the sibling name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifact_getA
Retrieve an artifact by id within the workspace resolved from project_root. Reads are audience-scoped: a caller outside the frozen audience gets NOT_FOUND (indistinguishable from a missing id).
| Name | Required | Description | Default |
|---|---|---|---|
| artifact_id | Yes | The artifact_id to retrieve. REQUIRED. | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses a genuinely important trait: reads are audience-scoped and an outside caller receives NOT_FOUND that is deliberately indistinguishable from a missing id. That is valuable context an agent cannot infer from the schema. It does not state permission/auth requirements or whether the operation is read-only in a broader sense, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with the retrieval scope front-loaded and the security caveat second. Every clause earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the description covers the non-obvious audience-scoping behavior. For a two-parameter read tool this is nearly complete; only auth/permission expectations remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, and the description largely restates the project_root-to-workspace relationship already captured in the schema. It adds no format, validation, or edge-case detail beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Retrieve) and resource (artifact by id) plus the scoping mechanism (workspace resolved from project_root). It is clearly distinguishable from the sibling artifact_put, which performs the inverse operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'Retrieve an artifact by id' but there is no explicit when-to-use statement, no exclusions, and no pointer to alternatives (e.g., artifact_put for writes, or a listing tool). The audience-scoping sentence hints at a failure mode but does not guide tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifact_putC
Register a file/text/json/markdown/html artifact in the resolved workspace. Docs: okto-nexus://reference/tool-docs/artifacts.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Human-friendly name/label for the artifact (optional). | |
| path | No | Filesystem path to import into Nexus-managed artifact storage (must stay within the workspace root). Provide this OR content - at least one REQUIRED. | |
| content | No | UTF-8 content stored outside the database (bounded by max_inline_bytes; json must be well-formed). Provide this OR path - at least one REQUIRED. | |
| metadata | No | Free-form JSON object stored with the artifact (optional). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| artifact_type | Yes | Artifact classification - one of: file, text, json, markdown, html. REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about overwrite/duplicate behavior, idempotency, required permissions, size limits, or failure modes. 'Resolved workspace' hints at a dependency on workspace_resolve but does not explain it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with a useful docs pointer and no filler. Slightly terse for a write operation, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a mutation tool with no annotations and no output schema referenced in the description, so ambiguity about what a successful registration returns or does to existing artifacts is significant. The schema covers inputs and the docs link covers unspecified detail, but the description itself leaves the agent without the behavioral context needed to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (including the path-or-content exclusivity rule and the artifact_type enum-in-prose) are already documented in the schema. The description adds no parameter meaning beyond what is structured, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Register') and resource ('artifact') plus the accepted artifact types and the workspace scope. It is clearly distinguishable from artifact_get by direction, though it doesn't explicitly name that sibling as the converse operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as shared_md_render or artifact_get, and no prerequisites or exclusions. The only pointer is an external docs URI, which an agent cannot resolve inline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capability_listB
List the capability catalog merged with owners: every registered name (with description; agent_count 0 if unowned), agents scoped to your comm reach. Normalised as capability routing matches.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It does disclose some useful behavior – unowned capabilities appear with agent_count 0, and agents are filtered to the caller's 'comm reach' – but says nothing about read-only nature, pagination, or what 'Normalised as capability routing matches' implies for the results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, which is good, but the remainder is a colon-glued run-on with parentheticals and a dangling fragment ('Normalised as capability routing matches.') that obscures meaning rather than saving space. It is dense without being clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and no parameters, the description only needs to frame the tool, which it does: it names the catalog, the owner merge, and the agent_count-0 edge case. The unexplained 'comm reach' and 'routing matches' phrasing are the only real gaps for such a low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema declares zero parameters at 100% description coverage, so there is nothing for the description to compensate for. Baseline 4 applies for a parameterless list tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('List the capability catalog') and clarifies that it merges owner/agent information, so the agent can tell what data comes back. However, the purpose is diluted by undefined terms like 'comm reach' and 'capability routing matches', and it does not differentiate itself from adjacent listing tools such as agent_list or tag_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool versus alternatives like agent_list, agent_get, or tag_list, nor any stated preconditions or exclusions. Usage must be inferred purely from the word 'List'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channel_createA
Create a channel by name (idempotent; created=false if it already existed). Channels are organizational labels, not ACLs.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Channel name (REQUIRED; short topic label, max 64 chars, unique per workspace). Idempotent by name (existing -> created=false). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses idempotency, the created=false outcome, and that channels are organizational labels rather than ACLs — a genuinely non-obvious semantic. However, it says nothing about permissions, workspace-scoping errors, or failure behavior, so meaningful gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and the idempotency caveat. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and both params are fully described. The description covers semantics and idempotency; only authorization/prerequisite context is absent, which is minor for a low-risk creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents both name (max 64 chars, unique per workspace, idempotent) and project_root. The description only restates the idempotency aspect, adding no syntax or format detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (channel) with scope (by name) and an idempotency qualifier. An agent can distinguish this from channel_list and other channel siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for when to call it — creating a channel by name, tolerating repeats via idempotency. It does not name alternatives (e.g. channel_list) or state exclusions, but the usage window is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channel_listB
Return the workspace channels (general is seeded by default).
| Name | Required | Description | Default |
|---|---|---|---|
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses one useful trait (a default 'general' channel is seeded), which implies a read-only listing, but says nothing about ordering, pagination, archived/hidden channels, or whether it reflects workspace scope. For a zero-annotation tool this leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. The scoping detail is embedded compactly in the parenthetical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure need not be explained, and the single required parameter is fully described in the schema. The description covers the default-seed behavior but omits ordering and whether non-default channels are included, which is minor for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and project_root is fully documented as the workspace scope in the schema itself. The description adds no further meaning about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Return the workspace channels'. An agent can distinguish it from channel_create by the listing verb, but the description never explicitly contrasts it with the sibling create tool. Clear purpose with no explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus channel_create or how the result should be used. The parenthetical about 'general' being seeded is context, not a usage condition. The agent must infer the use case entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
coordination_healthA
Windowed coordination-health report for the workspace: aggregated ok|warn status, 7 metric blocks and thresholds. Requires feature_health and health.read.
| Name | Required | Description | Default |
|---|---|---|---|
| window | No | Report window - one of: 1h, 24h, 7d (default: 24h). | 24h |
| agent_id | No | Your agent_id for permission evaluation in open stdio mode (optional; authenticated HTTP MCP uses the API-key identity). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose key traits: the output is an aggregated ok|warn status over 7 metric blocks with thresholds, and the call requires the feature_health feature and health.read permission. It stops short of explicitly stating it is read-only/no-side-effect, which an agent could still want confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler: the first front-loads what the tool produces and its scope, the second states the hard requirement. Nothing repeats the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description needn't enumerate return fields, and it already summarizes the return shape and permission prerequisites. The main remaining gap is the absence of usage context that would tell an agent when to reach for this versus other observability-style read tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the window enum values, agent_id purpose, and project_root scope are already fully documented in the schema. The description only echoes 'windowed' without adding syntax, format, or default guidance beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (coordination-health) and scope (workspace, windowed) plus the concrete output shape (aggregated ok|warn status, 7 metric blocks). No sibling tool overlaps this domain, so differentiation isn't a concern, but it stops just short of the model 5 by describing a report rather than an explicit action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is given. The only conditional information is a prerequisite (requires feature_health and health.read), which is an auth/permission note rather than usage direction. An agent must infer from the name that this is a health-check tool to call proactively.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
event_cursorA
Return the stream's CURRENT END as a cursor (O(1), no scan) - the pre-flight monitor anchor so you see only events appended after now. Returns {cursor} (0 for empty).
| Name | Required | Description | Default |
|---|---|---|---|
| stream | Yes | Event stream to read - one of: workspace, agent, handoff. message.created, the message.delivered/message.read receipts and artifact.created are on workspace. | |
| agent_id | Yes | Your agent_id; scopes per-event visibility (you only see events you may see). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden well: it discloses the O(1)/no-scan cost profile, the anchor/monotonic semantics (only events after now), and the sentinel return (0 for empty). It never states read-only or auth/permission needs, but for a pure cursor read this is a solid disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with the anchor semantics front-loaded and cost/return details in tight parentheses. Every clause earns its place (O(1), no scan, append-after-now scoping, empty sentinel).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the return shape need not be explained, and params are fully covered by the schema. The description supplies cost and usage framing, giving an agent enough to call it; only explicit sibling routing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so stream, agent_id, and project_root are already fully documented in the schema (including scope and enum-ish stream values). The description adds no parameter-level meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Return the stream's CURRENT END as a cursor") and adds the operational role ("pre-flight monitor anchor"). It implicitly separates this from event_get/event_wait by framing it as an anchor rather than a reader, but never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Pre-flight monitor anchor so you see only events appended after now" gives a clear when-to-use context: grab this before monitoring so subsequent reads are scoped to new events. No explicit when-not or named alternatives (e.g. event_get vs event_wait) are given, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
event_getA
Read a cursor-paginated page of the event log (non-blocking). Actors outside your comm scope are omitted; yours and system events always show. Docs: okto-nexus://reference/tool-docs/events.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max events per page (optional; default 100, clamped to the server maximum, default 1000). | |
| cursor | No | Pagination cursor: the last event_id you consumed; returns event_id > cursor (optional; omit/0 = from the beginning). | |
| stream | Yes | Event stream to read - one of: workspace, agent, handoff. message.created, the message.delivered/message.read receipts and artifact.created are on workspace. | |
| filters | No | Equality filters, AND-combined (optional). Keys: type, agent_id, task_id, handoff_id, trace_id. e.g. {"type":"message.created"}. | |
| profile | No | Response size profile - one of: default, summary, full (optional; summary trims per-call tokens). | |
| agent_id | Yes | Your agent_id; scopes per-event visibility (you only see events you may see). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses non-blocking behavior, the visibility/comm-scope filtering rule, and cursor pagination. It omits auth/prerequisite requirements and any rate-limit or error behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences plus a docs pointer; the core purpose and the most surprising behavior (non-blocking, scope filtering) are front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers pagination, blocking behavior, and visibility. What remains thin is guidance on choosing among the many sibling event/inbox readers, but nothing required to invoke the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters including cursor, limit, stream, and filters. The description adds only the high-level 'cursor-paginated page' framing, not parameter-specific semantics beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (a page of the event log) plus a distinguishing behavioral trait: non-blocking. It implicitly separates itself from event_wait, but it never names that sibling or event_cursor explicitly, so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '(non-blocking)' tag implies you use this when you don't want to wait, but no when-to-use/when-not statement or explicit alternative is given. The visibility rule ('actors outside your comm scope are omitted; yours and system events always show') is useful context but is about results, not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
event_waitA
Read the event log; optionally long-poll (0/omitted/null = snapshot; >0 blocks). Scoped like event_get (actors outside your comm scope omitted). Patterns: okto-nexus://reference/monitoring.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max events per page (optional; default 100, clamped to the server maximum, default 1000). | |
| cursor | No | Pagination cursor: the last event_id you consumed; returns event_id > cursor (optional; omit/0 = from the beginning). | |
| stream | Yes | Event stream to read - one of: workspace, agent, handoff. message.created, the message.delivered/message.read receipts and artifact.created are on workspace. | |
| filters | No | Equality filters, AND-combined (optional). Keys: type, agent_id, task_id, handoff_id, trace_id. e.g. {"type":"message.created"}. | |
| profile | No | Response size profile - one of: default, summary, full (optional; summary trims per-call tokens). | |
| agent_id | Yes | Your agent_id; scopes per-event visibility (you only see events you may see). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| timeout_seconds | No | Long-poll bound in SECONDS (optional; default 0). 0/omitted/null = an immediate non-blocking snapshot; >0 OPTS IN to a BLOCKING long-poll until an event arrives or the timeout elapses. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a solid job: it discloses the visibility scoping ("actors outside your comm scope omitted"), the blocking semantics of the timeout, and points to a monitoring reference. It stops short of permissions or rate-limit context, but the key behavioral traits of a long-poll read are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the primary action and the blocking rule, with no filler. Slightly dense but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter read tool with an output schema and full schema coverage, the description supplies the missing behavioral context (scoping, blocking, monitoring reference) without needing to explain return values. Adequate for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all eight parameters are already documented, including the 0/omitted/null snapshot semantics for timeout and the limit/cursor meaning. The description merely restates the timeout rule rather than adding new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Read the event log") and disambiguates itself from the sibling event_get by noting the scoping relationship and by being the tool that offers long-polling. An agent can tell what it does, though the distinction from event_get/event_cursor is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The blocking-vs-snapshot behavior is explained, which implies when you'd reach for this tool, but there is no explicit statement of when to prefer it over event_get (plain read) or event_cursor, and no exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_cancelA
Creator-only OPEN -> CANCELLED; retract a handoff nobody should take (e.g. a pool target matching zero agents). Only OPEN handoffs cancel. In strict mode pass session creds.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Human-readable reason recorded with the handoff.cancelled event (optional). | |
| agent_id | Yes | Your agent_id; must be the handoff's creator. REQUIRED. | |
| handoff_id | Yes | The handoff_id to act on. REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses the creator-only authorization requirement, the state precondition (OPEN only), and that strict-mode trust requires passing session creds. It omits failure behavior and whether cancellation is reversible, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely front-loaded and dense: state transition, purpose, precondition, and auth note in three short clauses. The telegraphic style is efficient, though the clipped phrasing slightly risks ambiguity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no explanation. For a mutating, auth-sensitive tool the description covers the state constraint, authorization model, and strict-mode requirement adequately, leaving only edge-case failure behavior unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents agent_id, handoff_id, session_id, and session_secret in detail. The description reinforces the creator constraint and strict-mode credential requirement, but adds little syntax or format detail beyond the schema, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with an explicit state transition ('OPEN -> CANCELLED') and a concrete purpose ('retract a handoff nobody should take'). An agent can distinguish it from handoff_reject and handoff_complete from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use context with a concrete example ('a pool target matching zero agents') and a precondition ('Only OPEN handoffs cancel'). It does not explicitly name the alternative tools (reject/complete) or when NOT to cancel, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_claimA
Atomically claim an OPEN handoff; single winner, others get a structured error. Returns the payload + claimed_by/lease_expires_at. In strict mode pass session_id + session_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Your agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED. | |
| handoff_id | Yes | The handoff_id to act on. REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses atomicity, the single-winner outcome, the structured error for losers, the auth requirement in strict mode, and the returned fields (payload + claimed_by/lease_expires_at), which implies a lease/ownership model. It stops short of stating lease duration, renewal expectations, or whether a claim can be released, which would matter for a state-mutating coordination primitive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core behavior, then the return contract, then the conditional auth note. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail need not be expanded, and the description correctly covers concurrency semantics, auth modes, and failure shape. What's missing is post-claim lifecycle context (what the lease implies, and that handoff_complete/reject/cancel are the follow-up actions), which would help an agent plan the next step.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema. The description's mention of session_id + session_secret in strict mode restates what the schema descriptions already say, adding no new parameter semantics. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (claim), resource (handoff), and precondition (OPEN), plus the key semantic property (atomic, single winner). An agent can immediately distinguish it from handoff_create, handoff_list_available, handoff_get, and handoff_complete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditional guidance for trust_mode=strict (pass session_id + session_secret), which tells the agent when extra credentials are needed. It does not, however, contrast itself against adjacent siblings like handoff_get or handoff_verify, so the when-not guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_completeA
Owner-only delivery of a CLAIMED handoff: -> COMPLETED, or -> VERIFYING when acceptance_criteria were set (verifier decides via handoff_verify). In strict mode pass session creds.
| Name | Required | Description | Default |
|---|---|---|---|
| result | No | Completion result (optional; string or JSON) persisted on the handoff, recorded with handoff.completed, and delivered to the creator inbox. | |
| agent_id | Yes | Your agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED. | |
| handoff_id | Yes | The handoff_id to act on. REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the ownership restriction and the strict-mode session-credential requirement, which is valuable. However, it omits irreversibility, error behavior, and the side effect that the result is persisted and delivered to the creator inbox (that detail only lives in the schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the key state-transition and ownership facts with no wasted words. The terse arrow notation is efficient, though slightly cryptic for an agent parsing on first read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating handoff tool with an output schema covering returns, the description supplies ownership, transition logic, the verify path, and auth mode. That is nearly sufficient, with only irreversibility/error semantics left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents agent_id, handoff_id, session_id/secret, result, and project_root. The description adds the strict-mode credential hint but otherwise repeats what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('delivery of a CLAIMED handoff') and spells out the resulting state transitions (-> COMPLETED, or -> VERIFYING). It is clearly distinguishable from siblings like handoff_verify, handoff_reject, and handoff_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the branching condition for COMPLETED vs VERIFYING and names handoff_verify as the decider when acceptance_criteria were set, plus the owner-only constraint. It stops short of explicitly stating when to prefer handoff_cancel or handoff_reject instead, so no full when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_createA
Create an OPEN handoff (validates target/visibility); emit handoff.created. After creating, poll handoff_get for status/result. Full docs: okto-nexus://reference/tool-docs/handoff.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | Routing target as a raw JSON object - which agents may CLAIM this handoff (REQUIRED). strategy one of: direct {"strategy":"direct","agent_id":"<id>"}; capability {"strategy":"capability","capability":"<cap>"}; role {"strategy":"role","role":"<role>"}; tag {"strategy":"tag","selector":{"<key>":["<value>",...]}} (flat: AND across keys, OR within values) or rich [{"key":"<k>","operator":"In|NotIn|Exists|DoesNotExist","values":["<v>",...]}] (ANDed; CAUTION: NotIn/DoesNotExist also match agents MISSING the key); broadcast {"strategy":"broadcast"}; mixed {"strategy":"mixed","rules":[<sub-target>,...]}; direct_with_fallback {"strategy":"direct_with_fallback","agent_id":"<id>","fallback_after_seconds":<n>}. Competing-consumers: the first to handoff_claim wins. Hierarchy/catalog rules, examples, edge-cases: okto-nexus://reference/target-grammar. | |
| payload | No | Inline work content (optional; raw JSON object/array, string, or null - NOT JSON-encoded). Returned only to the claimant by handoff_claim / claimant handoff_get. For large content pass an artifact_id. | |
| trace_id | No | Trajectory trace_id to stamp on this handoff (optional; non-empty string, max 128 chars). Needs the feature_trace flag ON, else accepted and ignored; omitted = generate one. | |
| verify_by | No | Who verifies (optional; only WITH acceptance_criteria; raw JSON object). one of: {"kind":"creator"} (default); {"kind":"agent","agent_id":"<id>"} (must be registered); {"kind":"capability","capability":"<name>"} (in the catalog; resolved at verify time). The claimant never verifies their own delivery. | |
| depends_on | No | Handoff ids this one depends on (optional; raw JSON list of 1..20 unique existing ids; IMMUTABLE). Blocked - unlisted/unclaimable - until ALL are COMPLETED. Needs feature_dag ON (rejected while OFF). | |
| session_id | No | Session_id attributing this operation to a specific open session of yours (optional). | |
| visibility | Yes | Who may SEE the handoff (separate from who may CLAIM it = target). one of: public, eligible, private. REQUIRED (case-insensitive). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| from_agent_id | Yes | Your agent_id (the creator); recorded as the handoff's originator - the owner is whichever agent later claims it (handoff_claim), not necessarily you. | |
| acceptance_criteria | No | Verification contract (optional; raw JSON list of 1..20 unique non-empty strings, max 500 chars each; IMMUTABLE). handoff_complete then parks the handoff in VERIFYING until the handoff_verify verdict. Needs feature_verification ON (rejected while OFF, never ignored). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that target/visibility are validated, that a 'handoff.created' event is emitted, and that status/result should be polled via handoff_get. However, it omits permissions, error handling, or side effects like ownership transfer, leaving significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tightly written sentences: front-loaded action, validation, event, next step, and a documentation link. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, no annotations, and an output schema that covers return values, the description is partially complete: it names the creation, validation, event emission, and follow-up poll. But it omits failure modes, required permissions, and other behavioral traits needed for a complex coordination tool, relying heavily on the external doc link.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters thoroughly. The description only reiterates that target/visibility are validated and adds no parameter syntax or meaning beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and resource ('handoff') with additional state ('OPEN') and validation scope ('validates target/visibility'). It clearly identifies the operation, though it does not explicitly differentiate from siblings like handoff_claim beyond the verb 'Create'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies when to use by stating the next step ('After creating, poll handoff_get for status/result'), but gives no explicit when-not conditions or alternative selection guidance relative to other handoff tools (claim, complete, etc.). Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_getA
Read a handoff by id: status, claimant, payload, result/rejected_reason + verification/dependency fields if set. The creator's path to the outcome. Full docs: okto-nexus://reference/tool-docs/handoff.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Your agent_id (REQUIRED). Creator and claimant always read; others are gated by the handoff visibility. | |
| handoff_id | Yes | The handoff_id to act on. REQUIRED. | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the burden. It signals this is a read and notes conditional fields ("if set"), but says nothing about the visibility gating described in the schema, permissions, or auth requirements. An output schema exists, so return-format explanation is not needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the verb and resource, with the field list and doc pointer kept short. The trailing documentation link is mildly expendable but harmless.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with a full output schema and complete parameter coverage, the description covers purpose and return fields adequately. The only real gap is the absence of any behavioral note on visibility/access, which the schema partially covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters are documented there, so the baseline is 3. The description adds no syntax or format detail beyond restating that the read is by id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Read a handoff by id") and enumerates the fields returned (status, claimant, payload, result/rejected_reason), which distinguishes it from list-oriented siblings like handoff_list_available. It doesn't explicitly name the sibling it contrasts with, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"The creator's path to the outcome" hints at context of use, but there is no explicit when-to-use guidance, no when-not, and no named alternatives (e.g. handoff_list_available for discovery). Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_list_availableA
Expire leases, then list OPEN handoffs visible+eligible to the caller (paginated); entries are metadata-only until claimed.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max handoffs per page (optional; default applied by the server, clamped to the maximum). | |
| cursor | No | Opaque pagination cursor: pass back the next_cursor from the previous page to continue (optional; omit or '0' to start from the beginning). | |
| agent_id | Yes | Your agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED. | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| timeout_seconds | No | Long-poll bound in SECONDS (optional). 0/omitted = non-blocking snapshot; >0 OPTS IN to blocking until a claimable handoff appears or the timeout elapses. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose the non-obvious side effect (expiring leases before listing) plus the 'metadata-only until claimed' constraint, which an agent would not infer from the name. It omits permission/auth requirements and any note on whether lease expiration is destructive or affects other callers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence, front-loaded with the side effect and the action, with no filler. The semicolon clause packs scope, pagination, and the metadata-only caveat efficiently, though it is a lot compressed into one sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return shape need not be explained. The description covers the key behavioral facts (lease expiration, eligibility scoping, pagination, metadata-only pre-claim state) needed to call it correctly; only auth/permission expectations are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so limit/cursor/agent_id/project_root/timeout_seconds are already fully documented; the description only reinforces 'paginated' and caller-scoped visibility. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (OPEN handoffs) plus a scope qualifier ('visible+eligible to the caller'), which separates it from handoff_get/handoff_claim. It also names a side effect ('Expire leases'), but does not explicitly contrast with sibling list/claim tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'visible+eligible to the caller' and the paginated metadata-only framing, but there is no explicit when-to-use versus handoff_get or handoff_claim, and no stated preconditions. The long-poll semantics are left to the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_rejectA
Reject a handoff (owner CLAIMED->REJECTED or direct-target OPEN->REJECTED). In trust_mode=strict pass session_id + session_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Human-readable reason (optional) persisted on the handoff, recorded with the rejection, and delivered to the creator inbox. | |
| agent_id | Yes | Your agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED. | |
| handoff_id | Yes | The handoff_id to act on. REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does disclose the auth requirement in trust_mode=strict. It does not cover reversibility, idempotency, error conditions, or what happens to the counterparty beyond what the schema's reason field already says.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action and state transitions, then the auth caveat. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with a full output schema, the description covers purpose, state transitions, and auth preconditions adequately. Remaining gaps (reversibility, side effects) are minor given the rich schema and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, including the trust_mode conditional on session_id/session_secret. The description merely restates that auth rule, adding no meaning beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reject) plus the resource (handoff) and even enumerates the exact state transitions it performs (CLAIMED->REJECTED, OPEN->REJECTED). This cleanly distinguishes it from siblings like handoff_cancel, handoff_complete, and handoff_verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The state transitions imply when the tool applies, and the second sentence gives an auth precondition for strict mode. However, there is no explicit guidance on when to reject versus cancel/complete a handoff, nor any stated exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_verifyB
Verifier-only verdict on a VERIFYING handoff: 'pass' -> COMPLETED (verified_by), 'fail' -> CLAIMED for rework (feedback + renewed lease). In strict mode pass session creds.
| Name | Required | Description | Default |
|---|---|---|---|
| verdict | Yes | The verdict. one of: pass (VERIFYING -> COMPLETED; handoff.completed carries verified_by), fail (VERIFYING -> CLAIMED for rework, lease renewed). REQUIRED (lowercase). | |
| agent_id | Yes | Your agent_id (the verifier); must satisfy the handoff's verify_by resolved AT VERIFY TIME. The claimant (claimed_by) is always refused. REQUIRED. | |
| feedback | No | Rework guidance (optional; max 2000 chars; only with verdict 'fail'). Persisted (each fail overwrites the previous) and delivered to the claimant's inbox. | |
| handoff_id | Yes | The handoff_id to act on. REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the two verdict outcomes, that feedback overwrites previous feedback and goes to the claimant's inbox, and the strict-mode credential requirement. It omits error/edge behavior (wrong state, idempotency, concurrency) that matters for a state-mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loads the core behavior, but the telegraphic style ('pass session creds', arrow notation) sacrifices readability, and the opening clause is dense enough that an agent must re-read it to parse the state transitions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers the primary path. For a 7-parameter, multi-mode (open vs strict trust) verification tool with no annotations, however, more operational context (preconditions, failure behavior) would be needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents every parameter including feedback and the session credentials. The description's only added param meaning is a restatement of the strict-mode session requirement ('pass session creds'), which the schema already states, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('verifier-only verdict on a VERIFYING handoff') and spells out both state-machine outcomes, which lets an agent distinguish it from handoff_complete, handoff_reject, and handoff_cancel. It is clear but relies on terse jargon rather than plainly naming the siblings it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Verifier-only' and the VERIFYING precondition imply when the tool applies, and the agent_id note ('the claimant is always refused') narrows the caller. But it never explicitly contrasts this with the sibling verify/complete/reject tools or states what to do when the handoff is in another state, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inbox_ackA
Acknowledge messages into history (read). Returns {acknowledged, read_message_ids}. Emits a message.read receipt to each sender.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Your agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| message_ids | Yes | The message_id(s) to acknowledge - a single id or a list. REQUIRED. Idempotent (already-read/foreign ids are no-ops). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose two real traits: it emits a message.read receipt to each sender (a visible side effect) and returns a specific result shape. It stops short of permissions, rate limits, or what happens to unacknowledged messages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, purpose and return shape and side effect front-loaded in that order. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, yet the description still names the response keys. Auth and idempotency are covered by a rich schema. The main gap is the absence of when-to-use routing among the inbox_* siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so agent_id, session_id, message_ids and session_secret are fully documented in the schema, including idempotency and trust-mode requirements. The description adds nothing about parameters, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: acknowledge messages into history, with a parenthetical clarifying it counts as a read. This is clear enough to distinguish it from inbox_pull/inbox_peek, but it never names those siblings or explains how ack differs from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to acknowledge versus pulling, peeking, or reading history. The idempotency note lives in the schema, not the description, and there is no mention of prerequisites or when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inbox_countA
Return your inbox lane sizes {unread, in_flight, read}. Cheap READ-ONLY between-turns check (pull when unread > 0). Expired in-flight leases count as unread; parked excluded.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Your agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it declares the operation read-only, characterizes cost ('cheap'), and discloses non-obvious counting semantics (expired in-flight leases count as unread; parked excluded). It doesn't cover auth requirements or rate limits, but for a read-only counter this is solid disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, zero filler, with the return shape front-loaded and the usage rule and edge-case semantics trailing. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be re-explained; the description supplies the selection context and the subtle counting rules an agent needs to interpret the numbers correctly. Complete for a single-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single agent_id parameter is fully documented in the schema (global inbox, reaches you in any workspace, REQUIRED). The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (inbox lane sizes) and enumerates the exact returned fields {unread, in_flight, read}. An agent can distinguish this counting tool from inbox_pull/inbox_peek/inbox_history without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage context ('Cheap READ-ONLY between-turns check') and a routing rule ('pull when unread > 0') that points to the sibling inbox_pull. No when-not-to-use guidance is given, but the positive selection rule is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inbox_extendA
Renew the lease on in-flight messages you pulled but have not finished (now + extend_seconds). All-or-nothing: if any id is not in-flight the call fails per-message and nothing is extended.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Your agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| message_ids | Yes | The message_id(s) whose lease to renew - id or list. REQUIRED. All must be in-flight; otherwise the call fails per-message and nothing is extended. | |
| extend_seconds | Yes | New lease duration in seconds, counted from now (REQUIRED; clamped to 10..3600). | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and discloses the key atomicity trait: all-or-nothing, with per-message failure and no extensions if any id is not in-flight. It also clarifies the lease is extended from now. It does not discuss authentication or return values, but the output schema handles returns and the input schema documents session/trust requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences with no filler. The core action is front-loaded, followed by the critical failure condition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 5-parameter mutation-like tool with 100% schema coverage, an output schema, and no annotations, the description covers the essential behavior: lease renewal and all-or-nothing failure. It does not mention trust-mode auth requirements, but those are fully documented in the schema, so the description is largely complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains all parameters, including extend_seconds clamping and session trust modes. The description adds only mild context via '(now + extend_seconds)' and the all-must-be-in-flight constraint, which is also in the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: renew the lease on in-flight messages. It also distinguishes this from ack/finish behavior by specifying messages 'you pulled but have not finished.' An agent can clearly identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: messages that are in-flight but not yet finished need their lease renewed. It does not explicitly name alternatives like inbox_ack or state when not to use this tool, so it falls short of full when/when-not guidance, but the implied usage is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inbox_historyA
List your acknowledged (read) messages, newest-first, keyset-paginated (stable pages even while you keep acknowledging). READ-ONLY.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max messages to return (optional; default 50, max 200; must be a positive integer). | |
| cursor | No | Opaque pagination cursor: pass back the previous page next_cursor (optional; omit = from the newest read messages). | |
| profile | No | Response size profile - one of: default, summary, full (optional; summary trims per-call tokens). | |
| agent_id | Yes | Your agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden: it declares READ-ONLY, states the sort order, and discloses a non-obvious behavioral guarantee (keyset pagination keeps pages stable even while new messages are acknowledged). It stops short of covering auth/permission requirements or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence; the operation and its scope come first, and the read-only declaration is a clean trailing flag. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and all four parameters are documented at 100%. The remaining gap is minor: no note on auth/permission expectations for reading the global inbox, but the essentials for correct invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents limit, cursor, profile and agent_id in detail. The description adds only the conceptual meaning of the cursor (keyset pagination), which is useful framing but not parameter-level detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (List), resource (acknowledged/read messages), ordering (newest-first) and pagination model. It implicitly separates itself from siblings like inbox_peek by scoping to already-acknowledged messages, but it never names an alternative tool explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (you call this to review what you've already acknowledged), and the parenthetical about stable pages while acknowledging hints at a concurrent-workflow use case. However there is no explicit when-to-use/when-not-to-use guidance or reference to a sibling alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inbox_peekA
Triage pending messages (unread + in-flight) WITHOUT consuming. READ-ONLY, envelope-only by default (body_preview + body_bytes). include_parked/include_bodies opt in.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max messages to return (optional; default 50, max 200; must be a positive integer). | |
| profile | No | Response size profile - one of: default, summary, full (optional; summary trims per-call tokens). | |
| agent_id | Yes | Your agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED. | |
| include_bodies | No | Return the FULL body instead of the envelope-only body_preview (first 200 chars) + body_bytes (optional; default false; inbox_pull is the way to consume a message). | |
| include_parked | No | Also show parked (dead-letter) messages that exhausted redelivery (optional; default false). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it declares READ-ONLY, states the default payload (body_preview + body_bytes), and flags which behaviors require explicit opt-in. It stops short of describing ordering, pagination, or what 'in-flight' concretely means for a caller.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense clauses with zero filler; the non-consuming read-only framing is front-loaded before the default and opt-in details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers the read-only posture, defaults, and opt-ins. It is nearly complete, missing only hints about result ordering/volume expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents limit, profile, agent_id, include_bodies and include_parked in detail. The description only restates the include_parked/include_bodies opt-in defaults, adding little beyond the structured fields — the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('triage/peek') and resource ('pending messages: unread + in-flight') and immediately scopes it as non-consuming, which cleanly separates it from the sibling inbox_pull. An agent can pick this over inbox_pull without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (triage without consuming, envelope-only by default) and notes the opt-in flags for deeper inspection. It never names inbox_pull as the consuming alternative in the description itself, so the when-not is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inbox_pullA
Take your unread messages into in-flight and return them WITH body (index-free; no cursor). At-least-once: unacked pulls are redelivered. Docs: okto-nexus://reference/tool-docs/inbox.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max messages to return (optional; default 50, max 200; must be a positive integer). | |
| profile | No | Response size profile - one of: default, summary, full (optional; summary trims per-call tokens). | |
| agent_id | Yes | Your agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED. | |
| session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict). | |
| lease_seconds | No | Lease (seconds) before pulled messages are redelivered (optional; default 300, clamped 10..3600); renew mid-turn with inbox_extend. | |
| session_secret | No | session_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key trait: at-least-once delivery with unacked pulls being redelivered, plus the fact that bodies are returned. It omits auth/trust-mode behavior and what happens on ack, but the redelivery contract is the critical disclosure and it is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses with no waste; the core semantics and the at-least-once caveat are front-loaded, and the docs pointer is a compact trailing reference instead of prose bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 6 params covered at 100%, an output schema present, and the delivery guarantee stated, an agent has what it needs to call this correctly. Only the trust-mode/auth nuance and how the returned messages differ from inbox_peek output are left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so limit, profile, lease_seconds, session_id, and session_secret are already fully documented in the schema. The description adds no parameter-level detail beyond implying the lease/renew flow, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('take your unread messages into in-flight and return them WITH body') and adds a differentiating qualifier ('index-free; no cursor') that separates it from cursor-based readers like event_cursor. It does not name the closest sibling (inbox_peek), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (consume unread mail as in-flight messages) and points to inbox_extend for mid-turn lease renewal, but never states when to prefer this over inbox_peek, inbox_history, or message_list. Usage is inferred rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
message_createA
Persist a message and fan it out to recipient inboxes; emit message.created. The response confirms delivery (recipients + delivered_count). Full docs: okto-nexus://reference/tool-docs/messages.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Message body (inline text). For large content, attach an artifact and keep body a short pointer. | |
| target | No | Routing target as a raw JSON object (optional; omit = broadcast to present agents). strategy one of: direct {"strategy":"direct","agent_id":"<id>"}; capability {"strategy":"capability","capability":"<cap>"}; role {"strategy":"role","role":"<role>"}; tag {"strategy":"tag","selector":{"<key>":["<value>",...]}} (flat: AND across keys, OR within values) or rich [{"key":"<k>","operator":"In|NotIn|Exists|DoesNotExist","values":["<v>",...]}] (ANDed; CAUTION: NotIn/DoesNotExist also match agents MISSING the key); broadcast {"strategy":"broadcast"}; mixed {"strategy":"mixed","rules":[<sub-target>,...]}. Hierarchy/catalog rules, examples, edge-cases: okto-nexus://reference/target-grammar. | |
| subject | Yes | Short message subject/title (one line). | |
| trace_id | No | Trajectory trace_id to stamp on this message (optional; non-empty string, max 128 chars). Needs the feature_trace flag ON, else accepted and ignored; omitted = inherit the reply parent's trace, or generate one. | |
| artifacts | No | List of artifact_id strings to attach (optional; at most 20, no exact duplicates; reference large content instead of inlining it in body). | |
| channel_id | No | Channel_id to post into (optional; omit = no channel). Organizational label only; does NOT decide recipients (the target does). | |
| project_root | Yes | Absolute path to the project (defines the workspace scope). | |
| from_agent_id | Yes | Your agent_id (the sender); recorded as the author - recipients reply by targeting it. | |
| session_secret | No | session_secret from session_open for from_session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode). | |
| from_session_id | No | Your session_id from session_open (optional in trust_mode=open; REQUIRED with session_secret in trust_mode=strict). | |
| parent_message_id | No | Message_id this is a reply to, to thread the conversation (optional). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well on side effects: it discloses persistence, fan-out to recipient inboxes, emission of a message.created event, and that the response confirms delivery with recipients + delivered_count. It omits auth/trust-mode requirements and any failure/partial-delivery semantics, so it falls short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses: the core action and side effect are front-loaded, the response contract follows, and the docs pointer is last. Every sentence earns its place with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be explained, and the schema richly covers routing, session, and artifact semantics. For an 11-parameter tool the description is somewhat thin on auth/trust-mode context, but the docs reference plus the schema make it usable end-to-end.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 11 parameters, including the complex target routing grammar. The description adds no parameter-level meaning beyond what the schema provides, which is the baseline 3 for full-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states specific verbs and a resource ('persist a message and fan it out to recipient inboxes; emit message.created'), so an agent immediately knows this is the write/send path. It does not explicitly name read siblings (message_get, message_list, message_wait) to differentiate, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-tool guidance. The description explains the mechanics of sending but never states the conditions under which an agent should call message_create rather than inbox_pull, message_status, or channel tools; it only points to external docs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
message_getA
MIGRATED (S3): replaced by inbox_pull / inbox_peek / inbox_history. Always returns ok:false code=MIGRATED with the replacement call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so: it discloses that the tool always fails, the exact error code, and that the response contains the replacement call. That is complete transparency for a tombstone tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence stating the migration, the replacements, and the return behavior. Every clause earns its place with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description already explains the only meaningful return behavior (ok:false code=MIGRATED with the replacement call). For a zero-parameter migrated stub, this is fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The schema has 100% coverage and the description correctly implies no input is needed, though there is no additional parameter meaning to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states precisely what the tool now is: a migrated stub that always returns ok:false with code=MIGRATED. It names the three replacement tools (inbox_pull / inbox_peek / inbox_history), so an agent can immediately distinguish it from those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent not to use this tool and to use the replacements instead, with the condition that any call will return a MIGRATED error. No inference is required about when or whether to call message_get.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
message_listA
MIGRATED (S3): replaced by inbox_peek / inbox_history (your messages) and event_get (bus traffic). Always returns ok:false code=MIGRATED.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and meets it: it discloses the exact return contract ('Always returns ok:false code=MIGRATED'), so an agent knows the call is a guaranteed no-op failure and can avoid it entirely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence beginning with 'MIGRATED' conveys status, replacements, and exact behavior with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a deprecated stub, the description covers everything an agent needs: deprecation status, replacement routing, and the deterministic failure response. An output schema exists, but the description still correctly summarizes the return contract rather than leaving the agent to discover it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty at 100% coverage, so there is nothing for the description to clarify. Baseline 4 applies; no parameter discussion is needed or given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what this tool now is: a migrated tombstone that always returns ok:false code=MIGRATED. It names the sibling replacements (inbox_peek, inbox_history, event_get) so an agent can immediately tell this tool apart from them without inspecting any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: it says the tool is replaced and maps each use case to its correct alternative — inbox_peek/inbox_history for 'your messages' and event_get for 'bus traffic'. There is no ambiguity about when not to call this tool and what to call instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
message_statusA
Track a message you SENT: per-recipient delivery states {recipient, status, attempts, read_at} (unread/delivered/read/parked). READ-ONLY.
| Name | Required | Description | Default |
|---|---|---|---|
| message_id | Yes | The message_id from message_create whose per-recipient delivery states you want to track. REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that the operation is READ-ONLY and enumerates the possible status values, which is genuinely useful. However it says nothing about permissions, whether status is eventually consistent, or how 'parked' differs operationally, leaving real behavioral gaps for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight clauses with the scope ('you SENT') and the read-only nature front-loaded; every element carries information. The inline set notation is dense but readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with an output schema, the description supplies scope, the return shape's key fields, and the read-only guarantee, which is enough to call it correctly. The output schema relieves it of explaining return values in detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter's origin (message_id from message_create) is already documented in the schema. The description adds no format or constraint details beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: tracking per-recipient delivery states for a message the caller SENT. The 'you SENT' scoping and the enumerated states (unread/delivered/read/parked) make it distinguishable from generic siblings like message_get or message_list, though no sibling is named directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Track a message you SENT' implies the usage context and requires a message_id from message_create, but there is no explicit when-to-use vs. when-not guidance and no alternative tool named for adjacent needs (e.g., reading message content vs. tracking its delivery).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
message_waitA
MIGRATED (S3): replaced by inbox_count polling (cheap) or event_wait (explicit blocking). Always returns ok:false code=MIGRATED.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well by disclosing the deterministic failure shape (ok:false, code=MIGRATED), so an agent knows calling it is pointless. It stops short of saying whether the tool is still callable at all or will be removed, but the essential behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, migration status and replacement named first, mechanical outcome second. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter deprecation shim with an output schema present, this is fully sufficient: it tells the agent what it is, what it returns, and where to go instead. Nothing needed to decide whether to invoke it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing to disambiguate. The description correctly implies no inputs are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool is (a migrated/deprecated stub) and what it does when called (always returns ok:false code=MIGRATED). It also names the two replacement tools, so an agent can distinguish this from event_wait and inbox_count without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to alternatives and gives the selecting condition for each: inbox_count for cheap polling, event_wait for explicit blocking. Effectively a when-not-to-use instruction, which is the strongest form of guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
nexus_infoA
Report server versions: package_version, schema_version, surface_revision, resource_versions, features (read-only {feature_*: bool}). Call when behaviour disagrees with cached schemas.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It implies a passive report and notes the features map is 'read-only {feature_*: bool}', but never states side-effect profile, permissions, or cost. With an output schema present the return values needn't be explained, so this is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the enumerated payload is front-loaded and the usage trigger follows. Dense and jargon-heavy, but every clause carries information and nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return shape is covered; the description still summarizes the payload and gives a trigger. For a zero-parameter read-only info tool this is close to complete, with only the safety/permission profile left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 (plus schema coverage is 100%). There is no parameter meaning the description could add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Report server versions') and enumerates the exact fields returned (package_version, schema_version, surface_revision, resource_versions, features). It is clearly distinguishable from siblings like capability_list or coordination_health, though it does not name those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete trigger: 'Call when behaviour disagrees with cached schemas.' That is a real when-to-use condition rather than implied usage. It lacks any when-not guidance or named alternatives, which keeps it out of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
poll_token_issueA
Issue an ephemeral read-only monitor bearer (nxsept_...) for this authenticated agent's session workspace. Store only in the background monitor; never persist the raw token.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | session_id returned by session_open. REQUIRED. | |
| session_secret | Yes | session_secret returned by session_open. REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does disclose real traits: the token is ephemeral, read-only in scope, prefixed nxsept_, bound to the authenticated session, and must not be persisted. It stops short of stating lifetime/expiry or that poll_token_renew exists for rotation, which are the remaining behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the security constraint second; nothing is wasted and no sentence restates the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers purpose, scope, and handling. It is only marginally incomplete in omitting token lifetime and the renewal path, which an agent would need to avoid misusing an expired token.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are described in the schema as values 'returned by session_open. REQUIRED.' The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Issue) plus resource (ephemeral read-only monitor bearer) and scope (this authenticated agent's session workspace). The nxsept_ prefix and 'ephemeral read-only' framing distinguish it from poll_token_revoke and poll_token_renew, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The sentence 'Store only in the background monitor; never persist the raw token' gives handling guidance but not selection guidance — it never says when to call this versus poll_token_renew or poll_token_revoke, nor what precondition (an open session) triggers it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
poll_token_renewA
Rotate and extend the active ephemeral poll token for this session. The previous raw nxsept_ bearer stops working immediately.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | session_id returned by session_open. REQUIRED. | |
| session_secret | Yes | session_secret returned by session_open. REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the single most important behavioral trait: "The previous raw nxsept_ bearer stops working immediately" — a specific, non-obvious invalidation side effect. It omits auth/permission requirements and failure modes (e.g. no active token), which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action and followed by the single highest-value behavioral fact. No filler, no redundancy, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 required params, full schema coverage, output schema present so return values need no explanation) and the description covers the key lifecycle consequence of renewing. It is close to complete, missing only explicit preconditions and when-to-use framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so the schema already documents session_id and session_secret (including their provenance from session_open). The description adds no format, constraint, or sensitivity detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb pair plus resource: "Rotate and extend the active ephemeral poll token for this session." That is well beyond a restatement of the name and clearly separates renew from issue/revoke. It stops short of naming those siblings explicitly, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase "the active ... token for this session" tells the agent a token must already exist, but there is no explicit when-to-call, no precondition statement, and no routing against poll_token_issue or poll_token_revoke. Minimal viable guidance, nothing more.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
poll_token_revokeB
Revoke the active ephemeral poll token for this session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | session_id returned by session_open. REQUIRED. | |
| session_secret | Yes | session_secret returned by session_open. REQUIRED. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state whether revocation is immediate or permanent, whether an already-revoked or missing token errors, that it requires the session_secret, or what effect revocation has on in-flight polling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every word contributes to identifying the action and its scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema removes the need to describe return values, and the tool is a simple two-parameter revoke. Still, with no annotations and no usage or behavioral context, the description is barely sufficient for an agent to know when revocation is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (session_id, session_secret) are documented in the schema as required values returned by session_open. The description adds no extra meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource (revoke the active ephemeral poll token) and scopes it to this session, so the action is unambiguous. It does not, however, distinguish itself from the sibling poll_token_issue and poll_token_renew tools, leaving the token lifecycle relationship implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no mention of the sibling tools poll_token_issue/poll_token_renew or when revoking should be preferred over letting a token expire. Usage can only be inferred from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_closeB
Close a session (idempotent); repeating returns ok and stays closed.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | The session_id returned by session_open. REQUIRED. | |
| workspace_id | No | Workspace_id scope guard (optional); when given it must match the session's workspace. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real behavioral value: idempotency and the exact repeat-call result ('returns ok and stays closed'). It omits the consequential side effects an agent needs before calling — whether closing invalidates tokens, terminates in-flight messages, or requires workspace ownership — so the disclosure is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the verb and resource lead and the idempotency caveat follows immediately. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the key idempotency trait. But with no annotations at all, it should also address the side effects of closing a session and any authorization or scope requirements, which it does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with session_id documented as 'returned by session_open' and workspace_id documented as a scope guard, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Close a session'), which is unambiguous on its own. However, it does nothing to distinguish the tool from close cousins in the sibling list such as session_open and session_heartbeat, so an agent must rely on the name alone to route between session lifecycle tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to close a session versus leaving it open, no prerequisites (e.g. must the session be open?), and no mention of the sibling tools that could be alternatives. The only contextual hint is that repeated calls are safe, which is a behavior note rather than usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_heartbeatA
Advance a session heartbeat and report the derived status; keeps you PRESENT (in the broadcast audience) and clear of the stale-session reaper.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | The session_id returned by session_open. REQUIRED. | |
| workspace_id | No | Workspace_id scope guard (optional); when given it must match the session's workspace. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose real behavioral traits: membership in the broadcast audience and eviction avoidance via the stale-session reaper. It omits failure behavior (e.g., what happens if the session is already reaped or expired), leaving a gap for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a semicolon-delimited elaboration; every clause earns its place and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't explain return values, and it correctly references the derived status. The only shortfall is the absence of any guidance on heartbeat cadence or on error/expired-session outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both session_id and workspace_id are already documented in the schema (including the workspace scope-guard rule). The description adds no parameter-level meaning beyond that, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Advance a session heartbeat') and its effect ('report the derived status'), which cleanly separates it from session_open/session_close in the sibling list. It doesn't explicitly name those siblings, so it stops just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'keeps you PRESENT ... and clear of the stale-session reaper' implies when to call it (to stay alive), but there is no explicit when/when-not guidance or cadence/frequency instruction. Usage must be inferred rather than read.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_openA
Open a session bound to (agent_id, workspace_id); returns a per-session session_secret (ONLY here - keep it; required by sensitive verbs in strict mode). Heartbeat to receive broadcasts.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Your agent_id; the session is bound to this identity. REQUIRED. | |
| metadata | No | Free-form JSON object stored with the session (optional). | |
| workspace_id | No | Workspace_id (from workspace_resolve) the session operates in (optional; omit for an unbound session). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose a genuinely important trait – that a per-session session_secret is emitted ONLY here, must be kept, and is required by sensitive verbs in strict mode – but omits session lifetime, idempotency on re-open, and any auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence covers the action, the key return value, and the follow-up step with no filler. Slightly cryptic in phrasing ('ONLY here - keep it') but appropriately front-loaded and compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return-value structure need not be explained, and the description adds the one thing an agent must know that the schema cannot express: the session_secret's singular issuance and downstream requirement. Only session lifetime and error behavior are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents agent_id, workspace_id, and metadata in detail. The description names the (agent_id, workspace_id) binding but adds no format or constraint detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Open) and resource (a session), and specifies the binding scope: (agent_id, workspace_id). An agent can distinguish it from session_heartbeat and session_close, though it does not explicitly contrast with them by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trailing clause 'Heartbeat to receive broadcasts' implies a lifecycle (open, then heartbeat) but gives no explicit when-to-use, prerequisites, or exclusions relative to siblings like agent_register or workspace_resolve. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tag_listA
List the global operator-managed tag catalog (keys + values) that agent tags, comm_scope and tag targets are validated against fail-closed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the catalog is global, operator-managed, and used in fail-closed validation, but does not state permissions, side effects, rate limits, or pagination behavior. 'List' strongly implies a read-only operation, which is useful but not sufficient for full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The trailing clause 'that agent tags, comm_scope and tag targets are validated against fail-closed' is dense but earns its place by explaining the catalog's role.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter list tool with an output schema, the description covers what the tool returns (keys + values) and why the catalog matters. It does not need to explain return values since an output schema exists; only explicit when-to-use guidance is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to clarify. Per the rubric, zero-parameter tools default to a baseline of 4 when the schema and description are otherwise coherent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (global operator-managed tag catalog keys + values), plus the validation domain (agent tags, comm_scope, tag targets). No sibling tool overlaps, so the purpose is unmistakable without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives implied usage by saying the catalog is what tags, comm_scope, and tag targets are validated against, suggesting an agent should consult it before those operations. However, it does not explicitly state when to use this tool, when not to use it, or what alternative exists (no similar sibling exists).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_listA
GLOBAL-ADMIN: enumerate ALL workspaces. Paths OMITTED by default (include_paths=true is an admin/ops opt-in). For discovery use agent_list / capability_list.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Your agent_id for permission evaluation in open stdio mode (optional; authenticated HTTP MCP uses the API-key identity). | |
| include_paths | No | Also return each workspace on-disk root_realpath (default false; paths OMITTED by default - opt-in defense-in-depth). For routine discovery use agent_list / capability_list. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that paths are omitted by default as a defense-in-depth measure, that include_paths is an admin/ops opt-in, and that this is a GLOBAL-ADMIN privileged operation. It stops short of stating rate limits or what the admin scope concretely gates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the scope/privilege constraint and the default behavior, with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers privilege scope, default behavior, and alternatives. Complete enough to call correctly, missing only finer detail on what the admin path reveals.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so baseline is 3; the description adds value by explaining WHY include_paths defaults to false (defense-in-depth, admin opt-in) rather than merely restating the default. It adds minor rationale beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('enumerate ALL workspaces') with an explicit scope qualifier (GLOBAL-ADMIN, ALL). An agent can distinguish this enumeration tool from sibling list/resolve tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes discovery use cases to agent_list / capability_list, giving a clear when-not condition. However it omits guidance on the closest sibling workspace_resolve, so the routing is not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_resolveB
Resolve a project_root to its deterministic workspace_id and upsert it.
| Name | Required | Description | Default |
|---|---|---|---|
| display_name | No | Human-friendly label to store/refresh for the workspace (optional). | |
| project_root | Yes | Absolute path to the project; the server derives workspace_id = sha256(realpath). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. 'Upsert it' usefully discloses that resolving also writes/refreshes a workspace record (a side effect not implied by 'resolve' alone), and 'deterministic' hints at idempotency. However it doesn't state what gets created or overwritten, whether display_name refresh is destructive, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the core action and its side effect are stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the schema covers both params. But for a tool that performs an upsert with no annotations, the description should say more about the write side effect (idempotency, what is created/updated) to be complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (project_root and display_name) are already fully documented, including the sha256(realpath) derivation. The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: resolve a project_root into a deterministic workspace_id, plus an upsert. This is clearly distinct from sibling workspace_list. It doesn't name an alternative explicitly, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as workspace_list. The agent must infer that this is the registration/derivation entry point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
43 tool updates
v0.1.10- First observed
agent_get - First observed
agent_list - First observed
agent_register - First observed
agent_whoami - First observed
artifact_get - First observed
artifact_put - First observed
capability_list - First observed
channel_create - First observed
channel_list - First observed
coordination_health - First observed
event_cursor - First observed
event_get - First observed
event_wait - First observed
handoff_cancel - First observed
handoff_claim - First observed
handoff_complete - First observed
handoff_create - First observed
handoff_get - First observed
handoff_list_available - First observed
handoff_reject - First observed
handoff_verify - First observed
inbox_ack - First observed
inbox_count - First observed
inbox_extend - First observed
inbox_history - First observed
inbox_peek - First observed
inbox_pull - First observed
message_create - First observed
message_get - First observed
message_list - First observed
message_status - First observed
message_wait - First observed
nexus_info - First observed
poll_token_issue - First observed
poll_token_renew - First observed
poll_token_revoke - First observed
session_close - First observed
session_heartbeat - First observed
session_open - First observed
shared_md_render - First observed
tag_list - First observed
workspace_list - First observed
workspace_resolve
TDQS
Scored across 43 tools
The tool set covers many distinct sub-domains (handoffs, inbox, events, sessions, etc.), and most tools have clearly differentiated purposes. However, some read-oriented tools overlap (event_get/event_wait, inbox_pull/inbox_peek) and three legacy message_* tools are deprecated but still present, which could momentarily confuse an agent.
All 43 tools use consistent snake_case with a domain_action pattern (e.g., handoff_create, inbox_pull, poll_token_issue). The only variation is minor compound forms like handoff_list_available, but the convention is predictable and uniform.
43 tools is well above the typical 3–15 range and even excluding the 3 deprecated message tools leaves 40. This heavy surface likely burdens agents and exceeds what the server’s purpose requires.
The surface covers CRUD-like lifecycle for handoffs, inbox, sessions, events, artifacts, agents, workspaces, and poll tokens, with no dead ends in core workflows. Minor gaps exist (no artifact list/delete, no agent/channel delete), but agents can work around them.
Maintenance
Related MCP Connectors
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
Agent-native notes, tasks, dev-docs, vaults, sync & handoffs. MCP + OpenAPI dual surface.
- OctopadOAuthapp.octopad
The back-office workspace for your team's AIs: tasks, knowledge and context shared over MCP.
Hosted MCP messaging across owners, tools, and machines, with readable transcripts.
Related MCP Servers
AlicenseNot gradedqualityDmaintenancePersistent memory and handoff intelligence layer for MCP agents. Most memory servers retrieve text — Memory Nexus compounds operational context, learning from usage and progressively synthesizing observations into higher-order intelligence across sessions and tools.MIT- AlicenseAqualityAmaintenanceNexus Memory gives every MCP-compatible agent one persistent, self-hosted shared memory with hybrid retrieval, drift detection, and anti-poisoning features.213MIT
- AlicenseNot gradedqualityAmaintenanceA real-time inter-agent switchboard, delivered as one centralized streamable-HTTP MCP server. Any MCP-capable agent can message, coordinate, and stay ambiently aware of others.1AGPL 3.0
- AlicenseNot gradedqualityDmaintenanceMCP server that wraps the nexus CLI, giving AI agents cross-session memory, semantic search, preference learning, and smart context injection.137 npmMIT