Skip to main content
Glama
OktoLabsAI

okto-nexus

by OktoLabsAI

Okto Nexus

Local-first coordination for teams of AI agents.

Okto Nexus is an MCP server and operator hub for agents working in the same repository. It gives them durable identities, presence, messages, inboxes, handoffs, artifacts, an event log, governance controls, and a live dashboard without requiring a cloud broker.

okto-nexus serve exposes the complete hub on one port:

  • /mcp — MCP over streamable HTTP, authenticated as an agent;

  • /api/v1 — operator REST APIs, read-only monitor endpoints, and SSE;

  • / — the bundled React dashboard.

The classic stdio transport remains supported. Coordination records live in one SQLite database in WAL mode. Optional metrics, referenced workspace files, and the derived shared.md view live outside that database.

Release fact

Value

Package

okto-nexus 0.1.10

Python

>=3.11

MCP surface

43 tools by default; 46 with memory enabled

MCP resources

12 versioned reference resources

MCP prompts

0

Surface revision

33

Database schema

28 migrations, 34 tables

Storage

local SQLite/WAL catalog + adapter-backed artifact payloads

Contents

Related MCP server: Nexus Memory

Why Nexus

  • One coordination space per project. An absolute project_root is canonicalized and hashed into a deterministic workspace_id. Every client that resolves the same real path joins the same workspace.

  • Durable delivery. Messages fan out into per-recipient inbox lanes with leases, redelivery, acknowledgements, delivery status, and optional read receipts.

  • Single-winner work dispatch. Handoffs support atomic claim, leases, rejection, cancellation, optional verification, and optional dependency graphs.

  • Explicit identity and presence. Operators create identities and API keys; agents open sessions and heartbeat to remain present.

  • Targeted routing. Direct, capability, role, tag, broadcast, mixed, and direct-with-fallback strategies share one validated grammar.

  • Governed communication. Permissions, communication scopes, versioned policies, quotas, guardrails, groups, and human approval can restrict writes without exposing the control plane to agents.

  • Observable by design. Monotonic event IDs, cursor reads, long-poll, replay export, SSE, health aggregates, and a live dashboard expose what the team is doing.

  • Local-first and fail-closed. Configuration, target grammars, catalogs, workspace paths, API keys, and state transitions are validated before writes.

  • Token-aware MCP docs. First-use guidance remains resident; deeper reference material is available through versioned MCP resources on demand.

Install

The recommended install includes the HTTP hub, dashboard, local embedding provider, and tokenizer:

uv tool install "okto-nexus[serve]"
okto-nexus serve

Equivalent with pipx:

pipx install "okto-nexus[serve]"

Available extras:

Extra

Includes

Use it when

none

stdio MCP core

You only need a lightweight local stdio server

serve-lite

FastAPI, Uvicorn, dashboard

You need HTTP without Torch/model dependencies

embeddings

sentence-transformers

You want the local embedding provider separately

serve

HTTP stack, embeddings, tokenizer

You want the complete supported hub

dev

pytest, FastAPI, Uvicorn, httpx

You are developing or testing Nexus

The published wheel and sdist contain the compiled dashboard. Node.js is only needed when rebuilding the frontend from source.

From a checkout:

git clone https://github.com/OktoLabsAI/okto-nexus.git
cd okto-nexus
uv sync --extra dev

# Complete HTTP build:
uv sync --extra serve --extra dev

Start the hub

okto-nexus serve

Defaults:

  • dashboard: http://127.0.0.1:8202/;

  • MCP: http://127.0.0.1:8202/mcp;

  • data directory: ~/.okto_nexus;

  • database: ~/.okto_nexus/nexus.db;

  • initial workspace context: the current directory. The dashboard keeps a saved selection when present and otherwise may open the all-workspaces view.

Useful variants:

okto-nexus serve --project-root /absolute/path/to/project
okto-nexus serve --host 0.0.0.0 --port 8202
okto-nexus serve --trust-mode strict
okto-nexus serve --embedding-mode local

On first use, open Agents → New agent in the dashboard. Create one identity per participant and copy its nxs_... key immediately: Nexus stores only the hash and shows plaintext only at creation or regeneration. Capability and tag values must first exist in Registry before an identity can use them.

Connect an MCP client

The dashboard generates snippets for Claude Code, Claude Desktop, Codex, Cursor, VS Code, Windsurf, and Cline. The generic streamable-HTTP URL is:

http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME

Credential extraction order is query api_key, x-api-key header, then Authorization: Bearer. Treat client configuration containing a query key as a secret.

Examples:

claude mcp add -t http okto-nexus \
  "http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME"

codex mcp add okto-nexus \
  --url "http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME"

Generic JSON:

{
  "mcpServers": {
    "okto-nexus": {
      "url": "http://127.0.0.1:8202/mcp?api_key=nxs_REPLACE_ME"
    }
  }
}

Stdio

Run okto-nexus without a subcommand for stdio:

{
  "mcpServers": {
    "okto-nexus": {
      "command": "okto-nexus",
      "args": [],
      "env": {
        "OKTO_NEXUS_HOME": "/absolute/path/to/nexus-home"
      }
    }
  }
}

Without OKTO_NEXUS_API_KEY, stdio preserves the cooperative anonymous model. Set that variable to an active nxs_... key to bind the process to the same authenticated identity rules as HTTP. An invalid configured key fails closed.

Agent pre-flight

Every authenticated agent should do this on its first turn:

  1. Call agent_whoami(), use the returned agent_id consistently, and treat its role and communication.content as the default operating contract.

  2. Call workspace_resolve(project_root=<absolute cwd>).

  3. Call session_open(agent_id=<you>, workspace_id=<resolved id>) and retain the returned session_id and one-time session_secret.

  4. Check inbox_count(agent_id=<you>); pull and acknowledge backlog.

  5. Anchor monitoring with event_cursor(project_root=..., agent_id=<you>, stream="workspace").

The full procedure is available at okto-nexus://reference/preflight.

The role guides responsibilities, operating perspective, and decision boundaries. The communication block guides tone, format, language, verbosity, structure, and agent-to-agent as well as user-facing communication. Follow both unless the user explicitly directs otherwise for the current task or interaction. That task-scoped override does not modify the Nexus profile or bypass permissions, policies, guardrails, approvals, communication scope, safety rules, or higher-priority host instructions. Capabilities are routing claims, not authorization or persona.

Direct message

{
  "project_root": "/absolute/path/to/project",
  "from_agent_id": "researcher",
  "subject": "API findings",
  "body": "The endpoint is idempotent; details are attached.",
  "target": {
    "strategy": "direct",
    "agent_id": "implementer"
  },
  "from_session_id": "ses_...",
  "session_secret": "..."
}

The response names the resolved recipients and delivery count. The recipient uses:

inbox_count(agent_id="implementer")
inbox_pull(agent_id="implementer", session_id="ses_...", session_secret="...")
inbox_ack(agent_id="implementer", message_ids=[...],
          session_id="ses_...", session_secret="...")

Messages are delivered through the inbox. event_get and event_wait are observability tools, not delivery.

Long-poll

event_wait is a snapshot when timeout_seconds is omitted, null, or 0. Long-poll is explicit:

event_wait(
  project_root="/absolute/path/to/project",
  agent_id="researcher",
  stream="workspace",
  cursor=123,
  timeout_seconds=25,
  profile="summary"
)

Always continue from next_cursor. The waiter uses SQLite PRAGMA data_version plus bounded sleep polling; HTTP runs the blocking wait in a worker thread so it does not block the shared event loop.

Architecture

MCP stdio             MCP HTTP              REST / SSE / SPA        CLI
    \                     |                         |                 /
     +---------------- inbound adapters and transport auth ----------------+
                                      |
                             application services
       identity · messages · inbox · handoffs · events · artifacts
       permissions · policies · approvals · guardrails · memory · health
                                      |
                         domain models and pure rules
                                      |
      +------------------------- outbound ports ---------------------------+
      | SQLite repositories | files/shared.md | waiter | telemetry | embed |
      +--------------------------------------------------------------------+

bootstrap() resolves configuration, creates the store, applies migrations, wires repositories, telemetry, embeddings, and approval execution, then seeds the reserved operator, backfills the capability catalog, and creates built-in permission/communication presets. create_server() lazily imports FastMCP and registers the effective tools and resources. serve wraps the same composition in FastAPI/Uvicorn; tail and admin are separate CLI adapters.

Important boundaries:

  • domain code contains state machines, routing, IDs, and invariants;

  • application services own the core coordination use cases; operator CRUD/maintenance routes may drive repositories and units of work directly;

  • inbound adapters translate MCP, HTTP, SSE, and CLI calls;

  • outbound adapters implement SQLite, files, telemetry, tokenization, embeddings, and waiting;

  • coordination truth is durable in SQLite; bounded process caches are implementation details, not authoritative state.

HTTP surfaces and authentication

Surface

Loopback bind

Non-loopback bind

SPA shell/assets, /healthz, info, license

Public

Public

REST data/control plane

Keyless operator trust

Active nxs_ key required

MCP /mcp

Active nxs_ key required

Active nxs_ key required

EPT monitor endpoints

Scoped nxsept_ accepted

Scoped nxsept_ accepted

MCP-over-HTTP connections always represent an agent; stdio may use the cooperative anonymous mode. The dashboard/REST loopback trust path represents the local operator. Browser-origin checks protect mutating operator routes, and binding beyond loopback removes keyless REST trust.

On a non-loopback bind, use the reserved operator identity's key for the dashboard/control plane. Participant keys authenticate requests but operator-only routes return PERMISSION_DENIED. When a store has no keys at all, startup creates the operator key and prints its plaintext once.

Permanent agent keys can authenticate REST and MCP, but helper monitors should receive only a short-lived ephemeral poll token (nxsept_...). EPTs are bound to the issuing session, agent, and workspace and are accepted only as Authorization: Bearer nxsept_...; query api_key and x-api-key are rejected for this token type. The bearer is valid only on:

  • GET /api/v1/events and GET /api/v1/events/cursor;

  • GET /api/v1/inbox/count and GET /api/v1/inbox/peek.

They cannot call MCP or mutate state.

Dashboard

The bundled dashboard provides:

  • Graph — toggle between detailed agent cards and compact activity-sized circles, with profile colours, live presence status badges, recent message flow, unread traffic, open handoffs, and claimed relationships;

  • Messages — inbox lanes, peer conversations, undelivered targeting outcomes, receipts, and optional semantic search;

  • Handoffs — a six-column Kanban including VERIFYING, claim details, dependency state, verification, cancellation, and results;

  • Meta-harness — an OpenAI-style chat for private or broadcast messages and handoffs, with agent filtering, aggregated acknowledgement flags, and replies/results in one timeline;

  • Artifacts — paginated deliverables with list/Explorer-style grid views, type-aware icons, rich previews, metadata, and managed-payload downloads;

  • Events — filtered event history, trace navigation, and live SSE updates;

  • Memory — durable memory browse/search/curation when feature_memory is enabled;

  • Workspaces — sessions, analytics, and coordination health;

  • Agents — identities, keys, activation, roles, capabilities, metadata, colors, permissions, tags, inbound/outbound audiences, communication style, and steering;

  • Registry — operator-managed capability and tag vocabularies;

  • Policies — versioned policies and per-agent bindings;

  • Guardrails — groups, versioned content rules, assignments, and scrubbed denial audit;

  • Communication — versioned communication presets and bindings;

  • Approvals — pending and decided human-in-the-loop actions;

  • Settings — runtime-manageable settings, feature flags, retention, and database maintenance; metrics use their own header-menu panel.

Semantic search requires embedding_mode=stub or local. off returns EMBEDDINGS_UNAVAILABLE on the REST search endpoint. The local provider needs the embeddings extra; stub is deterministic and is intended for tests or demonstrations.

Coordination model

Workspaces, agents, and sessions

  • Agents are global identities; workspaces represent canonical project roots.

  • Most coordination tools accept project_root.

  • session_open and shared_md_render consume a resolved workspace_id.

  • An authenticated agent can update only its own profile with agent_register, subject to identity.update_profile and identity.update_capabilities. Operators create identities.

  • Authenticated discovery is reachability-scoped. agent_list and agent_get hide unreachable peers; capability_list returns the complete catalog but filters owner identities.

  • workspace_list is permission-gated; absolute paths require a separate permission.

  • Cross-workspace errors are intentionally operation-specific: WORKSPACE_MISMATCH for ownership guards, NOT_FOUND for hidden artifact or memory reads, and DEPENDENCY_NOT_FOUND for dependency creation.

Presence is explicit. A session is considered present while its heartbeat is within presence_ttl_seconds. A trust-sensitive write advances the heartbeat only when it authenticates with that session's credentials; in trust_mode=open, a credential-free write advances no session. Read-only tools do not heartbeat. Call session_heartbeat during long read-only or idle periods and session_close when finished.

Messages and inboxes

message_create persists one message and resolves recipients at send time. Each recipient gets a durable delivery row in its global inbox. The lanes are:

Lane

Meaning

unread

Available to pull

delivered

Pulled and protected by an in-flight lease

read

Acknowledged

parked

Dead-lettered after exhausting delivery claims; not redelivered automatically

An expired delivered lease becomes pullable again. Delivery is therefore at-least-once until acknowledgement. message_status lets the sender inspect each recipient's lane. Every pull/redelivery emits message.delivered; ack emits message.read. By default, ack also sends one synthetic read-receipt message to the sender; receipts do not recursively create receipts.

Channels are organizational labels, not ACLs or delivery mechanisms. Access still intersects with permissions, policy, guardrails, and communication reachability. Message retention can remove aged messages and their deliveries, including unread or in-flight rows, so durability is bounded by configured retention and explicit database reset.

Communication intent

Choose the coordination mechanism by intended outcome: if another agent is expected to perform work or produce a deliverable, create a handoff. If the recipient only needs to know something or reply, use a message.

  • Handoff — executable work. Any request to execute, investigate, change, build, test, review, validate, or otherwise produce a deliverable must use handoff_create; a direct message must not be its sole record. Target the intended assignee directly when known, or use a capability, role, tag, mixed, broadcast, or direct-with-fallback target when the first eligible claimant should own the work. The handoff is the canonical, operator-visible record of ownership, lifecycle, result, and verification. Include the objective, context, scope, constraints, and expected deliverable; add acceptance criteria, dependencies, and a verifier when applicable and the corresponding feature_verification / feature_dag flag is enabled. Those fields are rejected while their feature is off.

  • Broadcast message — shared alignment. Use it for shared context, decisions, announcements, discoveries, risk or blocker alerts, and general alignment across every selected reachable recipient. It informs recipients but assigns no owner and creates no task lifecycle. If anyone is expected to act, create one or more handoffs.

  • Direct message — conversation and informal coordination. Use it for status checks, questions, clarifications, acknowledgements, focused context exchange, and informal coordination. It may discuss an existing handoff, but new executable work requires a handoff; reference that handoff in the conversation. Use handoff_get for canonical lifecycle state and direct messages for contextual updates or blocker explanations.

A broadcast message is informational fan-out to many recipients. A handoff with a broadcast target is a claim pool for one executor. If several agents must produce independent results, create separate handoffs.

Routing

There are seven strategies:

Strategy

Descriptor

Notes

direct

{"strategy":"direct","agent_id":"a"}

One named identity

capability

{"strategy":"capability","capability":"review"}

One or any of several registered capabilities

role

{"strategy":"role","role":"reviewer"}

Exact role match

tag

{"strategy":"tag","selector":{"team":["platform"]}}

Registered tag selector

broadcast

{"strategy":"broadcast"}

Present workspace agents for messages; globally registered eligible agents for handoffs

mixed

{"strategy":"mixed","rules":[...]}

Non-empty union of non-broadcast rules

direct_with_fallback

direct plus fallback_after_seconds and optional fallback

Handoffs only

Messages support direct, capability, role, tag, broadcast, and mixed. Omitting a message target means broadcast. Handoffs require an explicit target and additionally support direct-with-fallback.

Tag selectors use AND across keys and OR across values. Rich In/NotIn/Exists/DoesNotExist expressions are also supported. Capability and tag names fail closed against operator-managed catalogs.

Target resolution is intersected with presence where applicable and with the caller's effective communication reach/audience. Separate enforcement layers then allow, deny, limit, or intercept the write:

  • per-agent permissions and recipient/rate limits;

  • versioned policy action rules and quotas;

  • content guardrails;

  • optional HITL approval interception.

A channel itself adds no ACL. Visibility controls who may see an item; eligibility controls who may claim it.

Event log

Events are immutable after insertion during normal operation. event_id is globally monotonic and never reused. Retention may delete old rows, so retained history can contain gaps.

Streams are workspace, agent, and handoff. Supported filters are type, agent_id, task_id, handoff_id, and trace_id. Authenticated event reads require events.read and omit actors outside the caller's communication reach, except the caller's own and system events.

Response profiles:

Profile

Behavior

default

Safe trim: preserves contextual fields while removing empty or duplicated data; oversized event payloads can yield a follow-up hint

summary

Aggressive projection: omits heavy bodies/payloads and returns follow-up hints

full

Raw debugging escape hatch

event_get is non-blocking. event_cursor returns the current end in O(1). event_wait long-polls only when timeout_seconds > 0. Dashboard SSE already provides operator UI updates; agent monitoring remains MCP polling/long-poll or the EPT read-only REST plane.

Handoffs

OPEN      -> CLAIMED
OPEN      -> REJECTED | CANCELLED
CLAIMED   -> COMPLETED
CLAIMED   -> VERIFYING        when acceptance criteria exist
CLAIMED   -> REJECTED         claimant rejects
CLAIMED   -> OPEN             lease expires
VERIFYING -> COMPLETED        verifier passes
VERIFYING -> CLAIMED          verifier fails; executor reworks

Only one agent wins handoff_claim. A claim returns the confidential payload. Available-list responses omit payload; handoff_get includes it only for the claimant. Events and synthetic notifications never carry the payload.

With feature_verification=true, creation may include acceptance_criteria and verify_by. The executor cannot verify its own result. A failed verdict persists feedback, returns the handoff to CLAIMED, and renews the lease.

With feature_dag=true, creation may include depends_on. The persisted status remains OPEN, but blockedness is derived. Blocked items are excluded from handoff_list_available and claim returns DEPENDENCY_NOT_MET until every dependency is COMPLETED.

Artifacts and shared.md

Artifacts are published as text/json/markdown/html content or imported from a workspace-contained file. Path containment is checked after canonical resolution; escapes return PATH_OUTSIDE_WORKSPACE. Text content is bounded by max_inline_bytes.

Payload bytes and free-form metadata do not live in SQLite. The database keeps only the minimal searchable and authorization catalog. The default local ArtifactStore adapter writes each payload and its manifest.json below ~/.okto_nexus/artifacts/<workspace>/<agent>/<artifact-id>/. Imported files are copied into that managed location, so the artifact survives later changes to the original workspace file. The dashboard's Artifacts screen pages through the managed payload catalog, filters by one or more producing agents and a production date interval, and previews or downloads the selected payload.

The publisher's effective outbound audience is frozen on the artifact. The same audience controls artifact.created visibility and artifact_get. Unauthorized, cross-workspace, and missing reads all return NOT_FOUND.

shared_md_render atomically rewrites a derived workspace shared.md with four fixed sections: relevant agents/sessions, open tasks, open handoffs, and recent events. It never becomes the source of truth and never renders handoff payloads. Authenticated callers need shared_md.render.

Permissions, policies, guardrails, and approvals

The operator control plane manages:

  • permission presets and per-agent effective permissions;

  • capability/tag catalogs and communication scopes;

  • versioned attachable policies with deny-overrides and quota windows;

  • versioned communication guidance returned privately by agent_whoami;

  • agent groups and versioned guardrails assigned by scope and priority;

  • one-shot HITL approval decisions and operator steering.

Guardrail denials persist scrubbed metadata, not rejected raw content. When feature_hitl=true and a policy requires approval, message_create or handoff_create may return status: "pending_approval". Do not resend the action; watch the returned approval ID.

Opt-in features

All seven flags default to off:

Flag

Effect

feature_trace

Accept and project trace_id correlation

feature_hitl

Enable new require_approval interception

feature_verification

Enable handoff acceptance criteria and verdicts

feature_dag

Enable handoff dependencies

feature_memory

Register three memory MCP tools and show Memory UI; restart required

feature_health

Enable coordination_health MCP execution

feature_replay

Enable REST replay export

Nuances:

  • approval history and decisions remain available when interception is off;

  • disabling verification or DAG blocks new contracts/dependencies, while already persisted workflow state remains enforceable and decidable;

  • workspace health REST remains available when the MCP health flag is off;

  • CLI replay export is operator-shell access and is not gated by feature_replay;

  • memory REST supports operator curation independently, but MCP memory tools are registered only when the flag is on at startup.

MCP surface

The default server exposes 43 tools: 42 across tool modules plus nexus_info. Enabling feature_memory at startup adds three tools, for 46. Both transports expose the same effective tool/resource surface for the same configuration.

Area

Tools

Metadata

nexus_info

Identity/workspace

workspace_resolve, agent_register, agent_whoami, session_open, session_heartbeat, session_close, workspace_list, agent_list, agent_get, capability_list

Events

event_get, event_cursor, event_wait

Messages/channels

message_create, channel_create, channel_list, message_get, message_list, message_wait

Inbox

inbox_pull, inbox_ack, inbox_extend, inbox_peek, inbox_count, inbox_history, message_status

Handoffs

handoff_create, handoff_list_available, handoff_claim, handoff_complete, handoff_verify, handoff_reject, handoff_cancel, handoff_get

Artifacts

artifact_put, artifact_get

Derived view

shared_md_render

Health

coordination_health

Catalog

tag_list

Ephemeral monitor tokens

poll_token_issue, poll_token_renew, poll_token_revoke

Optional memory

memory_put, memory_get, memory_search

message_get, message_list, and message_wait are intentional migration shims. They return MIGRATED with replacements in the inbox/event surface instead of failing as unknown tools.

coordination_health stays registered but returns VALIDATION_ERROR while feature_health is off. memory_put/get/search are absent until feature_memory is enabled and the server restarts.

Session credentials are trust-sensitive on:

  • message_create;

  • handoff_claim/complete/verify/reject/cancel;

  • inbox_pull/ack/extend;

  • memory_put when published.

In trust_mode=open they are optional but validated if supplied. In trust_mode=strict they are required. poll_token_issue/renew/revoke always require a valid session_id and session_secret in both modes.

Versioned MCP resources

URI

Version

okto-nexus://reference/preflight

4

okto-nexus://reference/communication

3

okto-nexus://reference/monitoring

5

okto-nexus://reference/target-grammar

5

okto-nexus://reference/tool-docs/messages

3

okto-nexus://reference/tool-docs/inbox

2

okto-nexus://reference/tool-docs/events

2

okto-nexus://reference/tool-docs/handoff

4

okto-nexus://reference/tool-docs/identity

5

okto-nexus://reference/tool-docs/artifacts

4

okto-nexus://reference/governance

2

okto-nexus://reference/hitl

2

nexus_info reports package/schema/surface versions, the URI-to-version map, and effective feature flags. Use that live metadata instead of assuming a cached surface.

Response envelope

Successful tools return:

{
  "ok": true,
  "data": {}
}

Failures return:

{
  "ok": false,
  "error": {
    "code": "VALIDATION_ERROR",
    "message": "Human-readable explanation"
  }
}

Transient SQLite lock/busy failures use DB_ERROR with details.retryable=true.

Resident token footprint

For the 0.1.10 default surface:

Component

Characters

Server instructions

3,992

Tool docstrings

6,563

Parameter schemas/descriptions

15,720

Cuttable resident surface

26,275 (~6,568 tokens)

Total measured surface

37,345 (~9,336 tokens)

Deep explanations live in resources so clients load them only when needed. The current measured cuttable reduction against the frozen baseline is about 44.3%.

Configuration

For serve settings managed by the runtime catalog, effective precedence is:

CLI flag > environment variable > stored dashboard override > default

Within the core/stdio bootstrap there is no stored layer, so precedence is CLI > env > default. Unknown flags, missing values, invalid enums, and out-of-range numbers fail closed with CONFIG_ERROR. Boolean CLI flags take an explicit value such as --feature-trace true.

Core runtime

Environment

CLI

Default

Notes

OKTO_NEXUS_HOME

--home

~/.okto_nexus

Runtime data directory

OKTO_NEXUS_DB_PATH

--db-path

{home}/nexus.db

SQLite database

OKTO_NEXUS_BUSY_TIMEOUT_MS

--busy-timeout-ms

5000

Minimum 0

OKTO_NEXUS_POLL_INTERVAL_MS

--poll-interval-ms

200

Waiter interval; minimum 1

OKTO_NEXUS_MAX_WAIT_TIMEOUT_SECONDS

--max-wait-timeout-seconds

30

Server wait ceiling; minimum 0

OKTO_NEXUS_HANDOFF_LEASE_TTL_SECONDS

--handoff-lease-ttl-seconds

300

Minimum 1

OKTO_NEXUS_MAX_INLINE_BYTES

--max-inline-bytes

65536

Inline artifact/content limit

OKTO_NEXUS_INBOX_LEASE_TTL_SECONDS

--inbox-lease-ttl-seconds

300

Minimum 1

OKTO_NEXUS_SESSION_STALE_TTL_SECONDS

--session-stale-ttl-seconds

60

Derived stale threshold

OKTO_NEXUS_PRESENCE_TTL_SECONDS

--presence-ttl-seconds

1800

Broadcast/tag presence window

OKTO_NEXUS_SESSION_REAP_SECONDS

--session-reap-seconds

86400

Opportunistic stale close

OKTO_NEXUS_MAX_SHARED_MD_EVENTS

--max-shared-md-events

1000

Render ceiling

OKTO_NEXUS_MAX_EVENT_LIMIT

--max-event-limit

1000

Event page ceiling

OKTO_NEXUS_POLL_TOKEN_TTL_SECONDS

--poll-token-ttl-seconds

3600

Minimum 60

OKTO_NEXUS_TRUST_MODE

--trust-mode

open

open or strict

OKTO_NEXUS_EMBEDDING_MODE

--embedding-mode

off

off, stub, or local

OKTO_NEXUS_INBOX_READ_RECEIPTS

--inbox-read-receipts

true

Sender inbox receipts

OKTO_NEXUS_META_HARNESS_RECEIPT_DISPLAY

--meta-harness-receipt-display

inline

inline flags or legacy timeline messages

OKTO_NEXUS_EXPOSE_WORKSPACE_PATH

--expose-workspace-path

false

Operator REST/dashboard path disclosure

OKTO_NEXUS_AUTO_PRUNE_ON_START

--auto-prune-on-start

false

One bounded best-effort startup pass

Retention

Environment

CLI

Default

Minimum

OKTO_NEXUS_RETENTION_EVENTS_KEEP_DAYS

--retention-events-keep-days

30

0

OKTO_NEXUS_RETENTION_READ_DELIVERIES_KEEP_DAYS

--retention-read-deliveries-keep-days

14

0

OKTO_NEXUS_RETENTION_CLOSED_SESSIONS_KEEP_DAYS

--retention-closed-sessions-keep-days

7

0

OKTO_NEXUS_RETENTION_MESSAGES_KEEP_DAYS

--retention-messages-keep-days

30

7

Pruning removes aged events, read deliveries, closed sessions, and messages older than the message window. Message retention is pure-age: it can remove unread, in-flight, or parked messages and cascades to their deliveries and embeddings. Handoffs and other non-message live rows are not age-pruned.

Metrics

Environment

CLI

Default

Notes

OKTO_NEXUS_METRICS_MODE

--metrics-mode

disabled

disabled, local_only, anonymous_beacon

OKTO_NEXUS_METRICS_DIR

--metrics-dir

{home}/metrics

Local telemetry JSONL/state

OKTO_NEXUS_METRICS_BEACON_URL

--metrics-beacon-url

https://nexus-metrics.oktolabs.ai

Used only in beacon mode

OKTO_NEXUS_METRICS_RETENTION_DAYS

--metrics-retention-days

30

Minimum 0

OKTO_NEXUS_METRICS_PUBLISH_INTERVAL_SECONDS

--metrics-publish-interval-seconds

3600

Minimum 60

Metrics are opt-in. Local mode stores bounded per-event metadata; beacon mode publishes aggregate hourly counts only. Message bodies, prompts, workspace file paths, coordination IDs, keys, tokens, URLs, and stack traces are excluded. Local telemetry JSONL is not currently pruned automatically; operators must manage those files even though metrics_retention_days is validated and exposed in configuration.

Feature flags

Environment

CLI

Default

OKTO_NEXUS_FEATURE_TRACE

--feature-trace

false

OKTO_NEXUS_FEATURE_HITL

--feature-hitl

false

OKTO_NEXUS_FEATURE_VERIFICATION

--feature-verification

false

OKTO_NEXUS_FEATURE_DAG

--feature-dag

false

OKTO_NEXUS_FEATURE_MEMORY

--feature-memory

false

OKTO_NEXUS_FEATURE_HEALTH

--feature-health

false

OKTO_NEXUS_FEATURE_REPLAY

--feature-replay

false

feature_memory changes tool registration and requires restart/reconnect. The other flags gate live behavior.

Serve-only and transport-specific settings

Environment

CLI

Default

Scope

OKTO_NEXUS_PORT

--port

8202

serve

OKTO_NEXUS_HOST

--host

127.0.0.1

serve

OKTO_NEXUS_LOG_LEVEL

--log-level

warning

critical through trace

—

--project-root

.

Initial dashboard workspace

OKTO_NEXUS_NO_BANNER

—

unset

Suppress serve banner

OKTO_NEXUS_API_KEY

—

unset

Optional stdio authenticated identity

Operations

CLI commands

Command

Purpose

okto-nexus serve

Start MCP HTTP, REST, SSE, and dashboard

okto-nexus

Start MCP over stdio

okto-nexus tail

Operator NDJSON follower over the event service

okto-nexus admin prune

Enforce retention, optionally vacuum

okto-nexus admin issue-keys

Add keys to legacy keyless identities

okto-nexus admin export

Export a workspace replay stream as NDJSON

Use --help on every command for the full argument grammar.

Tail

okto-nexus tail \
  --project-root /absolute/path/to/project \
  --agent-id observer \
  --stream workspace \
  --from latest

tail applies per-agent event visibility. An optional --cursor-file belongs to that consumer only; corrupted checkpoints fail closed.

Retention

Start with a dry run:

okto-nexus admin prune \
  --project-root /absolute/path/to/project \
  --dry-run

Then execute:

okto-nexus admin prune \
  --project-root /absolute/path/to/project \
  --messages-keep-days 30 \
  --vacuum

Retention spans the whole shared store even though --project-root is validated as the command anchor. --vacuum is the only option that compacts freed pages on disk. There is no always-running coordination reaper; auto_prune_on_start is a bounded opportunistic pass.

Issue legacy keys

okto-nexus admin issue-keys \
  --project-root /absolute/path/to/project

This is additive and idempotent. Existing keys are never rotated. Newly issued plaintext keys are printed once.

Replay export

okto-nexus admin export \
  --project-root /absolute/path/to/project \
  --trace-id trc_... \
  --output nexus-events.ndjson

The first line is a manifest; subsequent lines are raw events ordered by event_id. CLI export is operator-shell access and remains available even when the REST replay flag is off.

Ephemeral monitor token

An authenticated agent can call poll_token_issue, give only the returned nxsept_... token and base URL to a read-only helper, renew it before expiry, and revoke it on teardown. The raw token is returned only on issue/renew.

Data model and migrations

The current schema contains 34 tables:

Area

Tables

Core coordination

schema_migrations, workspaces, agents, sessions, events, channels, messages, tasks, handoffs, artifacts, message_deliveries

Settings/security/catalogs

settings, permission_presets, tag_keys, tag_values, capability_names, ephemeral_poll_tokens

Search and memory

message_embeddings, memories, memory_embeddings

Governance/workflows

governance_policies, approvals, handoff_dependencies, policies, policy_versions, agent_policy_bindings, comm_presets, comm_preset_versions, agent_comm_binding, agent_groups, agent_group_members, guardrails, guardrail_versions, guardrail_assignments

Not every table is workspace-scoped: agents and catalogs are global, inbox deliveries are keyed by recipient identity, and bindings/control-plane records have their own ownership rules.

Migrations are embedded in the package and applied in order:

  • 001–008: core schema, close metadata, handoff payload/result, presence, durable inbox deliveries, leases, and session secrets;

  • 009–015: API keys, settings, permissions, embeddings, tags/scopes, capability catalog, and trace IDs;

  • 016–021: governance, approvals, verification, dependencies, memory, and health/event indexes;

  • 022–026: versioned attachable policies, communication presets, display colors, groups/guardrails, and ephemeral poll tokens.

Nexus refuses to run against an unsupported newer schema. Runtime SQLite databases and their WAL/SHM/journal sidecars are ignored and must not be committed.

Errors

The domain/MCP contract has a closed catalog of 29 canonical codes:

Area

Codes

Workspace

WORKSPACE_REQUIRED, WORKSPACE_UNRESOLVED, WORKSPACE_MISMATCH

Validation/identity

VALIDATION_ERROR, NOT_FOUND, NOT_OWNER, PERMISSION_DENIED

Catalog/control plane

TAG_IN_USE, CAPABILITY_IN_USE, POLICY_IN_USE, COMM_PRESET_IN_USE

Governance

POLICY_DENIED, QUOTA_EXCEEDED, GUARDRAIL_DENIED, CONFLICT

Dependencies/state

DEPENDENCY_NOT_FOUND, DEPENDENCY_NOT_MET, INVALID_TRANSITION, INVALID_STREAM

Handoffs

HANDOFF_ALREADY_CLAIMED, NOT_ELIGIBLE_TO_CLAIM

Content/path

CONTENT_TOO_LARGE, PATH_OUTSIDE_WORKSPACE

Infrastructure

CONFIG_ERROR, MIGRATION_ERROR, DB_ERROR, RENDER_ERROR

Compatibility/internal

MIGRATED, INTERNAL_ERROR

REST adapters also use transport-specific codes such as AUTH_FAILED, CROSS_ORIGIN_BLOCKED, INVALID_PARAM, INVALID_WINDOW, INVALID_SETTING, EMBEDDINGS_UNAVAILABLE, and INTERNAL.

Development

uv sync --extra dev
uv run pytest -q

For the complete HTTP/embedding environment:

uv sync --extra serve --extra dev
uv run pytest -q

Frontend:

cd frontend
npm ci
npm run build

The build writes packaged static assets under src/okto_nexus/adapters/inbound/http/static/.

Release checks:

uv lock --check
uv build --out-dir dist/release-0.1.10
uvx twine check \
  dist/release-0.1.10/okto_nexus-0.1.10-py3-none-any.whl \
  dist/release-0.1.10/okto_nexus-0.1.10.tar.gz

Publish only explicitly named current-version artifacts. The top-level dist/ may contain older builds.

Project layout

src/okto_nexus/
  adapters/
    inbound/
      cli/                  serve, tail, admin
      http/                 FastAPI, REST, SSE, packaged SPA
      mcp/                  server, resources, projections, 43/46 tools
    outbound/
      sqlite/               repositories and migrations adapter
      embedding/            optional semantic provider
      file/ sharedmd/       artifact and derived-view I/O
      telemetry/ tokenizer/ metrics support
  application/              use cases and ports
  domain/                   entities, routing, state machines, policies
  migrations/               001 through 027
  testing/                  reusable test/replay harnesses
frontend/                    React dashboard source
tests/                       unit, contract, integration, replay tests
docs/design/                 architecture and design records
pyproject.toml               package metadata and extras
uv.lock                      reproducible dependency lock

Contributing

See CONTRIBUTING.md for the supported Python and frontend development environments, validation commands, and pull-request workflow.

Troubleshooting

event_wait returns immediately

Pass timeout_seconds > 0. Omitted, null, and zero are snapshots by design.

An agent misses broadcasts

Check that it opened a session in the correct resolved workspace and continues to heartbeat. Read-only event/inbox checks do not advance presence.

A direct peer is missing from discovery

Authenticated discovery is filtered by communication reachability. Inspect the caller's outbound and the peer's inbound communication scopes/tags in the dashboard.

Capability or tag targeting fails

Create the capability/tag value in Registry first. Catalog validation is fail-closed.

Memory tools are absent

Set feature_memory=true and restart/reconnect. Unlike live behavior flags, this flag changes MCP tool registration.

Semantic search is unavailable

Use embedding_mode=stub or install the embeddings/serve extra and use embedding_mode=local. off intentionally disables search.

DB_ERROR reports a lock

If details.retryable=true, retry the same call after the competing writer commits. Avoid opening the runtime SQLite file with tools that hold long write transactions.

The hub says another server already owns the home

serve holds {home}/nexus.serve.lock and refuses a second server using that same home. Stop the other hub or choose a different --home; changing only --db-path does not change the lock scope.

Workspace path is rejected

Pass an existing absolute path. Nexus canonicalizes the real path before deriving the workspace ID and validating artifact containment.

A monitor cannot mutate state

That is expected for nxsept_ tokens. They are intentionally read-only and accepted only on the four monitor endpoints.

Cached docs appear stale

Call nexus_info and compare surface_revision and resource_versions before reusing cached MCP reference content.

Security and limitations

Report vulnerabilities privately according to SECURITY.md. Do not include vulnerability details, credentials, tokens, or private workspace data in a public issue.

  • Nexus is designed for local or controlled single-tenant coordination, not as a public multi-tenant broker.

  • It does not terminate TLS. Put an authenticated TLS reverse proxy in front of a remote bind.

  • The dashboard shell and public health/info/license assets remain public; data/control REST requires authentication outside loopback.

  • API keys are hash-only at rest and shown once. Regeneration invalidates the old key immediately.

  • Session secrets are stored in plaintext in the local SQLite database; anyone who can read that file is inside the session trust boundary.

  • Channels are labels, not security boundaries.

  • SQLite is the only built-in coordination store. There is no Redis, PostgreSQL, or cloud broker adapter.

  • There is no always-running coordination scheduler/reaper. Expiry is checked opportunistically and retention runs manually or at startup when enabled.

  • The HTTP server does use worker threads/tasks for operational needs such as blocking waits and telemetry; “no scheduler” does not mean “no threads.”

  • Message durability is bounded by message retention and explicit reset.

  • Memory is experimental and changes the MCP surface at startup.

  • Artifact file references are confined to the canonical workspace.

  • Avoid committing runtime databases, sidecars, metrics output, and local secrets to version control.

Future direction

Likely extension points are additional durable-store adapters, a push-backed waiter that removes internal sleep polling, stronger remote deployment packaging, and further generated documentation from the live MCP schemas. The delivered surface already includes permissions, catalogs, communication scopes, policies, guardrails, HITL, trace correlation, verified/DAG handoffs, memory, health, replay, ephemeral poll tokens, embeddings, metrics, REST, SSE, and the dashboard.

Release notes

0.1.10 — current

Validation-hardening release: lands the external PR backlog and closes the open validation bugs (#26–#30). The MCP contract remains at surface revision 33 and the latest database migration remains 028.

  • Bumped the package and distribution metadata from 0.1.9 to 0.1.10.

  • Capped message artifact references at 20 (MAX_ARTIFACTS) with {count, max} diagnostics on both MCP and REST; the dashboard Meta-harness route now defers to the domain cap (the legacy 10-attachment fence was removed) and VALIDATION_ERROR details survive the REST envelope.

  • Rejected exact-duplicate artifact references; the error names the offending index and the duplicated reference.

  • Capped tag selector value lists at 20 values per key with set-based de-duplication, removing an O(n²) scan reachable from message_create/handoff_create tag targets.

  • Added length caps to the routing target grammar identifier fields (agent_id/role/capability, 256 chars) and to each depends_on id (128 chars).

  • Resynced the dashboard ColorPicker draft when the value prop changes while the component stays mounted.

  • Consolidated the duplicated bounded-list count check into the shared domain helper check_list_size.

0.1.9

Dashboard presentation release. The MCP contract remains at surface revision 33 and the latest database migration remains 028.

  • Bumped the package and distribution metadata from 0.1.8 to 0.1.9.

  • Added list/grid switching to the Artifacts catalog, including an Explorer-style grid with type-aware file icons, names, and payload sizes.

  • Collapsed Meta-harness read-receipt notifications into acknowledgement flags on their original messages by default. Grey flags identify incomplete target sets; green flags mean every target acknowledged, and their detail modal identifies pending recipients and delivery/read timestamps.

  • Added the meta_harness_receipt_display interface setting so operators can restore separate receipt messages with the timeline mode.

0.1.8

Artifact storage and catalog release. The MCP contract is at surface revision 33 and the latest database migration is 028.

  • Moved artifact payloads and free-form metadata out of SQLite into an adapter-backed store, with a local filesystem adapter by default.

  • Imported path-based artifacts into managed storage so they remain available independently of the original workspace file.

  • Added the Artifacts catalog with paginated producer/date filtering, detail previews, expanded Raw/Rich rendering, and managed-payload downloads.

  • Added rendered metadata fields plus safe Rich previews for Markdown, JSON, and HTML artifacts.

0.1.7

Dashboard usability and operator-interaction release. The MCP surface remains unchanged; the latest database migration is 027.

  • Added the Meta-harness chat with independent Message/Handoff and Private/Broadcast controls.

  • Combined outgoing turns, incoming agent messages, and handoff completion or rejection outcomes in a live, agent-filterable timeline.

  • Rendered structured chat content as readable labels and lists instead of raw JSON.

  • Reworked Guardrails group composition and rule authoring, including agent auto-completion, capability-scoped assignments, regex assistance, and stricter validation.

  • Added completed-agent responses and rejection reasons to Handoff cards and details.

  • Removed nested page scrolling from the dashboard shell.

0.1.6

Startup compatibility maintenance release. The MCP contract and database schema are unchanged: surface revision remains 32 and the latest migration remains 026.

  • Rebuilt the MCP v1 FastMCP settings model after import, eliminating the unresolved lifespan forward-reference warning introduced by pydantic-settings 2.15.

  • Applied the compatibility path to both stdio and streamable-HTTP server construction.

  • Constrained the MCP SDK dependency to mcp>=1.0,<2 so adopting the breaking v2 API requires an explicit migration.

  • Added a regression test that promotes the startup warning to an error and verifies the settings model is complete.

  • Verified the installed executable with local MiniLM warm-up, a live health request, and the complete test suite.

0.1.5

Live MCP smoke-test maintenance release. The MCP contract and database schema are unchanged: surface revision remains 32 and the latest migration remains 026.

  • Restored the real stdio smoke test on clean stores after capability registration became fail-closed.

  • Documented when and how to run the isolated two-agent smoke flow on Unix and Windows, including its LIVE E2E RESULT: PASS completion signal.

  • Added semantic cards for structured kind messages in the Graph conversation drawer, with distinct read-receipt and handoff lifecycle treatments instead of raw JSON.

  • Synchronized pyproject.toml, uv.lock, and release commands at version 0.1.5.

0.1.4

Dashboard, observability, and coordination-guidance release. The MCP guidance contract advances to surface revision 32; the database schema is unchanged and the latest migration remains 026.

  • Added detailed and compact activity-based agent graph modes, profile colours, live status badges, richer relationship context, and graph conversations.

  • Added an event timeline, expanded event and handoff filters, message detail hydration, workspace display names, and catalog import/export workflows.

  • Improved dashboard views for agents, approvals, communication, events, guardrails, handoffs, messages, policies, and workspaces.

  • Extended observability APIs and repositories with message lookup, filtered handoff/event queries, and bucketed event timeline data.

  • Completed PyPI project URLs, keywords, and supported-Python classifiers, with a focused metadata regression test.

  • Added contributor and security policies, structured GitHub issue forms, and a dedicated documentation-assets location for product screenshots.

  • Clarified that agents follow the role and communication profile returned by agent_whoami unless the user supplies a task-scoped override, without bypassing platform governance.

  • Defined handoffs as the required, traceable mechanism for executable work, broadcasts as shared alignment or information fan-out, and direct messages as status, clarification, and informal coordination channels.

  • Rewrote this README against the current CLI, transport/authentication model, dashboard, 43/46-tool surface, seven routing strategies, message retention, governance features, verification/DAG workflows, 34-table schema, 29-code error catalog, operations, testing, and limitations.

  • Synchronized pyproject.toml and the root project entry in uv.lock at version 0.1.4, and aligned the package/dashboard license label with the addendum that is actually included in LICENSE.

  • Removed tracked runtime SQLite databases and added ignore coverage for database files and journal/WAL/SHM sidecars.

  • Verified release archives contain the dashboard and migrations without runtime SQLite databases or sidecars.

0.1.2

  • Documentation-accuracy sweep at surface revision 31.

  • Corrected pre-flight, monitoring, heartbeat, inbox, trace, policy, HITL, artifact-audience, and target-grammar documentation.

  • Kept 12 deep reference resources versioned and reduced the resident cuttable surface to about 26.3k characters.

  • Corrected the token-reduction gate to include experimental-surface growth.

0.1.1

  • Hardened authenticated self-only identity/session rules and permission checks.

  • Added permission-gated workspace paths, shared view, health, and memory.

  • Added configurable aggregate metrics telemetry.

0.1.0

  • Added ephemeral remote-monitor tokens and read-only monitor endpoints.

  • Added attachable policies, communication presets, guardrail/group administration, enriched agent graph cards, and loopback trust hardening.

  • Made memory a registration-time experimental MCP surface.

0.0.x

  • Built the MCP reference-resource system, HTTP hub, live dashboard, inbox receipts, semantic search, monitoring guidance, and dashboard observability waves.

License

Copyright 2026 Okto Labs.

Okto Nexus is distributed under the Elastic License 2.0 together with the project's SaaS, competing-service, internal-use, and branding addendum. It permits internal and qualifying single-tenant use and prohibits the specified multi-tenant, white-label/OEM, competing, and large-scale internal-platform uses; applicable notices and attribution remain required.

Read the complete LICENSE before use or redistribution. It is also included in the source distribution and served by the running hub at GET /api/v1/license.

Available Tools

43 tools
agent_getA

Return one agent's details incl. last_seen_at. Scoped by reachability: an agent outside your comm scope reads as NOT_FOUND, indistinguishable from a non-existent agent_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent_id to look up. REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose a genuinely important trait: out-of-scope agents return NOT_FOUND, indistinguishable from a missing id. That prevents an agent from misreading a permission failure as non-existence. It does not state read-only status or return shape, but the output schema covers the latter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler; the core action leads and the surprising scoping caveat follows. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema, the description covers the one non-obvious failure mode an agent must know. Return values are handled by the output schema, so nothing critical is missing, though read-only confirmation is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single agent_id parameter is fully documented in the schema (100% coverage, including 'REQUIRED'), so the baseline is 3. The description adds no format, naming, or resolution semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return one agent's details') and names a key returned field (last_seen_at), which is crisper than the bare name. It never explicitly contrasts with agent_list or agent_whoami, so sibling differentiation is only implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this rather than agent_list, agent_whoami, or agent_register. The only qualifying information is a behavioral scoping caveat, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_listA

List registered agents (global), each with role/capabilities and last_seen_at. Authenticated callers see only agents their comm scope can reach (plus themselves); anonymous callers see all.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses a caller-dependent visibility rule (authenticated callers see only agents within their comm scope plus themselves; anonymous callers see all). It omits pagination/ordering behavior, but the output schema exists to cover the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler; the scope ('global') and the visibility rule are front-loaded rather than buried. Every clause carries information the caller needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless list tool with an output schema and no annotations, the description supplies the one thing structured data cannot: who sees which agents. Only pagination/results-size behavior is left unstated, a minor gap given the schema covers the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there is nothing for the description to disambiguate. Baseline 4 applies for a zero-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List registered agents'), names the scope ('global'), and previews the returned fields (role/capabilities, last_seen_at). It does not explicitly distinguish itself from agent_get or agent_whoami, but 'List ... (global)' implies enumeration vs single-record retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to reach for this tool versus agent_get, agent_whoami, or agent_register, nor does it state any preconditions for calling it. Usage context is only weakly implied by the word 'List'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_registerA

Update YOUR OWN profile (role/capabilities/metadata); SELF-ONLY (else PERMISSION_DENIED). Capabilities are fail-closed against the central catalog. Docs: okto-nexus://reference/tool-docs/identity.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoLogical role, e.g. validator, worker (optional); matched exactly/case-sensitively by role-strategy targets.
agent_idYesThe logical agent identity (stable, opaque string); agents are GLOBAL, not per-workspace. REQUIRED.
metadataNoFree-form JSON object of extra attributes stored with the agent (optional).
capabilitiesNoWhat this agent can do - used by capability routing + capability_list discovery (optional). Accepts a flag-map ({"ocr":true}), a list (["ocr","pdf"]), or a single name string. Blank names dropped. FAIL-CLOSED: every name must already exist in the central capability catalog (discover with capability_list).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does meaningful work: it discloses the authorization model (self-only, PERMISSION_DENIED on violation) and the fail-closed capability policy against the central catalog. It omits whether unspecified fields are preserved or cleared, and whether updates are idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: the action, the self-only restriction, and the capability constraint each appear in their own short clause, with the doc link last. No filler, though the clipped style borders on telegraphic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, the description leaves two real gaps for a mutation tool: whether this creates an agent that doesn't yet exist (the name implies registration) and why a required agent_id is needed if the operation is strictly self-only. For a write tool with no annotations, these should be resolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already explains role matching, the global agent_id, free-form metadata, and the three capability input shapes with the fail-closed rule. The description restates the field list and the catalog constraint without adding syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: update the caller's own agent profile, with the affected fields (role/capabilities/metadata) enumerated. It is clearly distinguishable from read-oriented siblings like agent_get, agent_list, and agent_whoami. The tool name 'agent_register' suggests creation rather than update, which slightly muddies the read of the purpose, but the prose is unambiguous about the action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition (SELF-ONLY, otherwise PERMISSION_DENIED) and routes capability discovery to capability_list. It does not, however, state when an agent should call this versus agent_whoami (read) or when a first-time registration versus a subsequent edit applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_whoamiA

Return YOUR OWN profile: agent_id, role, capabilities, metadata, permissions, effective_policies + governance, plus communication style when set. Docs: okto-nexus://reference/tool-docs/identity.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose the return surface (agent_id, role, capabilities, metadata, permissions, effective_policies + governance) and notes communication style is conditional ('when set'). It stops short of stating auth requirements or that the call is strictly non-mutating, but the self-read framing makes that clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence plus a docs pointer; no waste. The field list is dense but each item earns its place by telling the agent what it will receive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless self-profile read backed by an output schema and a docs link, this is essentially complete. The only soft gap is the absent explicit contrast with agent_get/agent_list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4; there is nothing for the description to disambiguate. The enumerated fields are extra value, though the output schema already documents returns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and a precisely scoped resource ('YOUR OWN profile'), and enumerates the fields returned. The 'YOUR OWN' framing cleanly distinguishes it from sibling agent_get/agent_list, which fetch other agents' records.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the self-scoped wording rather than stated. It never explicitly says to use agent_get for other agents or when this is preferable, so the agent must infer the routing from the sibling name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

artifact_getA

Retrieve an artifact by id within the workspace resolved from project_root. Reads are audience-scoped: a caller outside the frozen audience gets NOT_FOUND (indistinguishable from a missing id).

ParametersJSON Schema
NameRequiredDescriptionDefault
artifact_idYesThe artifact_id to retrieve. REQUIRED.
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it discloses a genuinely important trait: reads are audience-scoped and an outside caller receives NOT_FOUND that is deliberately indistinguishable from a missing id. That is valuable context an agent cannot infer from the schema. It does not state permission/auth requirements or whether the operation is read-only in a broader sense, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with the retrieval scope front-loaded and the security caveat second. Every clause earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description covers the non-obvious audience-scoping behavior. For a two-parameter read tool this is nearly complete; only auth/permission expectations remain unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, and the description largely restates the project_root-to-workspace relationship already captured in the schema. It adds no format, validation, or edge-case detail beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Retrieve) and resource (artifact by id) plus the scoping mechanism (workspace resolved from project_root). It is clearly distinguishable from the sibling artifact_put, which performs the inverse operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'Retrieve an artifact by id' but there is no explicit when-to-use statement, no exclusions, and no pointer to alternatives (e.g., artifact_put for writes, or a listing tool). The audience-scoping sentence hints at a failure mode but does not guide tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

artifact_putC

Register a file/text/json/markdown/html artifact in the resolved workspace. Docs: okto-nexus://reference/tool-docs/artifacts.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoHuman-friendly name/label for the artifact (optional).
pathNoFilesystem path to import into Nexus-managed artifact storage (must stay within the workspace root). Provide this OR content - at least one REQUIRED.
contentNoUTF-8 content stored outside the database (bounded by max_inline_bytes; json must be well-formed). Provide this OR path - at least one REQUIRED.
metadataNoFree-form JSON object stored with the artifact (optional).
project_rootYesAbsolute path to the project (defines the workspace scope).
artifact_typeYesArtifact classification - one of: file, text, json, markdown, html. REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about overwrite/duplicate behavior, idempotency, required permissions, size limits, or failure modes. 'Resolved workspace' hints at a dependency on workspace_resolve but does not explain it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with a useful docs pointer and no filler. Slightly terse for a write operation, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with no annotations and no output schema referenced in the description, so ambiguity about what a successful registration returns or does to existing artifacts is significant. The schema covers inputs and the docs link covers unspecified detail, but the description itself leaves the agent without the behavioral context needed to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (including the path-or-content exclusivity rule and the artifact_type enum-in-prose) are already documented in the schema. The description adds no parameter meaning beyond what is structured, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Register') and resource ('artifact') plus the accepted artifact types and the workspace scope. It is clearly distinguishable from artifact_get by direction, though it doesn't explicitly name that sibling as the converse operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as shared_md_render or artifact_get, and no prerequisites or exclusions. The only pointer is an external docs URI, which an agent cannot resolve inline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capability_listB

List the capability catalog merged with owners: every registered name (with description; agent_count 0 if unowned), agents scoped to your comm reach. Normalised as capability routing matches.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It does disclose some useful behavior – unowned capabilities appear with agent_count 0, and agents are filtered to the caller's 'comm reach' – but says nothing about read-only nature, pagination, or what 'Normalised as capability routing matches' implies for the results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, which is good, but the remainder is a colon-glued run-on with parentheticals and a dangling fragment ('Normalised as capability routing matches.') that obscures meaning rather than saving space. It is dense without being clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and no parameters, the description only needs to frame the tool, which it does: it names the catalog, the owner merge, and the agent_count-0 edge case. The unexplained 'comm reach' and 'routing matches' phrasing are the only real gaps for such a low-complexity tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema declares zero parameters at 100% description coverage, so there is nothing for the description to compensate for. Baseline 4 applies for a parameterless list tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('List the capability catalog') and clarifies that it merges owner/agent information, so the agent can tell what data comes back. However, the purpose is diluted by undefined terms like 'comm reach' and 'capability routing matches', and it does not differentiate itself from adjacent listing tools such as agent_list or tag_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives like agent_list, agent_get, or tag_list, nor any stated preconditions or exclusions. Usage must be inferred purely from the word 'List'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

channel_createA

Create a channel by name (idempotent; created=false if it already existed). Channels are organizational labels, not ACLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesChannel name (REQUIRED; short topic label, max 64 chars, unique per workspace). Idempotent by name (existing -> created=false).
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses idempotency, the created=false outcome, and that channels are organizational labels rather than ACLs — a genuinely non-obvious semantic. However, it says nothing about permissions, workspace-scoping errors, or failure behavior, so meaningful gaps remain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and the idempotency caveat. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and both params are fully described. The description covers semantics and idempotency; only authorization/prerequisite context is absent, which is minor for a low-risk creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents both name (max 64 chars, unique per workspace, idempotent) and project_root. The description only restates the idempotency aspect, adding no syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Create) and resource (channel) with scope (by name) and an idempotency qualifier. An agent can distinguish this from channel_list and other channel siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for when to call it — creating a channel by name, tolerating repeats via idempotency. It does not name alternatives (e.g. channel_list) or state exclusions, but the usage window is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

channel_listB

Return the workspace channels (general is seeded by default).

ParametersJSON Schema
NameRequiredDescriptionDefault
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses one useful trait (a default 'general' channel is seeded), which implies a read-only listing, but says nothing about ordering, pagination, archived/hidden channels, or whether it reflects workspace scope. For a zero-annotation tool this leaves significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. The scoping detail is embedded compactly in the parenthetical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure need not be explained, and the single required parameter is fully described in the schema. The description covers the default-seed behavior but omits ordering and whether non-default channels are included, which is minor for a simple list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and project_root is fully documented as the workspace scope in the schema itself. The description adds no further meaning about the parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Return the workspace channels'. An agent can distinguish it from channel_create by the listing verb, but the description never explicitly contrasts it with the sibling create tool. Clear purpose with no explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus channel_create or how the result should be used. The parenthetical about 'general' being seeded is context, not a usage condition. The agent must infer the use case entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

coordination_healthA

Windowed coordination-health report for the workspace: aggregated ok|warn status, 7 metric blocks and thresholds. Requires feature_health and health.read.

ParametersJSON Schema
NameRequiredDescriptionDefault
windowNoReport window - one of: 1h, 24h, 7d (default: 24h).24h
agent_idNoYour agent_id for permission evaluation in open stdio mode (optional; authenticated HTTP MCP uses the API-key identity).
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose key traits: the output is an aggregated ok|warn status over 7 metric blocks with thresholds, and the call requires the feature_health feature and health.read permission. It stops short of explicitly stating it is read-only/no-side-effect, which an agent could still want confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler: the first front-loads what the tool produces and its scope, the second states the hard requirement. Nothing repeats the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, the description needn't enumerate return fields, and it already summarizes the return shape and permission prerequisites. The main remaining gap is the absence of usage context that would tell an agent when to reach for this versus other observability-style read tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the window enum values, agent_id purpose, and project_root scope are already fully documented in the schema. The description only echoes 'windowed' without adding syntax, format, or default guidance beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (coordination-health) and scope (workspace, windowed) plus the concrete output shape (aggregated ok|warn status, 7 metric blocks). No sibling tool overlaps this domain, so differentiation isn't a concern, but it stops just short of the model 5 by describing a report rather than an explicit action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is given. The only conditional information is a prerequisite (requires feature_health and health.read), which is an auth/permission note rather than usage direction. An agent must infer from the name that this is a health-check tool to call proactively.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

event_cursorA

Return the stream's CURRENT END as a cursor (O(1), no scan) - the pre-flight monitor anchor so you see only events appended after now. Returns {cursor} (0 for empty).

ParametersJSON Schema
NameRequiredDescriptionDefault
streamYesEvent stream to read - one of: workspace, agent, handoff. message.created, the message.delivered/message.read receipts and artifact.created are on workspace.
agent_idYesYour agent_id; scopes per-event visibility (you only see events you may see).
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden well: it discloses the O(1)/no-scan cost profile, the anchor/monotonic semantics (only events after now), and the sentinel return (0 for empty). It never states read-only or auth/permission needs, but for a pure cursor read this is a solid disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence with the anchor semantics front-loaded and cost/return details in tight parentheses. Every clause earns its place (O(1), no scan, append-after-now scoping, empty sentinel).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the return shape need not be explained, and params are fully covered by the schema. The description supplies cost and usage framing, giving an agent enough to call it; only explicit sibling routing is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so stream, agent_id, and project_root are already fully documented in the schema (including scope and enum-ish stream values). The description adds no parameter-level meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Return the stream's CURRENT END as a cursor") and adds the operational role ("pre-flight monitor anchor"). It implicitly separates this from event_get/event_wait by framing it as an anchor rather than a reader, but never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Pre-flight monitor anchor so you see only events appended after now" gives a clear when-to-use context: grab this before monitoring so subsequent reads are scoped to new events. No explicit when-not or named alternatives (e.g. event_get vs event_wait) are given, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

event_getA

Read a cursor-paginated page of the event log (non-blocking). Actors outside your comm scope are omitted; yours and system events always show. Docs: okto-nexus://reference/tool-docs/events.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax events per page (optional; default 100, clamped to the server maximum, default 1000).
cursorNoPagination cursor: the last event_id you consumed; returns event_id > cursor (optional; omit/0 = from the beginning).
streamYesEvent stream to read - one of: workspace, agent, handoff. message.created, the message.delivered/message.read receipts and artifact.created are on workspace.
filtersNoEquality filters, AND-combined (optional). Keys: type, agent_id, task_id, handoff_id, trace_id. e.g. {"type":"message.created"}.
profileNoResponse size profile - one of: default, summary, full (optional; summary trims per-call tokens).
agent_idYesYour agent_id; scopes per-event visibility (you only see events you may see).
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses non-blocking behavior, the visibility/comm-scope filtering rule, and cursor pagination. It omits auth/prerequisite requirements and any rate-limit or error behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences plus a docs pointer; the core purpose and the most surprising behavior (non-blocking, scope filtering) are front-loaded with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers pagination, blocking behavior, and visibility. What remains thin is guidance on choosing among the many sibling event/inbox readers, but nothing required to invoke the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 7 parameters including cursor, limit, stream, and filters. The description adds only the high-level 'cursor-paginated page' framing, not parameter-specific semantics beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (a page of the event log) plus a distinguishing behavioral trait: non-blocking. It implicitly separates itself from event_wait, but it never names that sibling or event_cursor explicitly, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(non-blocking)' tag implies you use this when you don't want to wait, but no when-to-use/when-not statement or explicit alternative is given. The visibility rule ('actors outside your comm scope are omitted; yours and system events always show') is useful context but is about results, not tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

event_waitA

Read the event log; optionally long-poll (0/omitted/null = snapshot; >0 blocks). Scoped like event_get (actors outside your comm scope omitted). Patterns: okto-nexus://reference/monitoring.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax events per page (optional; default 100, clamped to the server maximum, default 1000).
cursorNoPagination cursor: the last event_id you consumed; returns event_id > cursor (optional; omit/0 = from the beginning).
streamYesEvent stream to read - one of: workspace, agent, handoff. message.created, the message.delivered/message.read receipts and artifact.created are on workspace.
filtersNoEquality filters, AND-combined (optional). Keys: type, agent_id, task_id, handoff_id, trace_id. e.g. {"type":"message.created"}.
profileNoResponse size profile - one of: default, summary, full (optional; summary trims per-call tokens).
agent_idYesYour agent_id; scopes per-event visibility (you only see events you may see).
project_rootYesAbsolute path to the project (defines the workspace scope).
timeout_secondsNoLong-poll bound in SECONDS (optional; default 0). 0/omitted/null = an immediate non-blocking snapshot; >0 OPTS IN to a BLOCKING long-poll until an event arrives or the timeout elapses.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does a solid job: it discloses the visibility scoping ("actors outside your comm scope omitted"), the blocking semantics of the timeout, and points to a monitoring reference. It stops short of permissions or rate-limit context, but the key behavioral traits of a long-poll read are surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the primary action and the blocking rule, with no filler. Slightly dense but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter read tool with an output schema and full schema coverage, the description supplies the missing behavioral context (scoping, blocking, monitoring reference) without needing to explain return values. Adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all eight parameters are already documented, including the 0/omitted/null snapshot semantics for timeout and the limit/cursor meaning. The description merely restates the timeout rule rather than adding new syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Read the event log") and disambiguates itself from the sibling event_get by noting the scoping relationship and by being the tool that offers long-polling. An agent can tell what it does, though the distinction from event_get/event_cursor is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The blocking-vs-snapshot behavior is explained, which implies when you'd reach for this tool, but there is no explicit statement of when to prefer it over event_get (plain read) or event_cursor, and no exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_cancelA

Creator-only OPEN -> CANCELLED; retract a handoff nobody should take (e.g. a pool target matching zero agents). Only OPEN handoffs cancel. In strict mode pass session creds.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoHuman-readable reason recorded with the handoff.cancelled event (optional).
agent_idYesYour agent_id; must be the handoff's creator. REQUIRED.
handoff_idYesThe handoff_id to act on. REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
project_rootYesAbsolute path to the project (defines the workspace scope).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses the creator-only authorization requirement, the state precondition (OPEN only), and that strict-mode trust requires passing session creds. It omits failure behavior and whether cancellation is reversible, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely front-loaded and dense: state transition, purpose, precondition, and auth note in three short clauses. The telegraphic style is efficient, though the clipped phrasing slightly risks ambiguity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need no explanation. For a mutating, auth-sensitive tool the description covers the state constraint, authorization model, and strict-mode requirement adequately, leaving only edge-case failure behavior unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents agent_id, handoff_id, session_id, and session_secret in detail. The description reinforces the creator constraint and strict-mode credential requirement, but adds little syntax or format detail beyond the schema, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with an explicit state transition ('OPEN -> CANCELLED') and a concrete purpose ('retract a handoff nobody should take'). An agent can distinguish it from handoff_reject and handoff_complete from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context with a concrete example ('a pool target matching zero agents') and a precondition ('Only OPEN handoffs cancel'). It does not explicitly name the alternative tools (reject/complete) or when NOT to cancel, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_claimA

Atomically claim an OPEN handoff; single winner, others get a structured error. Returns the payload + claimed_by/lease_expires_at. In strict mode pass session_id + session_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesYour agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED.
handoff_idYesThe handoff_id to act on. REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
project_rootYesAbsolute path to the project (defines the workspace scope).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses atomicity, the single-winner outcome, the structured error for losers, the auth requirement in strict mode, and the returned fields (payload + claimed_by/lease_expires_at), which implies a lease/ownership model. It stops short of stating lease duration, renewal expectations, or whether a claim can be released, which would matter for a state-mutating coordination primitive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core behavior, then the return contract, then the conditional auth note. No filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail need not be expanded, and the description correctly covers concurrency semantics, auth modes, and failure shape. What's missing is post-claim lifecycle context (what the lease implies, and that handoff_complete/reject/cancel are the follow-up actions), which would help an agent plan the next step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema. The description's mention of session_id + session_secret in strict mode restates what the schema descriptions already say, adding no new parameter semantics. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (claim), resource (handoff), and precondition (OPEN), plus the key semantic property (atomic, single winner). An agent can immediately distinguish it from handoff_create, handoff_list_available, handoff_get, and handoff_complete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit conditional guidance for trust_mode=strict (pass session_id + session_secret), which tells the agent when extra credentials are needed. It does not, however, contrast itself against adjacent siblings like handoff_get or handoff_verify, so the when-not guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_completeA

Owner-only delivery of a CLAIMED handoff: -> COMPLETED, or -> VERIFYING when acceptance_criteria were set (verifier decides via handoff_verify). In strict mode pass session creds.

ParametersJSON Schema
NameRequiredDescriptionDefault
resultNoCompletion result (optional; string or JSON) persisted on the handoff, recorded with handoff.completed, and delivered to the creator inbox.
agent_idYesYour agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED.
handoff_idYesThe handoff_id to act on. REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
project_rootYesAbsolute path to the project (defines the workspace scope).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the ownership restriction and the strict-mode session-credential requirement, which is valuable. However, it omits irreversibility, error behavior, and the side effect that the result is persisted and delivered to the creator inbox (that detail only lives in the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the key state-transition and ownership facts with no wasted words. The terse arrow notation is efficient, though slightly cryptic for an agent parsing on first read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating handoff tool with an output schema covering returns, the description supplies ownership, transition logic, the verify path, and auth mode. That is nearly sufficient, with only irreversibility/error semantics left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents agent_id, handoff_id, session_id/secret, result, and project_root. The description adds the strict-mode credential hint but otherwise repeats what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('delivery of a CLAIMED handoff') and spells out the resulting state transitions (-> COMPLETED, or -> VERIFYING). It is clearly distinguishable from siblings like handoff_verify, handoff_reject, and handoff_cancel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the branching condition for COMPLETED vs VERIFYING and names handoff_verify as the decider when acceptance_criteria were set, plus the owner-only constraint. It stops short of explicitly stating when to prefer handoff_cancel or handoff_reject instead, so no full when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_createA

Create an OPEN handoff (validates target/visibility); emit handoff.created. After creating, poll handoff_get for status/result. Full docs: okto-nexus://reference/tool-docs/handoff.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesRouting target as a raw JSON object - which agents may CLAIM this handoff (REQUIRED). strategy one of: direct {"strategy":"direct","agent_id":"<id>"}; capability {"strategy":"capability","capability":"<cap>"}; role {"strategy":"role","role":"<role>"}; tag {"strategy":"tag","selector":{"<key>":["<value>",...]}} (flat: AND across keys, OR within values) or rich [{"key":"<k>","operator":"In|NotIn|Exists|DoesNotExist","values":["<v>",...]}] (ANDed; CAUTION: NotIn/DoesNotExist also match agents MISSING the key); broadcast {"strategy":"broadcast"}; mixed {"strategy":"mixed","rules":[<sub-target>,...]}; direct_with_fallback {"strategy":"direct_with_fallback","agent_id":"<id>","fallback_after_seconds":<n>}. Competing-consumers: the first to handoff_claim wins. Hierarchy/catalog rules, examples, edge-cases: okto-nexus://reference/target-grammar.
payloadNoInline work content (optional; raw JSON object/array, string, or null - NOT JSON-encoded). Returned only to the claimant by handoff_claim / claimant handoff_get. For large content pass an artifact_id.
trace_idNoTrajectory trace_id to stamp on this handoff (optional; non-empty string, max 128 chars). Needs the feature_trace flag ON, else accepted and ignored; omitted = generate one.
verify_byNoWho verifies (optional; only WITH acceptance_criteria; raw JSON object). one of: {"kind":"creator"} (default); {"kind":"agent","agent_id":"<id>"} (must be registered); {"kind":"capability","capability":"<name>"} (in the catalog; resolved at verify time). The claimant never verifies their own delivery.
depends_onNoHandoff ids this one depends on (optional; raw JSON list of 1..20 unique existing ids; IMMUTABLE). Blocked - unlisted/unclaimable - until ALL are COMPLETED. Needs feature_dag ON (rejected while OFF).
session_idNoSession_id attributing this operation to a specific open session of yours (optional).
visibilityYesWho may SEE the handoff (separate from who may CLAIM it = target). one of: public, eligible, private. REQUIRED (case-insensitive).
project_rootYesAbsolute path to the project (defines the workspace scope).
from_agent_idYesYour agent_id (the creator); recorded as the handoff's originator - the owner is whichever agent later claims it (handoff_claim), not necessarily you.
acceptance_criteriaNoVerification contract (optional; raw JSON list of 1..20 unique non-empty strings, max 500 chars each; IMMUTABLE). handoff_complete then parks the handoff in VERIFYING until the handoff_verify verdict. Needs feature_verification ON (rejected while OFF, never ignored).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that target/visibility are validated, that a 'handoff.created' event is emitted, and that status/result should be polled via handoff_get. However, it omits permissions, error handling, or side effects like ownership transfer, leaving significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tightly written sentences: front-loaded action, validation, event, next step, and a documentation link. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters, no annotations, and an output schema that covers return values, the description is partially complete: it names the creation, validation, event emission, and follow-up poll. But it omits failure modes, required permissions, and other behavioral traits needed for a complex coordination tool, relying heavily on the external doc link.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters thoroughly. The description only reiterates that target/visibility are validated and adds no parameter syntax or meaning beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create') and resource ('handoff') with additional state ('OPEN') and validation scope ('validates target/visibility'). It clearly identifies the operation, though it does not explicitly differentiate from siblings like handoff_claim beyond the verb 'Create'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use by stating the next step ('After creating, poll handoff_get for status/result'), but gives no explicit when-not conditions or alternative selection guidance relative to other handoff tools (claim, complete, etc.). Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_getA

Read a handoff by id: status, claimant, payload, result/rejected_reason + verification/dependency fields if set. The creator's path to the outcome. Full docs: okto-nexus://reference/tool-docs/handoff.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesYour agent_id (REQUIRED). Creator and claimant always read; others are gated by the handoff visibility.
handoff_idYesThe handoff_id to act on. REQUIRED.
project_rootYesAbsolute path to the project (defines the workspace scope).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the burden. It signals this is a read and notes conditional fields ("if set"), but says nothing about the visibility gating described in the schema, permissions, or auth requirements. An output schema exists, so return-format explanation is not needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that front-load the verb and resource, with the field list and doc pointer kept short. The trailing documentation link is mildly expendable but harmless.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with a full output schema and complete parameter coverage, the description covers purpose and return fields adequately. The only real gap is the absence of any behavioral note on visibility/access, which the schema partially covers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented there, so the baseline is 3. The description adds no syntax or format detail beyond restating that the read is by id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Read a handoff by id") and enumerates the fields returned (status, claimant, payload, result/rejected_reason), which distinguishes it from list-oriented siblings like handoff_list_available. It doesn't explicitly name the sibling it contrasts with, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"The creator's path to the outcome" hints at context of use, but there is no explicit when-to-use guidance, no when-not, and no named alternatives (e.g. handoff_list_available for discovery). Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_list_availableA

Expire leases, then list OPEN handoffs visible+eligible to the caller (paginated); entries are metadata-only until claimed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax handoffs per page (optional; default applied by the server, clamped to the maximum).
cursorNoOpaque pagination cursor: pass back the next_cursor from the previous page to continue (optional; omit or '0' to start from the beginning).
agent_idYesYour agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED.
project_rootYesAbsolute path to the project (defines the workspace scope).
timeout_secondsNoLong-poll bound in SECONDS (optional). 0/omitted = non-blocking snapshot; >0 OPTS IN to blocking until a claimable handoff appears or the timeout elapses.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose the non-obvious side effect (expiring leases before listing) plus the 'metadata-only until claimed' constraint, which an agent would not infer from the name. It omits permission/auth requirements and any note on whether lease expiration is destructive or affects other callers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, front-loaded with the side effect and the action, with no filler. The semicolon clause packs scope, pagination, and the metadata-only caveat efficiently, though it is a lot compressed into one sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be explained. The description covers the key behavioral facts (lease expiration, eligibility scoping, pagination, metadata-only pre-claim state) needed to call it correctly; only auth/permission expectations are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so limit/cursor/agent_id/project_root/timeout_seconds are already fully documented; the description only reinforces 'paginated' and caller-scoped visibility. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (OPEN handoffs) plus a scope qualifier ('visible+eligible to the caller'), which separates it from handoff_get/handoff_claim. It also names a side effect ('Expire leases'), but does not explicitly contrast with sibling list/claim tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'visible+eligible to the caller' and the paginated metadata-only framing, but there is no explicit when-to-use versus handoff_get or handoff_claim, and no stated preconditions. The long-poll semantics are left to the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_rejectA

Reject a handoff (owner CLAIMED->REJECTED or direct-target OPEN->REJECTED). In trust_mode=strict pass session_id + session_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoHuman-readable reason (optional) persisted on the handoff, recorded with the rejection, and delivered to the creator inbox.
agent_idYesYour agent_id (the worker); scopes visibility/eligibility and ownership. REQUIRED.
handoff_idYesThe handoff_id to act on. REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
project_rootYesAbsolute path to the project (defines the workspace scope).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden, and it does disclose the auth requirement in trust_mode=strict. It does not cover reversibility, idempotency, error conditions, or what happens to the counterparty beyond what the schema's reason field already says.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the core action and state transitions, then the auth caveat. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a full output schema, the description covers purpose, state transitions, and auth preconditions adequately. Remaining gaps (reversibility, side effects) are minor given the rich schema and output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema, including the trust_mode conditional on session_id/session_secret. The description merely restates that auth rule, adding no meaning beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reject) plus the resource (handoff) and even enumerates the exact state transitions it performs (CLAIMED->REJECTED, OPEN->REJECTED). This cleanly distinguishes it from siblings like handoff_cancel, handoff_complete, and handoff_verify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The state transitions imply when the tool applies, and the second sentence gives an auth precondition for strict mode. However, there is no explicit guidance on when to reject versus cancel/complete a handoff, nor any stated exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_verifyB

Verifier-only verdict on a VERIFYING handoff: 'pass' -> COMPLETED (verified_by), 'fail' -> CLAIMED for rework (feedback + renewed lease). In strict mode pass session creds.

ParametersJSON Schema
NameRequiredDescriptionDefault
verdictYesThe verdict. one of: pass (VERIFYING -> COMPLETED; handoff.completed carries verified_by), fail (VERIFYING -> CLAIMED for rework, lease renewed). REQUIRED (lowercase).
agent_idYesYour agent_id (the verifier); must satisfy the handoff's verify_by resolved AT VERIFY TIME. The claimant (claimed_by) is always refused. REQUIRED.
feedbackNoRework guidance (optional; max 2000 chars; only with verdict 'fail'). Persisted (each fail overwrites the previous) and delivered to the claimant's inbox.
handoff_idYesThe handoff_id to act on. REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
project_rootYesAbsolute path to the project (defines the workspace scope).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the two verdict outcomes, that feedback overwrites previous feedback and goes to the claimant's inbox, and the strict-mode credential requirement. It omits error/edge behavior (wrong state, idempotency, concurrency) that matters for a state-mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads the core behavior, but the telegraphic style ('pass session creds', arrow notation) sacrifices readability, and the opening clause is dense enough that an agent must re-read it to parse the state transitions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers the primary path. For a 7-parameter, multi-mode (open vs strict trust) verification tool with no annotations, however, more operational context (preconditions, failure behavior) would be needed for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents every parameter including feedback and the session credentials. The description's only added param meaning is a restatement of the strict-mode session requirement ('pass session creds'), which the schema already states, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('verifier-only verdict on a VERIFYING handoff') and spells out both state-machine outcomes, which lets an agent distinguish it from handoff_complete, handoff_reject, and handoff_cancel. It is clear but relies on terse jargon rather than plainly naming the siblings it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Verifier-only' and the VERIFYING precondition imply when the tool applies, and the agent_id note ('the claimant is always refused') narrows the caller. But it never explicitly contrasts this with the sibling verify/complete/reject tools or states what to do when the handoff is in another state, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inbox_ackA

Acknowledge messages into history (read). Returns {acknowledged, read_message_ids}. Emits a message.read receipt to each sender.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesYour agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
message_idsYesThe message_id(s) to acknowledge - a single id or a list. REQUIRED. Idempotent (already-read/foreign ids are no-ops).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose two real traits: it emits a message.read receipt to each sender (a visible side effect) and returns a specific result shape. It stops short of permissions, rate limits, or what happens to unacknowledged messages.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, no filler, purpose and return shape and side effect front-loaded in that order. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, yet the description still names the response keys. Auth and idempotency are covered by a rich schema. The main gap is the absence of when-to-use routing among the inbox_* siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so agent_id, session_id, message_ids and session_secret are fully documented in the schema, including idempotency and trust-mode requirements. The description adds nothing about parameters, so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: acknowledge messages into history, with a parenthetical clarifying it counts as a read. This is clear enough to distinguish it from inbox_pull/inbox_peek, but it never names those siblings or explains how ack differs from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to acknowledge versus pulling, peeking, or reading history. The idempotency note lives in the schema, not the description, and there is no mention of prerequisites or when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inbox_countA

Return your inbox lane sizes {unread, in_flight, read}. Cheap READ-ONLY between-turns check (pull when unread > 0). Expired in-flight leases count as unread; parked excluded.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesYour agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it declares the operation read-only, characterizes cost ('cheap'), and discloses non-obvious counting semantics (expired in-flight leases count as unread; parked excluded). It doesn't cover auth requirements or rate limits, but for a read-only counter this is solid disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, zero filler, with the return shape front-loaded and the usage rule and edge-case semantics trailing. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need not be re-explained; the description supplies the selection context and the subtle counting rules an agent needs to interpret the numbers correctly. Complete for a single-param read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single agent_id parameter is fully documented in the schema (global inbox, reaches you in any workspace, REQUIRED). The description adds nothing about the parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (inbox lane sizes) and enumerates the exact returned fields {unread, in_flight, read}. An agent can distinguish this counting tool from inbox_pull/inbox_peek/inbox_history without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage context ('Cheap READ-ONLY between-turns check') and a routing rule ('pull when unread > 0') that points to the sibling inbox_pull. No when-not-to-use guidance is given, but the positive selection rule is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inbox_extendA

Renew the lease on in-flight messages you pulled but have not finished (now + extend_seconds). All-or-nothing: if any id is not in-flight the call fails per-message and nothing is extended.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesYour agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
message_idsYesThe message_id(s) whose lease to renew - id or list. REQUIRED. All must be in-flight; otherwise the call fails per-message and nothing is extended.
extend_secondsYesNew lease duration in seconds, counted from now (REQUIRED; clamped to 10..3600).
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and discloses the key atomicity trait: all-or-nothing, with per-message failure and no extensions if any id is not in-flight. It also clarifies the lease is extended from now. It does not discuss authentication or return values, but the output schema handles returns and the input schema documents session/trust requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences with no filler. The core action is front-loaded, followed by the critical failure condition. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 5-parameter mutation-like tool with 100% schema coverage, an output schema, and no annotations, the description covers the essential behavior: lease renewal and all-or-nothing failure. It does not mention trust-mode auth requirements, but those are fully documented in the schema, so the description is largely complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all parameters, including extend_seconds clamping and session trust modes. The description adds only mild context via '(now + extend_seconds)' and the all-must-be-in-flight constraint, which is also in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: renew the lease on in-flight messages. It also distinguishes this from ack/finish behavior by specifying messages 'you pulled but have not finished.' An agent can clearly identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for use: messages that are in-flight but not yet finished need their lease renewed. It does not explicitly name alternatives like inbox_ack or state when not to use this tool, so it falls short of full when/when-not guidance, but the implied usage is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inbox_historyA

List your acknowledged (read) messages, newest-first, keyset-paginated (stable pages even while you keep acknowledging). READ-ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax messages to return (optional; default 50, max 200; must be a positive integer).
cursorNoOpaque pagination cursor: pass back the previous page next_cursor (optional; omit = from the newest read messages).
profileNoResponse size profile - one of: default, summary, full (optional; summary trims per-call tokens).
agent_idYesYour agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden: it declares READ-ONLY, states the sort order, and discloses a non-obvious behavioral guarantee (keyset pagination keeps pages stable even while new messages are acknowledged). It stops short of covering auth/permission requirements or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence; the operation and its scope come first, and the read-only declaration is a clean trailing flag. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and all four parameters are documented at 100%. The remaining gap is minor: no note on auth/permission expectations for reading the global inbox, but the essentials for correct invocation are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit, cursor, profile and agent_id in detail. The description adds only the conceptual meaning of the cursor (keyset pagination), which is useful framing but not parameter-level detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (List), resource (acknowledged/read messages), ordering (newest-first) and pagination model. It implicitly separates itself from siblings like inbox_peek by scoping to already-acknowledged messages, but it never names an alternative tool explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (you call this to review what you've already acknowledged), and the parenthetical about stable pages while acknowledging hints at a concurrent-workflow use case. However there is no explicit when-to-use/when-not-to-use guidance or reference to a sibling alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inbox_peekA

Triage pending messages (unread + in-flight) WITHOUT consuming. READ-ONLY, envelope-only by default (body_preview + body_bytes). include_parked/include_bodies opt in.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax messages to return (optional; default 50, max 200; must be a positive integer).
profileNoResponse size profile - one of: default, summary, full (optional; summary trims per-call tokens).
agent_idYesYour agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED.
include_bodiesNoReturn the FULL body instead of the envelope-only body_preview (first 200 chars) + body_bytes (optional; default false; inbox_pull is the way to consume a message).
include_parkedNoAlso show parked (dead-letter) messages that exhausted redelivery (optional; default false).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares READ-ONLY, states the default payload (body_preview + body_bytes), and flags which behaviors require explicit opt-in. It stops short of describing ordering, pagination, or what 'in-flight' concretely means for a caller.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense clauses with zero filler; the non-consuming read-only framing is front-loaded before the default and opt-in details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description covers the read-only posture, defaults, and opt-ins. It is nearly complete, missing only hints about result ordering/volume expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents limit, profile, agent_id, include_bodies and include_parked in detail. The description only restates the include_parked/include_bodies opt-in defaults, adding little beyond the structured fields — the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('triage/peek') and resource ('pending messages: unread + in-flight') and immediately scopes it as non-consuming, which cleanly separates it from the sibling inbox_pull. An agent can pick this over inbox_pull without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to use it (triage without consuming, envelope-only by default) and notes the opt-in flags for deeper inspection. It never names inbox_pull as the consuming alternative in the description itself, so the when-not is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inbox_pullA

Take your unread messages into in-flight and return them WITH body (index-free; no cursor). At-least-once: unacked pulls are redelivered. Docs: okto-nexus://reference/tool-docs/inbox.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax messages to return (optional; default 50, max 200; must be a positive integer).
profileNoResponse size profile - one of: default, summary, full (optional; summary trims per-call tokens).
agent_idYesYour agent_id - the GLOBAL inbox to read (a direct message reaches you in any workspace). REQUIRED.
session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED together with session_secret in trust_mode=strict).
lease_secondsNoLease (seconds) before pulled messages are redelivered (optional; default 300, clamped 10..3600); renew mid-turn with inbox_extend.
session_secretNosession_secret from session_open for session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the key trait: at-least-once delivery with unacked pulls being redelivered, plus the fact that bodies are returned. It omits auth/trust-mode behavior and what happens on ack, but the redelivery contract is the critical disclosure and it is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses with no waste; the core semantics and the at-least-once caveat are front-loaded, and the docs pointer is a compact trailing reference instead of prose bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 6 params covered at 100%, an output schema present, and the delivery guarantee stated, an agent has what it needs to call this correctly. Only the trust-mode/auth nuance and how the returned messages differ from inbox_peek output are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so limit, profile, lease_seconds, session_id, and session_secret are already fully documented in the schema. The description adds no parameter-level detail beyond implying the lease/renew flow, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('take your unread messages into in-flight and return them WITH body') and adds a differentiating qualifier ('index-free; no cursor') that separates it from cursor-based readers like event_cursor. It does not name the closest sibling (inbox_peek), so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (consume unread mail as in-flight messages) and points to inbox_extend for mid-turn lease renewal, but never states when to prefer this over inbox_peek, inbox_history, or message_list. Usage is inferred rather than guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_createA

Persist a message and fan it out to recipient inboxes; emit message.created. The response confirms delivery (recipients + delivered_count). Full docs: okto-nexus://reference/tool-docs/messages.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesMessage body (inline text). For large content, attach an artifact and keep body a short pointer.
targetNoRouting target as a raw JSON object (optional; omit = broadcast to present agents). strategy one of: direct {"strategy":"direct","agent_id":"<id>"}; capability {"strategy":"capability","capability":"<cap>"}; role {"strategy":"role","role":"<role>"}; tag {"strategy":"tag","selector":{"<key>":["<value>",...]}} (flat: AND across keys, OR within values) or rich [{"key":"<k>","operator":"In|NotIn|Exists|DoesNotExist","values":["<v>",...]}] (ANDed; CAUTION: NotIn/DoesNotExist also match agents MISSING the key); broadcast {"strategy":"broadcast"}; mixed {"strategy":"mixed","rules":[<sub-target>,...]}. Hierarchy/catalog rules, examples, edge-cases: okto-nexus://reference/target-grammar.
subjectYesShort message subject/title (one line).
trace_idNoTrajectory trace_id to stamp on this message (optional; non-empty string, max 128 chars). Needs the feature_trace flag ON, else accepted and ignored; omitted = inherit the reply parent's trace, or generate one.
artifactsNoList of artifact_id strings to attach (optional; at most 20, no exact duplicates; reference large content instead of inlining it in body).
channel_idNoChannel_id to post into (optional; omit = no channel). Organizational label only; does NOT decide recipients (the target does).
project_rootYesAbsolute path to the project (defines the workspace scope).
from_agent_idYesYour agent_id (the sender); recorded as the author - recipients reply by targeting it.
session_secretNosession_secret from session_open for from_session_id (optional in open mode but VALIDATED if supplied; REQUIRED in strict mode).
from_session_idNoYour session_id from session_open (optional in trust_mode=open; REQUIRED with session_secret in trust_mode=strict).
parent_message_idNoMessage_id this is a reply to, to thread the conversation (optional).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well on side effects: it discloses persistence, fan-out to recipient inboxes, emission of a message.created event, and that the response confirms delivery with recipients + delivered_count. It omits auth/trust-mode requirements and any failure/partial-delivery semantics, so it falls short of 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight clauses: the core action and side effect are front-loaded, the response contract follows, and the docs pointer is last. Every sentence earns its place with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and the schema richly covers routing, session, and artifact semantics. For an 11-parameter tool the description is somewhat thin on auth/trust-mode context, but the docs reference plus the schema make it usable end-to-end.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters, including the complex target routing grammar. The description adds no parameter-level meaning beyond what the schema provides, which is the baseline 3 for full-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states specific verbs and a resource ('persist a message and fan it out to recipient inboxes; emit message.created'), so an agent immediately knows this is the write/send path. It does not explicitly name read siblings (message_get, message_list, message_wait) to differentiate, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-tool guidance. The description explains the mechanics of sending but never states the conditions under which an agent should call message_create rather than inbox_pull, message_status, or channel tools; it only points to external docs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_getA

MIGRATED (S3): replaced by inbox_pull / inbox_peek / inbox_history. Always returns ok:false code=MIGRATED with the replacement call.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so: it discloses that the tool always fails, the exact error code, and that the response contains the replacement call. That is complete transparency for a tombstone tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence stating the migration, the replacements, and the return behavior. Every clause earns its place with no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, the description already explains the only meaningful return behavior (ok:false code=MIGRATED with the replacement call). For a zero-parameter migrated stub, this is fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The schema has 100% coverage and the description correctly implies no input is needed, though there is no additional parameter meaning to add.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool now is: a migrated stub that always returns ok:false with code=MIGRATED. It names the three replacement tools (inbox_pull / inbox_peek / inbox_history), so an agent can immediately distinguish it from those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent not to use this tool and to use the replacements instead, with the condition that any call will return a MIGRATED error. No inference is required about when or whether to call message_get.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_listA

MIGRATED (S3): replaced by inbox_peek / inbox_history (your messages) and event_get (bus traffic). Always returns ok:false code=MIGRATED.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and meets it: it discloses the exact return contract ('Always returns ok:false code=MIGRATED'), so an agent knows the call is a guaranteed no-op failure and can avoid it entirely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence beginning with 'MIGRATED' conveys status, replacements, and exact behavior with zero filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deprecated stub, the description covers everything an agent needs: deprecation status, replacement routing, and the deterministic failure response. An output schema exists, but the description still correctly summarizes the return contract rather than leaving the agent to discover it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty at 100% coverage, so there is nothing for the description to clarify. Baseline 4 applies; no parameter discussion is needed or given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what this tool now is: a migrated tombstone that always returns ok:false code=MIGRATED. It names the sibling replacements (inbox_peek, inbox_history, event_get) so an agent can immediately tell this tool apart from them without inspecting any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: it says the tool is replaced and maps each use case to its correct alternative — inbox_peek/inbox_history for 'your messages' and event_get for 'bus traffic'. There is no ambiguity about when not to call this tool and what to call instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_statusA

Track a message you SENT: per-recipient delivery states {recipient, status, attempts, read_at} (unread/delivered/read/parked). READ-ONLY.

ParametersJSON Schema
NameRequiredDescriptionDefault
message_idYesThe message_id from message_create whose per-recipient delivery states you want to track. REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose that the operation is READ-ONLY and enumerates the possible status values, which is genuinely useful. However it says nothing about permissions, whether status is eventually consistent, or how 'parked' differs operationally, leaving real behavioral gaps for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight clauses with the scope ('you SENT') and the read-only nature front-loaded; every element carries information. The inline set notation is dense but readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema, the description supplies scope, the return shape's key fields, and the read-only guarantee, which is enough to call it correctly. The output schema relieves it of explaining return values in detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter's origin (message_id from message_create) is already documented in the schema. The description adds no format or constraint details beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: tracking per-recipient delivery states for a message the caller SENT. The 'you SENT' scoping and the enumerated states (unread/delivered/read/parked) make it distinguishable from generic siblings like message_get or message_list, though no sibling is named directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Track a message you SENT' implies the usage context and requires a message_id from message_create, but there is no explicit when-to-use vs. when-not guidance and no alternative tool named for adjacent needs (e.g., reading message content vs. tracking its delivery).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_waitA

MIGRATED (S3): replaced by inbox_count polling (cheap) or event_wait (explicit blocking). Always returns ok:false code=MIGRATED.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well by disclosing the deterministic failure shape (ok:false, code=MIGRATED), so an agent knows calling it is pointless. It stops short of saying whether the tool is still callable at all or will be removed, but the essential behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, migration status and replacement named first, mechanical outcome second. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter deprecation shim with an output schema present, this is fully sufficient: it tells the agent what it is, what it returns, and where to go instead. Nothing needed to decide whether to invoke it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is nothing to disambiguate. The description correctly implies no inputs are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool is (a migrated/deprecated stub) and what it does when called (always returns ok:false code=MIGRATED). It also names the two replacement tools, so an agent can distinguish this from event_wait and inbox_count without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to alternatives and gives the selecting condition for each: inbox_count for cheap polling, event_wait for explicit blocking. Effectively a when-not-to-use instruction, which is the strongest form of guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexus_infoA

Report server versions: package_version, schema_version, surface_revision, resource_versions, features (read-only {feature_*: bool}). Call when behaviour disagrees with cached schemas.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It implies a passive report and notes the features map is 'read-only {feature_*: bool}', but never states side-effect profile, permissions, or cost. With an output schema present the return values needn't be explained, so this is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the enumerated payload is front-loaded and the usage trigger follows. Dense and jargon-heavy, but every clause carries information and nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape is covered; the description still summarizes the payload and gives a trigger. For a zero-parameter read-only info tool this is close to complete, with only the safety/permission profile left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 (plus schema coverage is 100%). There is no parameter meaning the description could add.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Report server versions') and enumerates the exact fields returned (package_version, schema_version, surface_revision, resource_versions, features). It is clearly distinguishable from siblings like capability_list or coordination_health, though it does not name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete trigger: 'Call when behaviour disagrees with cached schemas.' That is a real when-to-use condition rather than implied usage. It lacks any when-not guidance or named alternatives, which keeps it out of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

poll_token_issueA

Issue an ephemeral read-only monitor bearer (nxsept_...) for this authenticated agent's session workspace. Store only in the background monitor; never persist the raw token.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYessession_id returned by session_open. REQUIRED.
session_secretYessession_secret returned by session_open. REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load and does disclose real traits: the token is ephemeral, read-only in scope, prefixed nxsept_, bound to the authenticated session, and must not be persisted. It stops short of stating lifetime/expiry or that poll_token_renew exists for rotation, which are the remaining behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and the security constraint second; nothing is wasted and no sentence restates the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description covers purpose, scope, and handling. It is only marginally incomplete in omitting token lifetime and the renewal path, which an agent would need to avoid misusing an expired token.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are described in the schema as values 'returned by session_open. REQUIRED.' The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Issue) plus resource (ephemeral read-only monitor bearer) and scope (this authenticated agent's session workspace). The nxsept_ prefix and 'ephemeral read-only' framing distinguish it from poll_token_revoke and poll_token_renew, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Store only in the background monitor; never persist the raw token' gives handling guidance but not selection guidance — it never says when to call this versus poll_token_renew or poll_token_revoke, nor what precondition (an open session) triggers it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

poll_token_renewA

Rotate and extend the active ephemeral poll token for this session. The previous raw nxsept_ bearer stops working immediately.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYessession_id returned by session_open. REQUIRED.
session_secretYessession_secret returned by session_open. REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the single most important behavioral trait: "The previous raw nxsept_ bearer stops working immediately" — a specific, non-obvious invalidation side effect. It omits auth/permission requirements and failure modes (e.g. no active token), which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary action and followed by the single highest-value behavioral fact. No filler, no redundancy, every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (2 required params, full schema coverage, output schema present so return values need no explanation) and the description covers the key lifecycle consequence of renewing. It is close to complete, missing only explicit preconditions and when-to-use framing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the schema already documents session_id and session_secret (including their provenance from session_open). The description adds no format, constraint, or sensitivity detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb pair plus resource: "Rotate and extend the active ephemeral poll token for this session." That is well beyond a restatement of the name and clearly separates renew from issue/revoke. It stops short of naming those siblings explicitly, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the phrase "the active ... token for this session" tells the agent a token must already exist, but there is no explicit when-to-call, no precondition statement, and no routing against poll_token_issue or poll_token_revoke. Minimal viable guidance, nothing more.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

poll_token_revokeB

Revoke the active ephemeral poll token for this session.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYessession_id returned by session_open. REQUIRED.
session_secretYessession_secret returned by session_open. REQUIRED.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether revocation is immediate or permanent, whether an already-revoked or missing token errors, that it requires the session_secret, or what effect revocation has on in-flight polling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every word contributes to identifying the action and its scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema removes the need to describe return values, and the tool is a simple two-parameter revoke. Still, with no annotations and no usage or behavioral context, the description is barely sufficient for an agent to know when revocation is appropriate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (session_id, session_secret) are documented in the schema as required values returned by session_open. The description adds no extra meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource (revoke the active ephemeral poll token) and scopes it to this session, so the action is unambiguous. It does not, however, distinguish itself from the sibling poll_token_issue and poll_token_renew tools, leaving the token lifecycle relationship implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no mention of the sibling tools poll_token_issue/poll_token_renew or when revoking should be preferred over letting a token expire. Usage can only be inferred from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_closeB

Close a session (idempotent); repeating returns ok and stays closed.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe session_id returned by session_open. REQUIRED.
workspace_idNoWorkspace_id scope guard (optional); when given it must match the session's workspace.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose real behavioral value: idempotency and the exact repeat-call result ('returns ok and stays closed'). It omits the consequential side effects an agent needs before calling — whether closing invalidates tokens, terminates in-flight messages, or requires workspace ownership — so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the verb and resource lead and the idempotency caveat follows immediately. Nothing could be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the key idempotency trait. But with no annotations at all, it should also address the side effects of closing a session and any authorization or scope requirements, which it does not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with session_id documented as 'returned by session_open' and workspace_id documented as a scope guard, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Close a session'), which is unambiguous on its own. However, it does nothing to distinguish the tool from close cousins in the sibling list such as session_open and session_heartbeat, so an agent must rely on the name alone to route between session lifecycle tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to close a session versus leaving it open, no prerequisites (e.g. must the session be open?), and no mention of the sibling tools that could be alternatives. The only contextual hint is that repeated calls are safe, which is a behavior note rather than usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_heartbeatA

Advance a session heartbeat and report the derived status; keeps you PRESENT (in the broadcast audience) and clear of the stale-session reaper.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe session_id returned by session_open. REQUIRED.
workspace_idNoWorkspace_id scope guard (optional); when given it must match the session's workspace.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose real behavioral traits: membership in the broadcast audience and eviction avoidance via the stale-session reaper. It omits failure behavior (e.g., what happens if the session is already reaped or expired), leaving a gap for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a semicolon-delimited elaboration; every clause earns its place and nothing is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and it correctly references the derived status. The only shortfall is the absence of any guidance on heartbeat cadence or on error/expired-session outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both session_id and workspace_id are already documented in the schema (including the workspace scope-guard rule). The description adds no parameter-level meaning beyond that, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Advance a session heartbeat') and its effect ('report the derived status'), which cleanly separates it from session_open/session_close in the sibling list. It doesn't explicitly name those siblings, so it stops just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'keeps you PRESENT ... and clear of the stale-session reaper' implies when to call it (to stay alive), but there is no explicit when/when-not guidance or cadence/frequency instruction. Usage must be inferred rather than read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_openA

Open a session bound to (agent_id, workspace_id); returns a per-session session_secret (ONLY here - keep it; required by sensitive verbs in strict mode). Heartbeat to receive broadcasts.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesYour agent_id; the session is bound to this identity. REQUIRED.
metadataNoFree-form JSON object stored with the session (optional).
workspace_idNoWorkspace_id (from workspace_resolve) the session operates in (optional; omit for an unbound session).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a genuinely important trait – that a per-session session_secret is emitted ONLY here, must be kept, and is required by sensitive verbs in strict mode – but omits session lifetime, idempotency on re-open, and any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence covers the action, the key return value, and the follow-up step with no filler. Slightly cryptic in phrasing ('ONLY here - keep it') but appropriately front-loaded and compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return-value structure need not be explained, and the description adds the one thing an agent must know that the schema cannot express: the session_secret's singular issuance and downstream requirement. Only session lifetime and error behavior are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents agent_id, workspace_id, and metadata in detail. The description names the (agent_id, workspace_id) binding but adds no format or constraint detail beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Open) and resource (a session), and specifies the binding scope: (agent_id, workspace_id). An agent can distinguish it from session_heartbeat and session_close, though it does not explicitly contrast with them by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trailing clause 'Heartbeat to receive broadcasts' implies a lifecycle (open, then heartbeat) but gives no explicit when-to-use, prerequisites, or exclusions relative to siblings like agent_register or workspace_resolve. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shared_md_renderB

Render the per-workspace human-readable shared.md (atomic overwrite).

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoYour agent_id for permission evaluation in open stdio mode (optional; authenticated HTTP MCP uses the API-key identity).
limit_eventsNoHow many of the most recent events to include in the rendered timeline (optional; default: 50, clamped to the configured maximum).
workspace_idYesThe target workspace_id (from workspace_resolve) whose shared.md is rendered. REQUIRED - NB: this tool takes the workspace_id directly, not project_root.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does disclose one important trait: the write is an atomic overwrite. That is genuinely useful mutation context, but permissions, failure modes, and what happens to prior content beyond 'overwrite' are unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the key constraint (atomic overwrite) in a parenthetical. No filler, nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the schema fully documents parameters. But for a mutation tool with no annotations, the description should say more about permission requirements, overwrite consequences, and when this is the right tool over its siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so workspace_id, agent_id, and limit_events are already fully documented in the schema, including the note to use workspace_id rather than project_root. The description adds no parameter detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (render) and resource (per-workspace shared.md), which is enough to distinguish it from siblings like workspace_resolve or artifact_put. However, it doesn't clarify what rendering produces or how it relates to other workspace tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives such as artifact_put or whichever tool creates shared.md in the first place. The agent must infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tag_listA

List the global operator-managed tag catalog (keys + values) that agent tags, comm_scope and tag targets are validated against fail-closed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the catalog is global, operator-managed, and used in fail-closed validation, but does not state permissions, side effects, rate limits, or pagination behavior. 'List' strongly implies a read-only operation, which is useful but not sufficient for full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The trailing clause 'that agent tags, comm_scope and tag targets are validated against fail-closed' is dense but earns its place by explaining the catalog's role.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter list tool with an output schema, the description covers what the tool returns (keys + values) and why the catalog matters. It does not need to explain return values since an output schema exists; only explicit when-to-use guidance is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to clarify. Per the rubric, zero-parameter tools default to a baseline of 4 when the schema and description are otherwise coherent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (global operator-managed tag catalog keys + values), plus the validation domain (agent tags, comm_scope, tag targets). No sibling tool overlaps, so the purpose is unmistakable without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives implied usage by saying the catalog is what tags, comm_scope, and tag targets are validated against, suggesting an agent should consult it before those operations. However, it does not explicitly state when to use this tool, when not to use it, or what alternative exists (no similar sibling exists).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_listA

GLOBAL-ADMIN: enumerate ALL workspaces. Paths OMITTED by default (include_paths=true is an admin/ops opt-in). For discovery use agent_list / capability_list.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoYour agent_id for permission evaluation in open stdio mode (optional; authenticated HTTP MCP uses the API-key identity).
include_pathsNoAlso return each workspace on-disk root_realpath (default false; paths OMITTED by default - opt-in defense-in-depth). For routine discovery use agent_list / capability_list.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that paths are omitted by default as a defense-in-depth measure, that include_paths is an admin/ops opt-in, and that this is a GLOBAL-ADMIN privileged operation. It stops short of stating rate limits or what the admin scope concretely gates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the scope/privilege constraint and the default behavior, with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers privilege scope, default behavior, and alternatives. Complete enough to call correctly, missing only finer detail on what the admin path reveals.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline is 3; the description adds value by explaining WHY include_paths defaults to false (defense-in-depth, admin opt-in) rather than merely restating the default. It adds minor rationale beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('enumerate ALL workspaces') with an explicit scope qualifier (GLOBAL-ADMIN, ALL). An agent can distinguish this enumeration tool from sibling list/resolve tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes discovery use cases to agent_list / capability_list, giving a clear when-not condition. However it omits guidance on the closest sibling workspace_resolve, so the routing is not fully exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_resolveB

Resolve a project_root to its deterministic workspace_id and upsert it.

ParametersJSON Schema
NameRequiredDescriptionDefault
display_nameNoHuman-friendly label to store/refresh for the workspace (optional).
project_rootYesAbsolute path to the project; the server derives workspace_id = sha256(realpath).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'Upsert it' usefully discloses that resolving also writes/refreshes a workspace record (a side effect not implied by 'resolve' alone), and 'deterministic' hints at idempotency. However it doesn't state what gets created or overwritten, whether display_name refresh is destructive, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the core action and its side effect are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and the schema covers both params. But for a tool that performs an upsert with no annotations, the description should say more about the write side effect (idempotency, what is created/updated) to be complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (project_root and display_name) are already fully documented, including the sha256(realpath) derivation. The description adds nothing beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: resolve a project_root into a deterministic workspace_id, plus an upsert. This is clearly distinct from sibling workspace_list. It doesn't name an alternative explicitly, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as workspace_list. The agent must infer that this is the registration/derivation entry point.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 43 tool updatesv0.1.10
    • First observedagent_get
    • First observedagent_list
    • First observedagent_register
    • First observedagent_whoami
    • First observedartifact_get
    • First observedartifact_put
    • First observedcapability_list
    • First observedchannel_create
    • First observedchannel_list
    • First observedcoordination_health
    • First observedevent_cursor
    • First observedevent_get
    • First observedevent_wait
    • First observedhandoff_cancel
    • First observedhandoff_claim
    • First observedhandoff_complete
    • First observedhandoff_create
    • First observedhandoff_get
    • First observedhandoff_list_available
    • First observedhandoff_reject
    • First observedhandoff_verify
    • First observedinbox_ack
    • First observedinbox_count
    • First observedinbox_extend
    • First observedinbox_history
    • First observedinbox_peek
    • First observedinbox_pull
    • First observedmessage_create
    • First observedmessage_get
    • First observedmessage_list
    • First observedmessage_status
    • First observedmessage_wait
    • First observednexus_info
    • First observedpoll_token_issue
    • First observedpoll_token_renew
    • First observedpoll_token_revoke
    • First observedsession_close
    • First observedsession_heartbeat
    • First observedsession_open
    • First observedshared_md_render
    • First observedtag_list
    • First observedworkspace_list
    • First observedworkspace_resolve

TDQS

A3.5/5.0

Scored across 43 tools

Disambiguation4/5

The tool set covers many distinct sub-domains (handoffs, inbox, events, sessions, etc.), and most tools have clearly differentiated purposes. However, some read-oriented tools overlap (event_get/event_wait, inbox_pull/inbox_peek) and three legacy message_* tools are deprecated but still present, which could momentarily confuse an agent.

Naming Consistency5/5

All 43 tools use consistent snake_case with a domain_action pattern (e.g., handoff_create, inbox_pull, poll_token_issue). The only variation is minor compound forms like handoff_list_available, but the convention is predictable and uniform.

Tool Count2/5

43 tools is well above the typical 3–15 range and even excluding the 3 deprecated message tools leaves 40. This heavy surface likely burdens agents and exceeds what the server’s purpose requires.

Completeness4/5

The surface covers CRUD-like lifecycle for handoffs, inbox, sessions, events, artifacts, agents, workspaces, and poll tokens, with no dead ends in core workflows. Minor gaps exist (no artifact list/delete, no agent/channel delete), but agents can work around them.

Maintenance

ActivityActive
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Persistent memory and handoff intelligence layer for MCP agents. Most memory servers retrieve text — Memory Nexus compounds operational context, learning from usage and progressively synthesizing observations into higher-order intelligence across sessions and tools.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Nexus Memory gives every MCP-compatible agent one persistent, self-hosted shared memory with hybrid retrieval, drift detection, and anti-poisoning features.
    21
    3
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    A real-time inter-agent switchboard, delivered as one centralized streamable-HTTP MCP server. Any MCP-capable agent can message, coordinate, and stay ambiently aware of others.
    1
    AGPL 3.0
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server that wraps the nexus CLI, giving AI agents cross-session memory, semantic search, preference learning, and smart context injection.
    137 npm
    MIT