Skip to main content
Glama

NEXO Brain — Your AI Gets a Brain

npm F1 0.588 on LoCoMo +55% vs GPT-4 GitHub stars License: AGPL-3.0

Local cognitive runtime with a shared brain across Claude Code, Codex, Claude Desktop, and other MCP clients. Persistent memory, durable workflow runs, selectable terminal and automation backends, overnight learning, self-healing background jobs, startup preflight, and doctor diagnostics. 150+ MCP tools. Benchmarked on LoCoMo (F1 0.588, +55% vs GPT-4).

NEXO Brain transforms any MCP-compatible AI agent from a stateless assistant into a cognitive partner that remembers, learns, forgets, adapts, and builds a relationship with you over time.

Watch the overview video · Watch on YouTube · Open the infographic

Version 7.8.0 is the current packaged-runtime line. Minor release that closes the PostCompact continuity work Francisco requested after v7.7: src/hooks/post_compact.py is a real registered hook (part of the canonical 9-hook set, was 8), pre-compact.sh resolves the exact NEXO SID from CLAUDE_SESSION_ID instead of falling back to "latest active session" (that was actively wrong in multi-conversation Desktop), the sidecar moves from /tmp to $NEXO_HOME/runtime/data/compacting-sid.txt so two concurrent compactions on two conversations cannot race on /tmp, post-compact.sh removes its "latest checkpoint" fallback (fail-closed to a diagnostic systemMessage instead of restoring the wrong conversation), and the hook cross-checks the sidecar SID against the env-resolved one so a "SID mismatch" is logged as such. Pre- and post-compact now emit NDJSON events the engine drains on every periodic tick via _consume_pending_hook_events(); the queue file is truncated after read so an event never fires twice. A new contract test (tests/test_v78_compaction_continuity.py) pins 11 invariants across ten rails including the hook registration, the exact-SID resolution path, fail-closed behaviour, and that compaction_count only increments on real restore. Pytest 2086 passing (+16 vs v7.7). No Desktop bump — v0.27.0 continues to ship.

Previously in 7.7.0: minor release that closed the six gaps left partial after v7.6.0's constructor-guardian-90 pass 1 (autonomous detector for multi_step_task_detected, R16 vocabulary expansion, R_CATALOG extended to plain Edit/Write, new R_PRIMITIVE_CHOICE rule, R11_plugin_load_pre_inventory hardened, 12 new contract tests). Post-review hotfix on the same release wired task_open rearm properly (discarded from tools_called + per-instance pin cleared on task_close), added live on_event triggers in R14 and R16, and called on_tool_call_before before on_tool_call in run_with_enforcement so before_tool rules fire in Brain the same way Desktop fires onBeforeToolCall.

Previously in 7.6.0: minor release that closed the drift between tool-enforcement-map.json v2.2 and the two enforcement engines (Brain Python + Desktop JS), added per-instance after_tool satisfaction, tightened learning_add grace to 0 and task_open threshold to 4/must, hardened R15/R17/R22/R_CATALOG from soft to hard, and raised R34 from shadow to soft.

Previously in 7.5.0: minor release that promoted nexo_lifecycle_event from ledger + reconciliation authority to canonical authority of session-end. Brain now owns the prompt, the sequence, and the timing of diary+stop; Desktop v0.25.0 (closed-source companion) is the conduit that executes Brain's plan against the live Claude process. The new 2-call contract — nexo_lifecycle_event returns a versioned canonical_plan (resume_session → inject_prompt → stop_session, with stable ids and per-action timeouts) and nexo_lifecycle_complete_canonical confirms execution with a per-action results array — replaces polling with explicit acknowledgement. canonical_plan_id is deterministic: sha256(event_id + "|v" + plan_version)[:24], so retries reuse the same id. Migration m52 extends lifecycle_events with six canonical_* columns plus an index; pre-v7.5 rows simply carry NULL. session_diary is the dedupe key on re-delivery: if Desktop crashes between executing the inject and sending the complete call, the next nexo_lifecycle_event for the same event_id checks for a diary written after canonical_dispatched_at; if one exists, Brain short-circuits to already_processed and refuses to re-dispatch. The seven explicit delivery_status values (accepted, processed, canonical_pending, canonical_done, already_processed, retryable_error, rejected) give the pipeline a diffable state machine. switch and window-close stay observational (no plan ever issued, even with a live session_id). nexo lifecycle record now returns exit code 0 for canonical_pending; older wrappers that treated it as an error are incompatible with v7.5. MCP tool count: 262 → 263.

Previously in 7.4.1: patch release correcting the over-promise in v7.4.0's release notes and locking in the exact role of nexo_lifecycle_event as a ledger + reconciliation authority — NOT the canonical executor of diary+stop, which lived in Desktop. That responsibility moved to Brain in v7.5.

Previously in 7.2.0: minor release consolidating three parallel workstreams into a single Guardian-active-by-default train. Block K roadmap closure (G1 enforcer active, G3 SSH remote-write detector, src/guardian_runtime_config.py resolver, _persist_guardian_hard_defaults during nexo update). F0.6 hardening wave (nexo rollback f06 CLI, src/scripts/prune_runtime_backups.py promoted to core, docs/f06-layout-contract.md, three new doctor boot-tier checks, scripts/nexo-migrate-nora.sh + scripts/f0-safe-apply-remote.sh idempotent migration). Adaptive weights flipped from "14-day calendar wait" to "14 days OR (≥200 samples AND ≥2 days)" with auto-promotion during nexo update. Small-fixes batch: R34 bool("unknown")==True fix, classify_scripts_dir dedup, B10 module-level path constants lazy-evaluated, schedule override audit log, scripts/pre-release-verify.sh + docs/release-discipline.md, pre-commit hook that blocks commits when tool-enforcement-map.json drifts from src/plugins/.

Previously in 7.1.10: follow-up over v7.1.8 that shipped two rescue batches of WIP stashed aside during the v7.1.8 release window. First rescue: src/autonomy_mandate.py expanded the mandate-detection vocabulary (hazlo todo / no pares / estás al mando / te dejo al mando / sigue sin parar / haz el plan completo), added three honest flags on MandateState (execute_until_blocker, suppress_mid_task_menus, revalidate_after_compaction) with session filtering, wired post/pre-compact hooks that read those flags, surfaced them through protocol/workflow handlers and session payload, and introduced the new src/checkpoint_policy.py module with tests. Second rescue: scripts/verify_release_readiness.py gained a smoke-artifact contract pass that validates release-contracts/smoke/v<version>.json before any tag push, the release-final audit skill references the new contract, src/hook_guardrails.py + src/hooks/post_tool_use.py refine the post-tool protocol reminder path with a new contract test, and a couple of core prompts (task-close evidence, r14 correction learning) got wording polish.

Previously in 7.1.8: batch release over v7.1.7 consolidating the Block K Guardian/Enforcer roadmap (auto-drain of stale protocol_debt rows, destructive-command pre-tool gate, guard_check-required gate, inline guard ack on nexo_task_open, Guardian Health in the morning briefing) with Block D hardcode cleanup (classifier-backed backfill_task_owner, migration v50 supersedes the duplicate NEXO-product learning pair, new semantic-hardcodes audit) and Block E product guards (LaunchAgent plist protection, agent-name fallbacks no longer leak the product identity, francisco_emails removed from the email-config dict export, runner-health-check.py + nexo_personal_automation.py promoted from personal to core).

Previously in 7.0.1: hotfix over v7.0.0 (db._core.DB_PATH was only caller still hardcoded to legacy ~/.nexo/data/nexo.db; every shared-DB command silently returned empty results post-migration). Previously in 7.0.0: BREAKING — Plan Consolidado fase F0.6: physical separation of the runtime tree into ~/.nexo/{core,personal,runtime}/. The flat layout (~/.nexo/scripts/, brain/, data/, operations/, ...) is gone. Operators on v6.x are auto-migrated on first nexo update; fresh installs land directly in the new tree. New paths.py helpers are transition-aware.

Previously in 6.5.0: Plan Consolidado fase F0.2: operators can now nexo scripts enable|disable|status <name> any personal automation. The cron wrapper honours the flag at every tick (exit 0 with summary='[disabled]' while the LaunchAgent stays loaded). The companion NEXO Desktop client (a closed-source product, distributed separately) wires the same toggle into its Automatizaciones panel. See CHANGELOG for the full diff.

About NEXO Desktop. NEXO Desktop is a separate closed-source companion app distributed at nexo-desktop.com — its source does not live in this repo. When release notes mention Desktop they describe a coordinated client release that consumes the Brain's CLI / MCP contract; the Brain itself is fully usable on its own (terminal, Codex, Claude Code, or any MCP client). If you want the product edition rather than the open-source Brain alone, contact info@wazion.com and ask about NEXO Desktop.

Previously in 6.4.0: Plan Consolidado fase F1 — multi-tenant email accounts (email_accounts table, nexo email setup interactive wizard, nexo email add --password-stdin --json for machine consumers, idempotent migrator from legacy ~/.nexo/nexo-email/config.json). On post-F0.6 installs that legacy-looking path is only a compatibility alias/shim into ~/.nexo/runtime/nexo-email/config.json; it should never be treated as a second source of truth.

Previously in 6.3.1: privacy hotfix over v6.3.0. The nightly auditor caught that src/presets/entities_universal.json in v6.3.0 shipped operator-specific vhost_mapping entries (private IPs, hostnames, tenant names). v6.3.1 pulls those out into src/presets/entities_local.sample.json (template) + .gitignore'd ~/.nexo/brain/presets/entities_local.json (operator copy), and the installer drops the sample at nexo init. No behaviour change on the Guardian side.

Previously in 6.3.0 — Plan Consolidado wave 2, coordinated with NEXO Desktop v0.18.0. Closes the remaining Guardian roadmap items that do not require an invasive structure migration: extended cognitive_sentiment shape (is_correction/valence/intent), extended entities schema, 21 labelled rule fixtures with R13 spike gates, Fase F telemetry loops + Deep Sleep phase, pinned local zero-shot classifier skeleton (mDeBERTa), hook respects NEXO_MIGRATING=1, origin column on personal_scripts, and the T4 LLM gate wrapping R15/R23e/R23f/R23h (byte-parity Py ↔ JS). Two pre-release auditors flagged a CRITICAL in the first JS wire (method-name + async mismatch) and a HIGH (classifier bool conflated "no" with "unparseable"); both corrected with regression tests before merge.

Previously in 6.1.1: small fix to nexo --help so the Latest: vX line reliably appears when NEXO Desktop invokes the CLI via subprocess — unblocks the Desktop Brain auto-update banner that previously couldn't parse the version delta. No behaviour change for interactive terminal users; the 6-hour registry cache still rate-limits network calls. Bundles all v6.1.0 Protocol Enforcer Fase 2 + multi-claude-sid hotfix content.

Previously in 6.0.2: adds the reserved caller prefix personal/* so scripts living in ~/.nexo/scripts/ can invoke the automation backend with their own caller id without editing src/resonance_map.py. New kwarg tier ("maximo" / "alto" / "medio" / "bajo") on run_automation_prompt, run_automation_interactive, nexo_helper.run_automation_text, nexo_helper.run_automation_json, and nexo-agent-run.py --tier. Precedence for personal/* callers: explicit tier= → explicit reasoning_effort= → calibration.preferences.default_resonance → DEFAULT_RESONANCE (alto). Registered callers keep their behaviour unchanged. New guide: docs/personal-scripts-guide.md.

Previously in 6.0.1: hotfix on top of the 6.0.0 release. protocol_settings.py now treats the process as interactive when either stdin+stdout are TTYs or NEXO_INTERACTIVE=1 is exported — closes the gap where NEXO Desktop 0.12.0 spawned claude through pipes and Brain fell back to lenient even with a human in the loop. The PostToolUse hook also gains an inbox autodetect stage: when the session has unread nexo_send messages and has gone 60s+ without a heartbeat, it emits a systemMessage asking the agent to run nexo_heartbeat and consume them. Rate-limited to one reminder per minute per SID (new hook_inbox_reminders table, migration m42). Added sessions.last_heartbeat_ts, stamped by every successful heartbeat. NEXO_INTERACTIVE is an internal Brain↔Electron contract — not user-facing, not a resurrection of the removed NEXO_PROTOCOL_STRICTNESS.

Previously in 6.0.0: BREAKING tier-only setup. Onboarding asks for one resonance tier (maximo/alto/medio/bajo) and that choice drives every backend via src/resonance_tiers.json; the per-backend model/effort prompts are gone and the legacy client_runtime_profiles.{claude_code,codex}.{model,reasoning_effort} are silently purged from schedule.json on upgrade. Protocol strictness is no longer configurable — interactive TTY sessions run strict, non-TTY (crons, pipes, tests) run lenient; NEXO_PROTOCOL_STRICTNESS env, preferences.protocol_strictness, and the default/normal/off/warn/soft aliases are all removed. preferences.show_pending_at_start moves to NEXO Desktop's electron-store. The seven core hooks are now unified behind src/hooks/manifest.json (plugin and npm modes read the same file), two new hooks ship (Notification for live-session activity and SubagentStop for auto-closing stale protocol_tasks), and auto_capture.py is wired to both UserPromptSubmit and PostToolUse with a persistent 1h dedup table plus an automatic nexo_learning_add on correction matches. ~/.nexo/hooks_status.json is published after every registerAllCoreHooks() so NEXO Desktop ≥0.12.0 can render Hooks activos X/Y. New nexo-brain --skip flag aliases --yes/--defaults. Full suite 1057 passed, 1 skipped.

Previously in 5.10.2: auto-bootstraps brain/profile.json from brain/calibration.json on nexo update when the profile file is missing, empty, or corrupt AND calibration carries at least one of meta.role, meta.technical_level, name, language. NEXO Desktop's Preferencias → Avanzado tab used to render an empty {} for that block when the onboarding flow had been interrupted; now it either shows the seeded profile or a friendly explanation of what each file is for, paired with Desktop v0.11.2 which adds header descriptions to both JSON blocks. Never overwrites a populated profile, never raises, idempotent. Also fixes a latent host-filesystem leak in test_user_facing_caller_with_no_user_default_uses_alto exposed by the v5.10.1 migration.

Previously in 5.10.1: silent, one-shot migration that recovers legacy reasoning_effort="max" (written by nexo preferences --reasoning-effort max before v5.9.0) into the new preferences.default_resonance map — any user who had configured max before v5.9.0 and never touched the new selector was silently falling back to DEFAULT_RESONANCE="alto" on interactive calls since the v5.10.0 update. _run_runtime_post_sync() runs _migrate_effort_to_resonance() exactly once: max→maximo, xhigh→alto, high→medio, medium→bajo. No-op when calibration or schedule already declares an explicit default_resonance; idempotent; conservative; never raises.

Previously in 5.10.0: fixes the deep-sleep extract bloat that made Session 1 take ~57 minutes on some installs (new bare_mode on run_automation_prompt wires claude --bare for JSON-only extractor callers — ~4.3× faster per child, sourced from ANTHROPIC_API_KEY env or ~/.claude/anthropic-api-key.txt). caller= is now mandatory on run_automation_prompt — no silent fallback; every automation subprocess traces back to a registered caller with a deliberate tier. Five personal scripts (personal/email-monitor, personal/github-monitor, personal/post-x, personal/followup-runner, personal/orchestrator-v2) joined the resonance map with tiers picked per caller based on what each one does. gbp/* marketing posts bumped from medio to alto (public-facing copy, quality first over speed). 65 legacy protocol debts bulk-resolved as part of the audit — the patterns that generated them are structurally closed by mandatory caller= + unified session log + bare_mode.

Previously in 5.9.1: adds default_resonance to brain/calibration.json via the Desktop-facing schema (nexo schema --json), so NEXO Desktop's Preferences dialog renders a select with Máximo / Alto (recomendado) / Medio / Bajo automatically — no Desktop release needed. resolve_tier_for_caller reads calibration first and falls back to the legacy schedule.json location. nexo preferences --resonance writes both. The UI control only affects interactive sessions (nexo chat, Desktop new conversation, interactive nexo update); crons and background processes stay pinned per caller in resonance_map.py.

Previously in 5.9.0: every Claude/Codex invocation now flows through a central resonance map and a unified session log. Four tiers (MAXIMO / ALTO / MEDIO / BAJO) each resolve to a concrete (model, reasoning_effort) pair per backend. User-facing callers (nexo chat, Desktop new conversation, interactive nexo update) honour the user's default_resonance preference; system-owned callers (deep-sleep, evolution, catchup, GBP posts, …) run at a fixed tier chosen per caller in src/resonance_map.py — the user's preference never downgrades a cron we decided needs MAXIMO. Unknown callers raise UnregisteredCallerError. Migration #41 adds caller, session_type, started_at, ended_at, pid, resonance_tier to automation_runs; interactive sessions record a row at spawn (with ended_at=NULL) and update it on close, so the Brain now has a single source of truth for every Claude/Codex call regardless of origin. New nexo preferences --resonance CLI. New MCP tools nexo_session_log_create / nexo_session_log_close let NEXO Desktop (which spawns claude directly from its TypeScript process) feed the same log.

Previously in 5.8.2: the Brain core no longer auto-classifies followups and reminders on behalf of agents. v5.8.0's classify_task() heuristic (NEXO-specific ID prefixes NF-PROTOCOL-* / NF-DS-* / NF-AUDIT-*, Spanish user-verbs debes / revisar / firmar, agent keywords monitor / auditoría diaria / checkpoint) was fine for NEXO's own DB but bled convention into every third-party agent plugged into the shared Brain. The core now persists internal=0 and owner=NULL when the caller omits them, and clients that want automatic classification (NEXO Desktop does, via its _legacyClassifyOwner helpers) compute it themselves and pass the result. Migration #40 keeps the columns + indexes; rows already backfilled by v5.8.0 keep their values. normalise_owner still explicitly rejects the string "nexo" so legacy hardcoding cannot sneak back in.

Previously in 5.8.1: closes a self-reinforcing launchctl kickstart -k loop in the watchdog that wedged deep-sleep Phase 2 between 2026-04-14 and 2026-04-17. The cron wrapper now INSERTs an in-flight row (ended_at=NULL) at start and traps SIGTERM/INT/HUP to close it with exit_code=143 instead of vanishing from cron_runs. The watchdog interprets in-flight rows as "currently running" and only re-executes after verifying the worker process is dead. extract.py classifies CLI failures into transient (overloaded_error, rate-limit, timeout, signal — retried next run) and deterministic (skipped after MAX_POISON_ATTEMPTS), and passes a slim shared-context (200 head lines + metadata) instead of the full 400+ KB dump. A new auto_update._heal_deep_sleep_runtime() repairs existing installs silently on the next nexo update: poisoned checkpoints, stale locks, dangling cron_runs rows, and bloated .watchdog-fails counters.

Previously in 5.8.0: first-class internal and owner columns on followups and reminders. Migration #40 adds both fields with an idempotent one-shot backfill, so the "who does this task belong to?" classification moves from client-side regex (Desktop) to persistent storage every MCP client shares. Taxonomy is intentionally generic — owner in {user, waiting, agent, shared} — so third-party agents plugging into the shared Brain can render whatever assistant label they carry without inheriting NEXO branding. nexo_reminder_create, nexo_reminder_update, nexo_followup_create, and nexo_followup_update gain optional internal and owner parameters that win over the default heuristic.

Previously in 5.7.0: nexo update now keeps Claude Code and Codex CLIs in lockstep with NEXO Brain itself. When the global @anthropic-ai/claude-code or @openai/codex packages are installed, the updater checks the npm registry and runs npm install -g <pkg>@latest in-line — so the terminal boot model stays aligned with the settings NEXO already wrote to ~/.claude/settings.json. Packages the operator never installed are skipped silently. Pass nexo update --no-clis to keep the terminal CLIs pinned.

Previously in 5.6.1: update-path hardening — 0-byte .db orphans from interrupted installs are now purged from ~/.nexo/ and ~/.nexo/data/ before the pre-update backup, and sync_claude_code_model() propagates the NEXO-recommended model into ~/.claude/settings.json whenever heal_runtime_profiles() migrates the claude_code default.

Previously in 5.5.5: data-loss guardrails + automatic self-heal. The updater now refuses to capture an already-wiped nexo.db into a pre-update-* snapshot (validated sqlite3.backup + pre-flight wipe guard + post-migration row-count gate), and an auto-heal restores data/nexo.db from the newest hourly backup on the next server boot when a wipe is detected. New nexo recover CLI + nexo_recover MCP tool.

Previously in 5.5.4: Deep Sleep no longer blocks on unparseable sessions — reduced retries, added a JSON escape hatch, and unified the automation subprocess timeout to 3h across all scripts via a single shared constant.

Previously in 5.5.3: CLAUDE.md CORE teaches the model to trust the Protocol Enforcer, so aligned backends stop rejecting heartbeat, diary, and checkpoint injections as suspected prompt injection.

Start here:

Every time you close a session, everything is lost. Your agent doesn't remember yesterday's decisions, repeats the same mistakes, and starts from zero. NEXO Brain fixes this with a cognitive architecture modeled after how human memory actually works.

Shared Brain Across Clients

Shared brain is now the baseline:

  • Claude Code remains the recommended path because it still has the deepest hook integration and the most battle-tested headless automation surface.

  • Codex is supported both as an interactive terminal client and as the background automation backend.

  • Claude Desktop can point at the same local brain through MCP.

That means NEXO now manages not only the shared runtime and MCP wiring, but also the startup layer around it:

  • nexo chat opens the configured client instead of assuming Claude Code forever.

  • Claude Code and Codex both get managed bootstrap files:

    • ~/.claude/CLAUDE.md

    • ~/.codex/AGENTS.md

  • Those files now use an explicit CORE / USER contract, so NEXO can update product rules in CORE while preserving operator-specific instructions in USER.

  • For Codex specifically, nexo chat and Codex headless automation inject the current bootstrap explicitly, so Codex starts as NEXO even when plain global Codex startup is inconsistent about global instructions.

  • Deep Sleep now reads both Claude Code and Codex transcript stores, so overnight analysis still works even when the user spends the day in Codex.

Versions 2.6.14 through 2.7.0 established the practical shared-brain baseline: managed Claude/Codex bootstrap, Codex config sync, transcript-aware Deep Sleep, 60-day long-horizon analysis, weekly/monthly summary artifacts, retrieval auto-mode, and the first measured engineering loop.

Versions 3.0.0 and 3.0.1 close the next execution gap:

  • protocol discipline is now a runtime contract, not just instructions:

    • nexo_task_open

    • nexo_task_close

    • persistent protocol_debt

    • enforceable Cortex gates

  • durable execution is now first-class:

    • resumable workflow runs

    • checkpoints

    • replay

    • retries

    • durable goals

  • conditioned learnings on critical files are now real guardrails across Claude hooks, Codex transcript audits, and headless automation prompts

  • repair/correction work now routes through canonical learning capture instead of depending on the model to remember to document after the fact

  • runtime truth is stricter:

    • no more healthy-looking warning storms

    • no more silent Deep Sleep schema drift

    • keep-alive jobs report alive/degraded/duplicated honestly

  • public proof is stronger:

    • measured compare scorecard

    • external and internal ablations

    • cost_per_solved_task

    • SDK/API/quickstart surface

Versions 3.1.7 through 3.2.0 close the recent-memory gap:

  • recent operational continuity is now first-class through hot context and recent events

  • the runtime can build a reusable pre-action bundle instead of reconstructing the last few hours from diaries and durable recall only

  • when even that misses, NEXO now exposes raw transcript fallback tools for Claude Code and Codex session stores

  • NEXO can now inspect itself through a live system catalog derived from canonical sources instead of relying only on stale docs or operator memory

Version 5.3.11 hardens protocol and Cortex contracts: malformed outcome, task_type, and impact_level values now fail explicitly instead of being coerced into other valid states, so persisted task history, debt, hot context, and decision telemetry stay faithful to what the caller actually asked for. Version 5.3.10 tightened the packaged-runtime truth layer again: installs and updates now keep ~/.nexo/package.json aligned with the published npm package so runtime metadata and doctor evidence no longer drift to an old version, nexo doctor --tier deep treats a missing self-audit-summary.json as a pending bootstrap artifact when the runtime was just installed or updated instead of reporting a false degradation, weekly Evolution now asks for explicit dimension_scores / score_evidence so telemetry can persist instead of staying blank, and daily synthesis only ingests update-last-summary.json when it carries actionable runtime signals. Version 5.3.9 is the packaged core-artifact manifest heal for 5.3.8: packaged updates now rebuild runtime-core-artifacts.json from the canonical npm package src/ tree instead of scanning the live ~/.nexo/scripts directory, script classification prefers that canonical packaged source when available, and runtime doctor syncs personal scripts before LaunchAgent inventory so personal automations recover cleanly instead of being mistaken for unknown core drift. Version 5.3.8 was the immediate packaged-migration hotfix for 5.3.7: the installer/runtime migrator now discovers all top-level runtime Python modules from src/ dynamically instead of relying on a manual allowlist, so new product surfaces like nexo export / nexo import actually arrive in ~/.nexo after update instead of being present only in the published npm tarball. Version 5.3.7 closed the remaining packaged-runtime happy-path gap and finally exposed portable user-data migration commands: packaged nexo update now self-heals cron definitions and LaunchAgents after a successful npm bump, new nexo export / nexo import commands move operator data as a safe bundle instead of leaving that flow implicit, and runtime doctor now distinguishes tracked historical Codex drift from an actually broken runtime so cleaned installs stop staying red for stale transcript debt alone. Version 5.3.6 hardened the Claude Code bootstrap path and related runtime hygiene: managed client sync now writes the NEXO MCP server where current Claude Code actually reads it (~/.claude.json), script classification is stricter about core-vs-personal runtime artifacts, schedule status distinguishes genuinely running jobs from broken ones, and retroactive learnings stop opening keyword-only false positives outside their declared applies_to scope. Version 5.3.5 already keeps CLI version visibility honest right after nexo update: if the cached npm version lags behind the runtime you just installed, nexo / nexo chat now clamp Latest to the installed version and refresh the cache instead of showing a stale older release. Version 5.3.4 already cleaned up legacy core alias leakage and added the version-status banner. Version 5.3.3 closed the remaining packaged-runtime doctor mismatch: the built-in hourly backup helper is now inventoried as a core LaunchAgent, so clean installs no longer get a false unknown-LaunchAgent warning. Version 5.3.2 already hardened the runtime boundary by persisting which runtime scripts/hooks are core product artifacts, keeping nexo scripts from mixing those into the personal bucket, and migrating the legacy Claude Code heartbeat wrappers into managed core hooks.

Version 5.3.1 normalizes packaged npm installs so they behave like packaged npm installs: nexo update now keeps the runtime anchored to ~/.nexo, refreshes packaged bootstrap/client artifacts after upgrade, avoids repo-only release-artifact drift in installed runtimes, and keeps personal scripts on the canonical packaged path.

Version 5.3.0 adds nexo uninstall — a CLI command that cleanly separates runtime from user data. It stops all crons, removes the MCP server config, and preserves databases, learnings, and personal scripts for safe reinstall.

Version 5.2.1 fixes the Deep Sleep datetime regression and closes the decision-to-outcome feedback gap:

  • _parse_any_datetime in apply_findings.py now strips timezone info before comparison, fixing the offset-aware/offset-naive crash that was breaking Deep Sleep verification work.

  • cortex_decide() now auto-creates a decision_outcome when none is linked yet, so the outcome-checker cron can verify real decisions instead of leaving the loop open.

Version 5.2.0 closes two focused gaps in the Cortex layer that were left open by the v5.1 audit — the high-stakes response-contract detector was English-only, and the nexo-cortex-cycle cron was writing a quality snapshot that no reader ever consumed:

  • HIGH_STAKES_KEYWORDS_ES adds ~45 Spanish keywords to the high-stakes detector with accented and unaccented variants, so a goal written in Spanish (migrar la base de datos de producción) trips the same gate as its English twin.

  • NEGATION_PATTERNS suppresses false positives when the user explicitly disclaims touching the sensitive area (sin afectar producción, no tocar prod, without touching production, don't modify). The raw keyword being present is no longer enough to flag the task.

  • evaluate_response_confidence accepts two new optional kwargs, pre_action_context_hits (+up to 10) and area_has_atlas_entry (+5), so the score can finally reward tasks that loaded real context instead of only punishing unprepared ones. Both signals are capped and cannot override a real risk penalty.

  • A monotonic numeric safeguard layers on top of the boolean decision tree: answer downgrades to verify when final_score < 50, and verify downgrades to defer when high_stakes and final_score < 30. The safeguard can only make response discipline stricter, never looser.

  • handle_cortex_quality in src/plugins/cortex.py now reads $NEXO_HOME/operations/cortex-quality-latest.json when the requested window (7 or 1 days) is fresh (<6h 30m) and the schema matches — silent fallback to the live SQL computation on any failure. The handler's JSON response now includes "source": "cache" | "live" for observability.

Version 5.1.0 lands the full NEXO-AUDIT-2026-04-11 roadmap as a single minor bump — every open evolution / adaptive / cognitive / skills loop now closes under itself, the knowledge graph exports cleanly, OpenTelemetry spans can be turned on without a hard dependency, and every PR has to clear lint, security, coverage, and release-readiness gates before it can merge:

  • Evolution cycle now auto-applies user-approved proposals on the next run (backed by the new idempotent migration m38), adaptive learned-weight rollbacks surface as visible followups, outcome patterns auto-promote to draft skills, and a Voyager-style detector exposes co-occurring skill pairs as composite-skill candidates via nexo_skill_compose_candidates.

  • cognitive._search.search() now accepts dream_weight and reranks dream-insights through it, somatic markers fold into the same reranking path (max +0.10 boost), state watchers open and auto-resolve deterministic NF-WATCHER-{id} followups, and correction fatigue opens a visible followup instead of only decaying memory.

  • A new Cortex quality cron (every 6h) watches accept rate / linked-success / override gap and opens NF-CORTEX-QUALITY-DROP idempotently when the decision engine starts drifting between cycles.

  • Adding a new learning now walks recent decisions through retroactive_learnings.apply_learning_retroactively() and opens deterministic NF-RETRO-L<id>-D<id> followups for every decision the learning would have changed (exposed via nexo_learning_apply_retroactively).

  • Hook lifecycle observability: new hook_runs table (migration m39) + nexo_hook_runs tool expose recent hook runs, failure streaks, and a health summary. Hook drops are no longer invisible.

  • Knowledge graph bitemporal export: nexo_kg_export emits JSON-LD (with an nexo:* vocabulary) or GraphML, and accepts an as_of ISO timestamp that replays the historical snapshot through kg_edges.valid_from / valid_until for igraph, Gephi, NetworkX, and Cytoscape.

  • OpenTelemetry integration: new src/observability.py soft-imports opentelemetry and only activates when OTEL_EXPORTER_OTLP_ENDPOINT or OTEL_SERVICE_NAME is set. tool_span() becomes a real span when enabled and stays a no-op context manager when disabled.

  • CI gates on every PR: new workflows enforce ruff (E9 / F63 / F7 / F82 / F821), bandit at high severity / high confidence, coverage baselines, and verify_release_readiness.py --ci. A PR that breaks the release contract fails loudly instead of waiting until tag push.

  • Safer update path: auto_update is guarded by a POSIX flock with stale-steal at 10 minutes, and on macOS it now launchctl unloads and reloads every com.nexo.*.plist after a version bump so long-lived crons pick up the new codebase immediately.

Version 5.0.4 tightens the local runtime bridge and trims false-positive doctor noise:

  • vendorable nexo_helper.py now resolves NEXO_HOME and the nexo CLI path robustly, so personal scripts and subprocess flows stop depending on a lucky PATH

  • doctor no longer degrades because of advisory-only self-audit warnings or a single missing usage-telemetry row

  • managed Claude Code and Codex bootstraps now force an immediate first answer after simple email/diary/reminder/followup reads instead of feeling hung while chaining extra lookups

Version 5.0.3 closes the next post-5.0 runtime gap:

  • nexo chat now boots Claude Code and Codex with an explicit NEXO startup prompt instead of opening cold or leaking the target path as a fake prompt

  • terminal launches now use the requested working directory as real cwd, so the selected project path stops behaving like chat text

  • the vendorable nexo_helper.py bridge now bounds helper calls with a timeout instead of letting personal-script subprocess flows wait forever

  • the doctor hardening from 5.0.2 remains validated on a real upgraded runtime after sync

Version 5.0.2 closes the small post-5.0.1 doctor drift:

  • deep doctor now reads the live learnings schema correctly whether the install uses status or the older archived flag

  • a real upgraded runtime was revalidated with nexo update, nexo doctor --tier deep, nexo doctor --tier all, and a fresh Claude Code startup smoke

Version 5.0.1 hardens the live 5.0 upgrade path:

  • managed Claude Code hooks are now cleaned up when an older release left obsolete core-managed entries behind

  • upgrades no longer preserve the stale heartbeat-guard.sh path that could create warning storms and fake "hung" symptoms after nexo update

  • the corrected path has been revalidated on a real install with nexo clients sync, Codex/Claude Code headless runtime access, email-monitor recovery, and a full nexo update

Version 5.0.0 closes the loop between memory, decisions, outcomes, and reusable behavior:

  • goal profiles are now explicit and auditable instead of living as hidden heuristics

  • the Cortex can rank alternatives with goals, outcomes, overrides, and structured penalties

  • repeated outcome patterns can become durable learnings that influence later decisions

  • outcome-backed evidence can seed, promote, demote, or retire reusable skills

  • the runtime benchmark pack now shows the operator/runtime advantage with checked-in artifacts instead of relying only on prose

  • personal-script/core runtime paths, protocol debt maintenance, and release doctoring are now strong enough that the live install path can be audited honestly before release

Client Capability Matrix

Capability

Claude Code

Codex

Claude Desktop

Shared brain / MCP runtime

Yes

Yes

Yes

Managed bootstrap document

~/.claude/CLAUDE.md

~/.codex/AGENTS.md

Not applicable

Global startup bootstrap sync

Native via hooks + bootstrap

Managed via bootstrap + Codex config initial_messages + mcp_servers.nexo

Managed MCP-only shared-brain metadata

nexo chat terminal client

Yes

Yes

No

Background automation backend

Recommended

Supported

No

Raw transcript source for Deep Sleep

Yes

Yes

No

Native hook depth

Deepest

Partial, compensated

None

Runtime doctor parity audit

Yes

Yes

Shared-brain only

Recommended today

Yes

Supported

Shared-brain companion

Supported Clients

Client

Status

Integration style

Notes

Claude Code

First-class

Managed install + hooks + bootstrap

Deepest NEXO parity today

Codex

First-class

Managed install + bootstrap + transcript parity

Best non-Claude terminal path

Claude Desktop

Companion

MCP-only shared brain

Useful as read/chat companion

Cursor

Documented companion

MCP + .cursor/rules

Good editor pairing; no Deep Sleep transcript parity yet

Windsurf

Documented companion

MCP + .windsurf/rules or repo AGENTS.md

Native MCP support, manual companion mode

Gemini CLI

Adapter included

MCP + GEMINI.md

Best when you want Gemini as a shared-brain companion, not the primary NEXO runtime

Related MCP server: mcp-memory

The Problem

AI coding agents are powerful but amnesic:

  • No memory — closes a session, forgets everything

  • Repeats mistakes — makes the same error you corrected yesterday

  • No context — can't connect today's work with last week's decisions

  • Reactive — waits for instructions instead of anticipating needs

  • No learning — doesn't improve from experience

  • No safety — stores anything it's told, including poisoned or redundant data

The Solution: A Cognitive Architecture

NEXO Brain implements the Atkinson-Shiffrin memory model from cognitive psychology (1968) — the same model that explains how human memory works:

What you say and do
    |
    +---> Sensory Register (raw capture, 48h)
    |       |
    |       +---> Attention filter: "Is this worth remembering?"
    |               |
    |               v
    +---> Short-Term Memory (7-day half-life)
    |       |
    |       +---> Used often? --> Consolidate to Long-Term Memory
    |       +---> Not accessed? --> Gradually forgotten
    |
    +---> Long-Term Memory (60-day half-life)
            |
            +---> Active: instantly searchable by meaning
            +---> Dormant: faded but recoverable ("oh right, I remember now!")
            +---> Near-duplicates auto-merged to prevent clutter

This isn't a metaphor. NEXO Brain literally implements Ebbinghaus forgetting curves, rehearsal-based reinforcement, and memory consolidation during automated "sleep" processes.

What Makes NEXO Brain Different

Without NEXO Brain

With NEXO Brain

Memory gone after each session

Persistent across sessions with natural decay and reinforcement

Repeats the same mistakes

Checks "have I made this mistake before?" before every action

Keyword search only

Finds memories by meaning, not just words

Starts cold every time

Resumes from the mental state of the last session

Same behavior regardless of context

Adapts tone and approach based on your mood

No relationship

Trust score that evolves — makes fewer redundant checks as alignment grows

Stores everything blindly

Prediction error gating rejects redundant information at write time

Vulnerable to memory poisoning

4-layer security pipeline scans every memory before storage

No proactive behavior

Context-triggered reminders fire when topics match, not just by date

How the Brain Works

Memory That Forgets (And That's a Feature)

NEXO Brain uses Ebbinghaus forgetting curves — memories naturally fade over time unless reinforced by use. This isn't a bug, it's how useful memory works:

  • A lesson learned yesterday is strong. If you never encounter it again, it fades — because it probably wasn't important.

  • A lesson accessed 5 times in 2 weeks gets promoted to long-term memory — because repeated use proves it matters.

  • A dormant memory can be reactivated if something similar comes up — the "oh wait, I remember this" moment.

On top of that baseline, NEXO now keeps a lightweight per-memory profile:

  • stability slows decay for memories that keep surviving retrieval and reinforcement

  • difficulty speeds decay slightly for memories that tend to be weak, noisy, or harder to reuse correctly

That keeps the core Ebbinghaus model, but makes decay more individual and less purely global.

Semantic Search (Finding by Meaning)

NEXO Brain doesn't search by keywords. It searches by meaning using vector embeddings (fastembed, 768 dimensions).

Example: If you search for "deploy problems", NEXO Brain will find a memory about "SSH connection timeout on production server" — even though they share zero words. This is how human associative memory works.

Retrieval is now also smarter by default:

  • HyDE auto mode expands conceptual or ambiguous queries when that improves recall

  • Spreading activation auto mode adds a shallow associative boost for concept-heavy searches

  • Exact lookup heuristics keep both off for literal file paths, IDs, stack traces, and other precision-sensitive queries

Metacognition (Thinking About Thinking)

Before every code change, NEXO Brain asks itself: "Have I made a mistake like this before?"

It searches its memory for related errors, warnings, and lessons learned. If it finds something relevant, it surfaces the warning BEFORE acting — not after you've already broken production.

Cognitive Dissonance

When you give an instruction that contradicts established knowledge, NEXO Brain doesn't silently obey or silently resist. It verbalizes the conflict:

"My memory says you prefer Tailwind over plain CSS, but you're asking me to write inline styles. Is this a permanent change or a one-time exception?"

You decide: paradigm shift (permanent change), exception (one-time), or override (old memory was wrong).

Sibling Memories

Some memories look identical but apply to different contexts. "How to deploy" for Project A is different from Project B. NEXO Brain detects discriminating entities (different OS, platform, language) and links them as siblings instead of merging them:

"Applying the Linux deploy procedure. Note: there's a sibling for macOS that uses a different port."

Trust Score (0-100)

NEXO Brain tracks alignment with you through a trust score:

  • You say thanks --> score goes up --> reduces redundant verification checks

  • Makes a mistake you already taught it --> score drops --> becomes more careful, checks more thoroughly

  • The score doesn't control permissions — you're always in control. It's a mirror that helps calibrate rigor.

Sentiment Detection

NEXO Brain reads your tone (keywords, message length, urgency signals) and adapts:

  • Frustrated? --> Ultra-concise mode. Zero explanations. Just solve the problem.

  • In flow? --> Good moment to suggest that backlog item from last Tuesday.

  • Urgent? --> Immediate action, no preamble.

Sleep Cycle

Like a human brain, NEXO Brain has automated processes that run while you're not using it:

Time

Process

Human Analogy

03:00

Decay + memory consolidation + merge duplicates + dreaming

Deep sleep consolidation

04:00

Clean expired data, prune redundant memories

Synaptic pruning

07:00

Self-audit, health checks, metrics

Waking up + orientation

23:30

Process day's events, extract patterns

Pre-sleep reflection

Boot

Catch-up: run anything missed while computer was off

--

If your Mac was asleep during any scheduled process, NEXO Brain catches up in order when it wakes.

Deep Sleep now also mixes recent context with older context across a 60-day horizon. Instead of only looking at the immediate past, it can surface:

  • recurring multi-week themes

  • cross-domain links between older learnings and current failures

  • stale followups and topics that keep being mentioned but never formalized

  • weighted project pressure based on diary activity, followups, learnings, and decision outcomes

It now also writes weekly and monthly Deep Sleep summaries so the overnight system can reuse higher-horizon signals instead of rediscovering everything from scratch every day.

Cognitive Cortex

The Cortex is a middleware cognitive layer that makes the agent think before acting. It implements architectural inhibitory control — the agent cannot bypass reasoning.

User message → Fast Path check → Simple chat? → Respond directly
                                → Action needed? → Cortex activates
                                                    ↓
                                              Generate cognitive state
                                              (goal, plan, unknowns, evidence)
                                                    ↓
                                              Middleware validates
                                              ├─ Unknowns? → ASK mode (tools blocked)
                                              ├─ No plan? → PROPOSE mode (read-only)
                                              └─ Plan + evidence → ACT mode (full access)

Feature

What It Does

Inhibitory Control

Physically restricts tools based on reasoning quality. Unknowns → can only ask. No plan → can only propose. Evidence + verification → can act.

Event-Driven Activation

Only activates on tool intent, ambiguity, destructive actions, or retries. Simple chat has zero overhead.

Trust-Gated Escalation

Low trust score → requires more evidence before allowing "act" mode. Trust builds through successful execution.

Core Rules Injection

Automatically surfaces relevant behavioral rules based on task type.

Activation Metrics

Tracks modes, inhibition rates, and task types for continuous improvement.

The Cortex was designed through a 3-way AI debate (Claude Opus 4.6 + GPT-5.4 + Gemini 3.1 Pro) and validated against 6 months of real production failures.

Durable Workflow Runtime

Memory and guardrails are not enough if long work still restarts from zero.

NEXO now ships a durable workflow runtime for multi-step and cross-session execution:

  • nexo_workflow_open creates a persistent run with step metadata, idempotency key, priority, and shared state

  • nexo_workflow_update records replayable checkpoints, retry metadata, approval gates, and the current actionable state

  • nexo_workflow_resume tells the agent what to do next without guessing

  • nexo_workflow_replay reconstructs the recent execution history honestly instead of pretending the run is still in memory

  • nexo_workflow_list keeps active and blocked work visible so it does not disappear into reminders or prose notes

This is the bridge between "good memory" and "reliable execution": tasks can now preserve state, retries, approval gates, and next action across interruptions.

Context Continuity (Auto-Compaction)

NEXO Brain automatically preserves session context when Claude Code compacts conversations. Using PreCompact and PostCompact hooks:

  • PreCompact: Saves a complete session checkpoint to SQLite (task, files, decisions, errors, reasoning thread, next step)

  • PostCompact: Re-injects a structured Core Memory Block into the conversation, so the session continues seamlessly

This means long sessions (8+ hours) feel like one continuous conversation instead of restarting after each compaction.

How it works:

  1. Configure the hooks in your Claude Code settings.json

  2. NEXO Brain's heartbeat automatically maintains the checkpoint

  3. When compaction happens, the PreCompact hook reads the checkpoint and injects a recovery block

  4. The session continues from exactly where it left off

Setup:

{
  "hooks": {
    "PreCompact": [{
      "matcher": "*",
      "hooks": [{"type": "command", "command": "bash $NEXO_HOME/hooks/pre-compact.sh", "timeout": 10}]
    }],
    "PostCompact": [{
      "matcher": "*",
      "hooks": [{"type": "command", "command": "bash $NEXO_HOME/hooks/post-compact.sh", "timeout": 10}]
    }]
  }
}

2 new MCP tools: nexo_checkpoint_save (manual or hook-triggered checkpoint), nexo_checkpoint_read (retrieves the latest checkpoint for context injection).

Cognitive Features

NEXO Brain provides 150+ MCP tools across 23 categories. These features implement cognitive science concepts that go beyond basic memory:

Input Pipeline

Feature

What It Does

Prediction Error Gating

Only novel information is stored. Redundant content that matches existing memories is rejected at write time, keeping your memory clean without manual curation.

Security Pipeline

4-layer defense against memory poisoning: injection detection, encoding analysis, behavioral anomaly scoring, and credential scanning. Every memory passes through all four layers before storage.

Quarantine Queue

New facts enter quarantine status and must pass a promotion policy before becoming trusted knowledge. Prevents unverified information from influencing decisions. Automated nightly processing promotes, rejects, or expires items.

Secret Redaction

Auto-detects and redacts API keys, tokens, passwords, and other sensitive data before storage. Secrets never reach the vector database.

Memory Management

Feature

What It Does

Pin / Snooze / Archive

Granular lifecycle states for memories. Pin = never decays (critical knowledge). Snooze = temporarily hidden (revisit later). Archive = cold storage (searchable but inactive).

Intelligent Chunking

Adaptive chunking that respects sentence and paragraph boundaries. Produces semantically coherent chunks instead of arbitrary token splits, reducing retrieval noise.

Adaptive Decay

Decay rate still follows Ebbinghaus as the base model, but now also adapts per memory using stability and difficulty profiles. Frequently reinforced memories become stickier; fragile memories fade sooner.

Auto-Migration

Formal schema migration system (schema_migrations table) tracks all database changes. Safe, reversible schema evolution for production systems — upgrades never lose data.

Auto-Merge Duplicates

Batch cosine deduplication during the 03:00 sleep cycle. Respects sibling discrimination — similar memories about different contexts are kept separate.

Memory Dreaming

Discovers hidden connections between recent memories during the 03:00 sleep cycle and now feeds a 60-day long-horizon Deep Sleep blend, so older patterns can reappear when they become relevant again.

Operational Continuity

Feature

What It Does

Hot Context 24h

Keeps active topics, blockers, and waiting states fresh across sessions, clients, cron ticks, and channel changes. This is the shared recent-memory substrate for operational continuity.

Pre-Action Context Bundle

Loads recent contexts, recent events, related reminders, and related followups before acting, so continuity is explicit instead of prompt-only.

Transcript Fallback

When recent-memory capture is thin or missing, NEXO can now search and read recent Claude Code / Codex transcripts directly through MCP instead of pretending the conversation is lost.

Live System Catalog

NEXO can now inspect its own current surface — core tools, plugin tools, skills, scripts, crons, projects, and artifacts — through a live catalog derived from canonical sources at read time.

Retrieval

Feature

What It Does

HyDE Query Expansion

Generates hypothetical answer embeddings for richer semantic search. NEXO now auto-enables HyDE for conceptual or ambiguous queries while keeping literal lookups conservative.

Hybrid Search (FTS5+BM25+RRF)

Combines dense vector search with BM25 keyword search via Reciprocal Rank Fusion. Outperforms pure semantic search on precise terminology and code identifiers.

Cross-Encoder Reranking

After initial vector retrieval, a cross-encoder model rescores candidates for precision. The top-k results are reordered by true semantic relevance before being returned to the agent.

Multi-Query Decomposition

Complex questions are automatically split into sub-queries. Each component is retrieved independently, then fused for a higher-quality answer — improves recall on multi-faceted prompts.

Temporal Indexing

Memories are indexed by time in addition to semantics. Time-sensitive queries ("what did we decide last Tuesday?") use temporal proximity scoring alongside semantic similarity.

Spreading Activation

Graph-based co-activation network. NEXO now auto-enables a shallow spreading pass for concept-heavy queries, improving contextual recall without turning every exact lookup into a fuzzy search.

Recall Explanations

Transparent score breakdown for every retrieval result. Shows exactly why a memory was returned: semantic similarity, recency, access frequency, and co-activation bonuses.

Proactive

Feature

What It Does

Prospective Memory

Context-triggered reminders that fire when conversation topics match, not just by date. "Remind me about X when we discuss Y" works naturally.

Hook Auto-capture

Extracts decisions, corrections, and factual statements from conversations automatically. You don't need to explicitly say "remember this" — the system detects what's worth storing.

Session Summaries

Automatic end-of-session summarization that distills key decisions, errors, and follow-ups into a compact diary entry. The next session starts with full context — not a cold slate.

Smart Startup

Pre-loads relevant cognitive memories at session boot by composing a query from pending followups, due reminders, and last session's topics. Every session starts with the right context — not a cold search.

Context Packets

Bundles all area knowledge (learnings, recent changes, active followups, preferences, cognitive memories) into a single injectable packet for subagent delegation. Subagents never start blind again.

Auto-Prime by Topic

Heartbeat detects project/area keywords in conversation and automatically surfaces the most relevant learnings. No explicit memory query needed — context arrives proactively.

Benchmark: LoCoMo (ACL 2024)

NEXO Brain was evaluated on LoCoMo (ACL 2024), a long-term conversation memory benchmark with 1,986 questions across 10 multi-session conversations.

System

F1

Adversarial

Hardware

NEXO Brain v0.5.0

0.588

93.3%

CPU only

GPT-4 (128K full context)

0.379

—

GPU cloud

Gemini Pro 1.0

0.313

—

GPU cloud

LLaMA-3 70B

0.295

—

A100 GPU

GPT-3.5 + Contriever RAG

0.283

—

GPU

+55% vs GPT-4. Running entirely on CPU.

Key findings:

  • Outperforms GPT-4 (128K full context) by 55% on F1 score

  • 93.3% adversarial rejection rate — reliably says "I don't know" when information isn't available

  • 74.9% recall across 1,986 questions

  • Open-domain F1: 0.637 | Multi-hop F1: 0.333 | Temporal F1: 0.326

  • Runs on CPU with 768-dim embeddings (BAAI/bge-base-en-v1.5) — no GPU required

  • First MCP memory server benchmarked on a peer-reviewed dataset

Full results in benchmarks/locomo/results/.

Nervous System (v2.0.0)

NEXO Brain doesn't just respond — it runs 13 core recovery-aware background jobs plus optional helpers, like a biological nervous system. They handle maintenance, health monitoring, and self-improvement without any user interaction:

Script

Schedule

What It Does

cognitive-decay

03:00 daily

Ebbinghaus decay + memory consolidation + duplicate merging + dreaming

sleep

04:00 daily

Synaptic pruning, expired data cleanup

deep-sleep

04:30 daily

4-phase overnight pipeline: Collect→Extract→Synthesize→Apply. Analyzes all sessions, detects emotional patterns, abandoned projects, productivity issues, and auto-creates learnings

self-audit

07:00 daily

Health checks, guard stats, trust score review, metrics

postmortem

23:30 daily

Session consolidation, extract patterns from day's events

catchup

On boot

Runs any missed scheduled processes (Mac was off/asleep)

tcc-approve

On boot (macOS)

Auto-approve macOS permissions for Claude Code updates

prevent-sleep

Always (daemon)

Keeps machine awake for nocturnal processes (caffeinate/systemd-inhibit)

evolution

Weekly (Sun)

Self-improvement proposals — NEXO suggests and applies enhancements

followup-hygiene

Weekly (Sun)

Normalizes statuses, flags stale followups, cleans orphans

learning-housekeep

03:15 daily

Dedup learnings, adjust weights by usage, process overdue reviews, reconcile decision outcomes

immune

Every 30 min

Quarantine processing, memory promotion/rejection, synaptic pruning

impact-scorer

05:45 daily

Scores active followups so queues can prioritize by expected impact

synthesis

06:00 daily

Memory synthesis — discovers cross-memory patterns

outcome-checker

08:00 daily

Verifies tracked outcomes and marks them met, pending, or missed

watchdog

Every 30 min

Monitors services, LaunchAgents, and infrastructure health

auto-close-sessions

Every 5 min

Cleans stale sessions

Core processes are defined in src/crons/manifest.json and auto-synced to your system by nexo_update. On macOS they run via LaunchAgents; on Linux via systemd user timers. tcc-approve, prevent-sleep, and backup are platform/personal helpers — not in the manifest but listed above for completeness. Personal crons (your own scripts) are never touched by the sync. If your Mac was asleep during a scheduled process, the catch-up script re-runs everything in order when it wakes.

Deep Sleep v2 — Overnight Learning (v2.1.0)

Deep Sleep is a 4-phase pipeline that runs at 4:30 AM and makes NEXO smarter while you sleep:

Phase 1: COLLECT (Python)
├── Reads all session transcripts from the day
├── Splits each session into individual .txt files
└── Gathers DB state (followups, learnings, trust)

Phase 2: EXTRACT (Opus, one call per session)
├── 8 types of findings per session:
│   ├── Uncaptured corrections (user corrected agent, no learning saved)
│   ├── Self-corrected errors (knowledge gaps to fix)
│   ├── Unformalised ideas (mentioned but never tracked)
│   ├── Missed commitments (promised but no followup)
│   ├── Protocol violations (guard_check, heartbeat, change_log)
│   ├── Emotional signals (frustration, flow, satisfaction)
│   ├── Abandoned projects (started but not finished)
│   └── Productivity patterns (corrections, proactivity, tool efficiency)
└── Outputs per-session JSON with findings + emotional timeline

Phase 3: SYNTHESIZE (Opus, one call)
├── Cross-session patterns (same error in 5 sessions = systemic)
├── Daily mood arc with score (0.0 = terrible day, 1.0 = great day)
├── Recurring triggers (what causes frustration vs flow)
├── Productivity analysis (corrections, tool efficiency)
├── Abandoned project detection
├── Morning agenda (prioritized)
└── Calibration recommendations

Phase 4: APPLY (Python)
├── Auto-creates learnings from high-confidence findings
├── Creates followups for unfinished work
├── Updates mood_history in calibration.json (30-day rolling)
├── Generates session-tone.json (emotional guidance for next session)
└── Writes morning-briefing.md

Session Tone — Emotional Intelligence

Deep Sleep generates a session-tone.json that tells NEXO how to behave next morning:

  • Agent made many mistakes yesterday → Acknowledge them, show what was learned, demonstrate improvement

  • User had a bad day (mood < 40%) → Supportive approach, lighter start, avoid known frustration triggers

  • User had a great day (mood > 70%) → Reinforce momentum, reference wins, push ambitious goals

  • Agent was too reactive → Be proactive today, don't wait for instructions

This is read by nexo_smart_startup and injected into every session's context. NEXO adapts its personality based on real behavioral data, not just configuration.

Cron Manifest & Scheduler (v2.4.0)

All core crons are defined in src/crons/manifest.json. When you run nexo_update, the sync script:

  • Installs new crons from the manifest

  • Updates changed schedules/intervals

  • Removes crons no longer in the manifest (only core ones)

  • Never touches personal crons you created yourself

Every cron execution is tracked in the cron_runs table via a universal wrapper. Use nexo_schedule_status to see what ran overnight:

✅ deep-sleep: 1/1 OK, 4523s avg — 37 sessions, 259 findings
✅ immune: 48/48 OK, 2s avg
❌ evolution: 0/1 OK — CLI timeout

Add personal crons from conversation with nexo_schedule_add — generates LaunchAgent (macOS) or systemd timer (Linux) automatically.

Skill Auto-Creation (v2.4.0)

Deep Sleep automatically extracts reusable procedures from successful multi-step tasks and stores them as skills with full procedural content (steps, gotchas, markdown).

Pipeline: trace → draft → published → archived. Trust rises with successful use, decays without it. No human approval gates.

7 MCP tools: nexo_skill_create, nexo_skill_match, nexo_skill_get, nexo_skill_result, nexo_skill_list, nexo_skill_merge, nexo_skill_stats.

Dashboard (v1.6.0)

A web interface at localhost:6174 with 6 interactive pages for visual insight into your brain's state:

Page

What It Shows

Overview

System health at a glance — memory counts, trust score, active sessions, recent changes

Graph

Interactive D3.js visualization of the knowledge graph (nodes, edges, clusters)

Memory

Browse and search all memory stores (STM, LTM, sensory, archived)

Somatic

Pain map per file/area — see which parts of your codebase cause the most errors

Adaptive

Personality signals, learned weights, and current mode

Sessions

Active and historical sessions with timeline and diary entries

Built with FastAPI backend and D3.js frontend. Dashboard files are installed to NEXO_HOME/dashboard/ but must be started manually:

python3 ~/.nexo/dashboard/app.py

This opens localhost:6174 in your browser. Add --port 8080 to change the port or --no-browser to skip auto-opening.

Full Orchestration System

Memory alone doesn't make a co-operator. What makes the difference is the behavioral loop — the automated discipline that ensures every session starts informed, runs with guardrails, and ends with self-reflection.

Automated Hooks

7 hooks fire automatically at key moments in every Claude Code session:

Hook

When

What It Does

SessionStart (timestamp)

Session opens

Writes session timestamp for staleness detection

SessionStart (briefing)

Session opens

Generates briefing from SQLite: overdue reminders, today's tasks, pending followups, active sessions. Cleans up post-mortem flags.

Stop

Session ends

Mandatory post-mortem: self-critique (5 questions), session buffer entry, followup creation, proactive seeds for next session

PostToolUse (capture)

After each tool call

Captures meaningful mutations to the Sensory Register + auto-diary every 10 tool calls

PostToolUse (inbox)

After each tool call

Inter-terminal inbox delivery between parallel sessions

PreCompact

Before context compression

Saves full session checkpoint to SQLite — task, files, decisions, errors, reasoning thread + emergency diary

PostCompact

After context compression

Re-injects Core Memory Block so the session continues seamlessly from where it left off

The Session Lifecycle

Session starts
    ↓
SessionStart hook generates briefing
    ↓
Operator reads diary, reminders, followups
    ↓
Heartbeat on every interaction (sentiment, context shifts)
    ↓
Guard check before every code edit
    ↓
PreCompact hook saves full checkpoint if conversation is compressed
    ↓
PostCompact hook re-injects Core Memory Block → session continues seamlessly
    ↓
Stop hook refreshes the diary draft and approves immediately:
  - Latest changes and decisions stay attached to the active session
  - Session buffer keeps structured tool activity for downstream processing
  - Followups and closing synthesis happen inline when the agent detects real closing intent
  - No mid-conversation blocking from the hook itself
    ↓
Nocturnal post-mortem consolidator processes the buffer mechanically
    ↓
Nocturnal processes: decay, consolidation, self-audit, dreaming

Reflection Engine

NEXO still ships nexo-reflection.py as a standalone analyzer for session_buffer.jsonl. It is not currently auto-triggered by the stop hook:

  • Extracts recurring tasks, error patterns, mood trends

  • Updates user_model.json with observed behavior

  • No LLM required — runs as pure Python

Auto-Migration

Existing users upgrading from any previous version:

npx nexo-brain  # detects current version, migrates automatically
  • Updates hooks, core files, plugins, scripts, and LaunchAgent templates

  • Runs database schema migrations automatically

  • Never touches your data (memories, learnings, preferences)

  • Saves updated CLAUDE.md as reference (doesn't overwrite customizations)

Runtime CLI (v2.6.0)

NEXO Brain includes a local CLI that runs independently of any single terminal client:

  • nexo chat — launch a NEXO terminal client; if both Claude Code and Codex are available, it asks every time which one to open and puts the last-used client first

  • nexo update — sync runtime from source, run migrations, reconcile schedules

  • nexo doctor --tier runtime — boot/runtime/deep diagnostics with --fix mode

  • nexo scripts list — list all personal scripts and their status

  • nexo scripts reconcile — align declared schedules with actual LaunchAgents/systemd

  • nexo -v — show installed runtime version

The CLI lives at NEXO_HOME/bin/nexo and is added to your PATH during install.

Personal Scripts Registry (v2.6.0)

Scripts in NEXO_HOME/scripts/ are first-class managed entities:

  • Tracked in SQLite with metadata, categories, and schedule associations

  • Inline metadata in scripts declares name, runtime, schedule, and recovery policy

  • nexo scripts create NAME scaffolds a new script with the correct template

  • nexo scripts reconcile creates/repairs LaunchAgents from declared metadata

  • nexo scripts sync discovers filesystem state and updates the registry

  • nexo doctor --tier runtime detects orphaned schedules, missing plists, and drift

Personal scripts are completely separate from core NEXO processes. The crons/manifest.json defines core; everything in NEXO_HOME/scripts/ is personal.

If you need to decide between a personal script, skill, plugin, or schedule, use docs/personal-artifacts-manual.md. That is the canonical operational guide.

Recovery-Aware Background Jobs (v2.6.2)

Core and personal jobs now declare explicit recovery contracts in crons/manifest.json:

Field

Purpose

recovery_policy

catchup, restart, restart_daemon, or skip

run_on_boot

Re-run when the machine starts

run_on_wake

Re-run after sleep/resume

idempotent

Safe to re-run without side effects

max_catchup_age

Maximum age of a missed window to still catch up

If the Mac was asleep during a scheduled window, catchup detects the gap from cron_runs (not a state file) and re-executes eligible jobs once. Interval-based personal scripts get a single recovery run, not repeated ticks.

For personal daemon-style helpers, recovery_policy=restart_daemon plus schedule_required=true declares an official KeepAlive schedule. NEXO can now reconcile and repair those daemons instead of treating them as unmanaged legacy LaunchAgents.

Startup Preflight (v2.6.2)

Before nexo chat or MCP server start, NEXO runs a preflight check:

  1. Apply power policy (caffeinate on macOS, systemd-inhibit on Linux)

  2. Run safe local migrations and backfills

  3. Sync personal scripts registry

  4. For dev-linked runtimes: check if source repo is behind, pull if safe, sync to runtime

This replaces the old "blind startup" where NEXO entered without verifying runtime health.

Knowledge Graph (v0.8)

A bi-temporal entity-relationship graph with 988 nodes and 896 edges. Entities and relationships carry both valid-time (when the fact was true) and system-time (when it was recorded), enabling temporal queries like "what did we know about X last Tuesday?". BFS traversal discovers multi-hop connections between concepts. Event-sourced edges with smart dedup (ADD/UPDATE/NOOP) prevent redundant writes while preserving full history.

4 MCP tools: nexo_kg_query (SPARQL-like queries), nexo_kg_path (shortest path between entities), nexo_kg_neighbors (direct connections), nexo_kg_stats (graph metrics).

Cross-Platform Support

Full Linux support and Windows via WSL. The installer detects the platform and configures the appropriate process manager (LaunchAgents on macOS, catch-up on startup for Linux). PEP 668 compliance (venv on Ubuntu 24.04+). Session keepalive prevents phantom sessions during long tasks. Opportunistic maintenance runs cognitive processes when resources are available.

Windows users: NEXO Brain requires WSL (Windows Subsystem for Linux). Install WSL first, then run npx nexo-brain inside the Ubuntu/WSL terminal.

Storage Router

A new abstraction layer routes storage operations through a unified interface, making the system multi-tenant ready. Each operator's data is isolated while sharing the same cognitive engine.

Learned Weights & Somatic Markers (v0.7.0)

Adaptive Learned Weights

Signal weights learn from real user feedback via Ridge regression. A 2-week shadow mode observes before activating. Weight momentum (85/15 blend) prevents personality whiplash. Automatic rollback if correction rate doubles.

Somatic Markers (Pain Memory)

Files and areas that cause repeated errors accumulate a risk score (0.0–1.0). The guard system warns on HIGH RISK (>0.5) and CRITICAL RISK (>0.8), lowering thresholds for more paranoid checking. Clean guard checks reduce risk multiplicatively (×0.7). Nightly decay (×0.95) ensures old pain fades.

Adaptive Personality v2

6 weighted signals: vibe, corrections, brevity, topic, tool errors, git diff. Emergency keywords bypass hysteresis. Severity-weighted decay. Manual override via nexo_adaptive_override.

Quick Start

Claude Code (Primary)

npx nexo-brain

The installer handles everything and syncs the same nexo MCP brain into Claude Code, Claude Desktop, and Codex when those clients are present:

  How should I call myself? (default: Nova) > Atlas

  Can I explore your workspace to learn about your projects? (y/n) > y

  Keep Mac awake so my cognitive processes run on schedule? (y/n) > y

  Installing cognitive engine dependencies...
  Setting up NEXO home...
  Scanning workspace...
    - 3 git repositories
    - Node.js project detected
  Configuring MCP server...
  Setting up nervous system...
    15 core recovery-aware jobs configured.
    Dashboard configured at localhost:6174.
  Caffeinate enabled.
  Generating operator instructions...

  +----------------------------------------------------------+
  |  Atlas is ready. Type 'atlas' to start.                  |
  +----------------------------------------------------------+

Docker Compose

NEXO now ships a root-level docker-compose.yml for a persistent containerized runtime. It does two things at once:

  • keeps NEXO_HOME on a named volume

  • exposes a remote MCP endpoint at http://localhost:8000/mcp for IDEs that support HTTP/SSE MCP

Start it with:

docker compose up -d

For Claude Code and Codex, keep using stdio and point the MCP command at the running container:

docker compose exec -T nexo python src/server.py

That gives you the same persistent brain in the container while keeping terminal clients on their native stdio transport. The full step-by-step flow, health checks, and config examples live in docs/docker-setup.md.

Starting a Session

After install, use the runtime CLI:

nexo chat          # Launch a NEXO terminal client (asks if both Claude Code and Codex are available)
nexo doctor        # Check runtime health
nexo update        # Pull latest version and sync
nexo clients sync  # Re-sync Claude Code/Desktop/Codex to the same brain
nexo scripts list  # See your personal scripts

During install, NEXO now asks which interactive clients you want to connect, which one nexo chat should suggest first when multiple terminal clients are available, whether to enable background automation, which backend should run that automation, and which model profile each active terminal/backend should use. Shared brain stays on in every mode.

Public entry points for the mental model now stay intentionally small:

  • nexo_remember

  • nexo_memory_recall

  • nexo_consolidate

  • nexo_run_workflow

  • nexo_pre_action_context

  • nexo_transcript_search

  • nexo_system_catalog

If you want the shell or Python wrappers instead of raw MCP tools:

The model you pick during install is used everywhere — interactive sessions, automation scripts, and all task profiles. Change it once in your preferences and every part of the system follows. Default: Opus 4.7 with 1M context.

Or use the shell alias created during install (e.g. atlas), which now runs nexo chat . so it opens the terminal client you pick for that session, with the last-used option shown first.

Your operator will greet you immediately — adapted to the time of day, resuming from where you left off. No cold starts.

Contributing

NEXO is being hardened in public, and the best contributions now are not only code changes but also real workflow feedback:

  • Open issues when a client flow feels asymmetric across Claude Code, Codex, Claude Desktop, OpenClaw, or other MCP environments.

  • Send PRs for docs, install UX, tests, compatibility checks, and public-facing copy.

  • If you use NEXO in production-like daily work, include exact runtime symptoms and commands in bug reports. This project improves fastest when the operational reality is concrete.

The project still recommends Claude Code as the primary path, but contributions that improve Codex, client parity, installer clarity, and ecosystem integrations are especially valuable.

Maintainers and contributors touching startup, bootstrap, Deep Sleep, or shared-brain behavior should also use the client parity checklist:

What Gets Installed

Component

What

Where

Cognitive engine

Python: fastembed, numpy, vector search

pip packages

MCP server

150+ tools for memory, cognition, learning, guard

NEXO_HOME/

Claude Code Plugin

Marketplace-ready (packaging verified)

.claude-plugin/

Plugins

Guard, episodic memory, cognitive memory, entities, preferences, update, etc.

Code: src/plugins/, Personal: NEXO_HOME/plugins/

Hooks (7)

SessionStart, Stop, PostToolUse, PreCompact, PostCompact

NEXO_HOME/hooks/

Nervous system

13 core recovery-aware jobs + optional helpers (dashboard, prevent-sleep)

NEXO_HOME/scripts/

Dashboard

Web UI at localhost:6174 (23 modules, dark theme) — opt-in, always-on

NEXO_HOME/dashboard/

Runtime CLI

nexo command: scripts, doctor, skills, update

NEXO_HOME/bin/

Doctor

Unified diagnostics: boot/runtime/deep tiers, --fix mode

src/doctor/

Skills v2

Executable skills with guide/execute/hybrid modes, approval levels

NEXO_HOME/skills/

Startup Preflight

Health checks before every nexo chat or server start

Built into CLI

CLAUDE.md

Complete operator instructions (Codex, hooks, guard, trust, memory)

~/.claude/CLAUDE.md

Schedule config

schedule.json with customizable process times and timezone

NEXO_HOME/config/

Auto-update

Non-blocking startup check (5s max), opt-out via schedule.json

Built into server startup

CLAUDE.md tracker

Version-tracked core sections with safe updates preserving customizations

Built into auto-update

Shared client sync

Same nexo MCP entry wired into Claude Code, Claude Desktop, and Codex

User config dirs

Client/backend preferences

Selected interactive clients, default terminal client, automation backend, and model/reasoning profiles per client

NEXO_HOME/config/schedule.json

Auto-diary

3-layer system: PostToolUse every 10 calls, PreCompact emergency, heartbeat DIARY_OVERDUE

Built into hooks

Claude Code config

MCP server + 7 hooks + 15 managed processes registered

~/.claude/settings.json

Runtime CLI

After installation or auto-update, NEXO adds NEXO_HOME/bin to your shell PATH. Open a new terminal and the nexo command provides operational tools:

# Personal Scripts
nexo scripts list              # List your personal scripts
nexo scripts run my-script     # Run a script with injected NEXO env
nexo scripts doctor            # Validate all personal scripts
nexo scripts call nexo_learning_search --input '{"query":"cron"}' # Call any MCP tool

# Skills v2
nexo skills sync               # Sync filesystem skill definitions into SQLite
nexo skills list               # List published/stable skills
nexo skills get SK-...         # Inspect a skill definition
nexo skills apply SK-... --dry-run --json  # Resolve guide/execute/hybrid without running it
nexo skills approve SK-... --execution-level local --approved-by Francisco  # Optional metadata override
nexo skills evolution          # Show text→script and improvement candidates

# Unified Doctor
nexo doctor                    # Quick boot diagnostics
nexo doctor --tier all         # Full system check (boot + runtime + deep)
nexo doctor --tier runtime --json  # Machine-readable health report
nexo doctor --fix              # Apply deterministic repairs

Personal scripts live in NEXO_HOME/scripts/ with inline metadata. Their Python templates now include run_automation_text(...), which routes work through the configured NEXO automation backend instead of hardcoding claude -p or provider-specific model names. nexo-agent-run.py now also supports task profiles (fast, balanced, deep) plus safe backend fallback, so automations can prefer cheaper/faster Codex paths or deeper Claude paths without hardcoding one provider forever. See docs/writing-scripts.md for details and docs/personal-artifacts-manual.md for the canonical artifact decision guide.

Skills v2 combine procedural guides with optional executable scripts. Personal skills live in NEXO_HOME/skills/, packaged core skills live in NEXO_CODE/skills/ during development and NEXO_HOME/skills-core/ in installed environments, and staged runtime copies live in NEXO_HOME/skills-runtime/. Execution is fully autonomous: Deep Sleep can evolve mature guide skills into executable drafts automatically, and runtime execution no longer waits for manual approval. See docs/skills-v2.md for the full model and docs/personal-artifacts-manual.md for the boundary between skills, scripts, plugins, and schedules.

The Doctor system reads existing health artifacts (immune, watchdog, self-audit) without triggering repairs in default mode.

Requirements

  • macOS or Linux (Windows via WSL)

  • Node.js 18+ (for the installer)

  • Claude Code is the primary recommended client. It remains the most mature NEXO path: native hooks, the most battle-tested automation contract, and the clearest parity with historical production behavior.

  • Model: You pick your model during install and every component uses it. Default is Opus 4.7 with 1M context. Scripts and automation profiles read from a single preference — no hardcoded model strings.

  • Python 3, Homebrew, and the selected required client/backend can be installed automatically when NEXO has a supported installer path for that dependency.

Architecture

Unified Code/Data Separation (v2.0.0)

NEXO Brain separates code (immutable, in the repo or npm package) from data (personal, in NEXO_HOME):

Path

Contents

src/ (or npm package)

Server, plugins, hooks, scripts — never modified at runtime

NEXO_HOME/ (default ~/.nexo/)

Database, config, personal plugins, schedule, backups

NEXO_HOME/config/schedule.json

Customizable process schedules, timezone, auto_update flag

NEXO_HOME/plugins/

Personal plugins that override or extend repo plugins

NEXO_HOME/data/

SQLite databases (nexo.db, cognitive.db), migration state

The plugin loader scans src/plugins/ first (base), then NEXO_HOME/plugins/ (personal override by filename). This dual-directory approach lets you extend NEXO without forking the repo. The client sync layer points Claude Code, Claude Desktop, and Codex at the same runtime and NEXO_HOME, so all three clients share one brain instead of drifting into separate local memories.

150+ MCP Tools across 23 Categories

Category

Count

Tools

Purpose

Cognitive

8

retrieve, stats, inspect, metrics, dissonance, resolve, sentiment, trust

The brain — memory, RAG, trust, mood

Cognitive Input

5

prediction_gate, security_scan, quarantine, promote, redact

Input pipeline — gating, security, quarantine

Cognitive Advanced

8

hyde_search, spread_activate, explain_recall, dream, prospect, hook_capture, pin, archive

Advanced retrieval, proactive, lifecycle

Guard

3

check, stats, log_repetition

Metacognitive error prevention

Episodic

10

change_log/search/commit, decision_log/outcome/search, review_queue, diary_write/read, recall

What happened and why

Sessions

4

startup, heartbeat, stop, status

Session lifecycle + context shift detection + inter-terminal auto-inbox

Coordination

7

track, untrack, files, send, ask, answer, check_answer

Multi-session file coordination + messaging

Reminders

5

list, create, update, complete, delete

User's tasks and deadlines

Followups

4

create, update, complete, delete

System's autonomous verification tasks

Learnings

5

add, search, update, delete, list

Error patterns and prevention rules

Credentials

5

create, get, update, delete, list

Local credential storage (plaintext SQLite — protect with filesystem permissions)

Task History

3

log, list, frequency

Execution tracking and overdue alerts

Menu

1

menu

Operations center with box-drawing UI

Entities

5

search, create, update, delete, list

People, services, URLs

Preferences

4

get, set, list, delete

Observed user preferences

Agents

5

get, create, update, delete, list

Agent delegation registry

Backup

3

now, list, restore

SQLite data safety

Evolution

5

propose, approve, reject, status, history

Self-improvement proposals

Adaptive & Somatic

4

adaptive_weights, adaptive_override, somatic_check, somatic_stats

Learned signal weights + pain memory per file

Knowledge Graph

4

kg_query, kg_path, kg_neighbors, kg_stats

Bi-temporal entity-relationship graph

Context Continuity

2

checkpoint_save, checkpoint_read

Auto-compaction session preservation

Personal Scripts

9

sync, list, create, remove, schedules, unschedule, reconcile, classify, ensure_schedules

Script lifecycle management

Skills

12

match, create, get, list, apply, approve, result, stats, evolution_candidates, merge, sync, featured

Reusable procedure library

Schedule

2

add, status

Personal cron scheduling

Doctor

1

doctor

Runtime diagnostics with --fix

Update

1

update

Pull latest code, backup, migrate, verify (with rollback)

Plugin System

NEXO Brain supports hot-loadable plugins with a dual-directory loader. Base plugins live in src/plugins/ (repo). Personal plugins go in NEXO_HOME/plugins/ and can override base plugins by filename. Drop a .py file in NEXO_HOME/plugins/:

# my_plugin.py
def handle_my_tool(query: str) -> str:
    """My custom tool description."""
    return f"Result for {query}"

TOOLS = [
    (handle_my_tool, "nexo_my_tool", "Short description"),
]

Reload without restarting: nexo_plugin_load("my_plugin.py")

Use a personal plugin only when you need a new MCP tool in the runtime surface. If the real need is autonomous execution or scheduling, use a personal script plus managed schedule instead. The canonical decision guide is docs/personal-artifacts-manual.md.

Data Privacy

  • Everything stays local. All data in ~/.nexo/, never uploaded anywhere.

  • No telemetry. No analytics. No phone-home.

  • No cloud dependencies. Vector search runs on CPU (fastembed), not an API.

  • Auto-update is resilient. NEXO checks for updates on startup. If an update fails, it continues with the current version and notifies you. Local migrations (database schema, configuration) always run. Network updates (git pull) can be disabled by setting auto_update: false in NEXO_HOME/config/schedule.json.

  • Secret redaction. API keys and tokens are stripped before they ever reach memory storage.

The Psychology Behind NEXO Brain

NEXO Brain isn't just engineering — it's applied cognitive psychology:

Psychological Concept

How NEXO Brain Implements It

Atkinson-Shiffrin (1968)

Three memory stores: sensory register --> STM --> LTM

Ebbinghaus Forgetting Curve (1885)

Exponential decay: strength = strength * e^(-lambda * time)

Rehearsal Effect

Accessing a memory resets its strength to 1.0

Memory Consolidation

Nightly process promotes frequently-used STM to LTM

Prediction Error

Only surprising (novel) information gets stored — redundant input is gated

Spreading Activation (Collins & Loftus, 1975)

Retrieving a memory co-activates related memories through an associative graph

HyDE (Gao et al., 2022)

Hypothetical document embeddings improve semantic recall

Prospective Memory (Einstein & McDaniel, 1990)

Context-triggered intentions fire when cue conditions match

Metacognition

Guard system checks past errors before acting

Cognitive Dissonance (Festinger, 1957)

Detects and verbalizes conflicts between old and new knowledge

Theory of Mind

Models user behavior, preferences, and mood

Synaptic Pruning

Automated cleanup of weak, unused memories

Associative Memory

Semantic search finds related concepts, not just matching words

Memory Reconsolidation

Dreaming process discovers hidden connections during sleep

Integrations

Claude Code (Primary)

NEXO Brain is designed as an MCP server. Claude Code remains the primary recommended client and the most complete integration path:

npx nexo-brain

All 150+ tools are available immediately after installation. The installer configures Claude Code's ~/.claude/settings.json automatically. The recommended Claude profile is Opus 4.7 with 1M context.

Claude Desktop

When Claude Desktop is installed, nexo-brain, nexo update, and nexo clients sync keep claude_desktop_config.json pointed at the same local NEXO runtime and NEXO_HOME.

Codex

When Codex CLI is available, nexo-brain, nexo update, and nexo clients sync register the same nexo MCP server via codex mcp add, so Codex uses the same local memory store as Claude Code and Claude Desktop. If selected during install, nexo chat can open Codex directly and background automation can also run through Codex. Interactive nexo chat launches use Codex's aggressive no-confirmation mode so the session does not stall on repetitive approval prompts. Codex uses the same model you configured during install — no separate model override is needed. Runtime Doctor also audits recent Codex sessions for NEXO startup markers and conditioned-file protocol discipline so parity drift does not hide behind the lack of native Claude-style hooks.

Cursor

Cursor works well as a documented companion client. Point Cursor at the same local nexo MCP server and add a project rule that forces nexo_startup, nexo_heartbeat, and the protocol path on real work. See docs/integrations/cursor.md.

Windsurf

Windsurf/Cascade supports MCP plus durable repo rules. Use the same local nexo server and add NEXO startup/protocol instructions in .windsurf/rules/ or your repo AGENTS.md. See docs/integrations/windsurf.md.

Gemini CLI

Gemini CLI can share the same local NEXO brain through mcpServers in ~/.gemini/settings.json plus a repo GEMINI.md. NEXO now ships a starter adapter in adapters/gemini/README.md.

OpenClaw

NEXO Brain also works as a cognitive memory backend for OpenClaw:

MCP Bridge (Zero Code)

Add NEXO Brain to your OpenClaw config at ~/.openclaw/openclaw.json:

{
  "mcp": {
    "servers": {
      "nexo-brain": {
        "command": "python3",
        "args": ["~/.nexo/server.py"],
        "env": {
          "NEXO_HOME": "~/.nexo"
        }
      }
    }
  }
}

Or via CLI:

openclaw mcp set nexo-brain '{"command":"python3","args":["~/.nexo/server.py"],"env":{"NEXO_HOME":"~/.nexo"}}'
openclaw gateway restart

ClawHub Skill

npx clawhub@latest install nexo-brain

Native Memory Plugin

npm install @wazionapps/openclaw-memory-nexo-brain
{
  "plugins": {
    "slots": {
      "memory": "memory-nexo-brain"
    }
  }
}

This replaces OpenClaw's default memory system with NEXO Brain's full cognitive architecture.

Any MCP Client

NEXO Brain works with any application that supports the MCP protocol. Configure it as an MCP server pointing to server.py inside NEXO_HOME (default ~/.nexo/server.py), with the NEXO_HOME env var set to the same directory.

Listed On

Directory

Type

Link

npm

Package

nexo-brain

Glama

MCP Directory

glama.ai

mcp.so

MCP Directory

mcp.so

mcpservers.org

MCP Directory

mcpservers.org

OpenClaw

Native Plugin

openclaw.com

dev.to

Technical Article

How I Applied Cognitive Psychology to AI Agents

Claude Code

Plugin (marketplace-ready)

Packaging verified, included in npm tarball

nexo-brain.com

Official Website

nexo-brain.com

Support the Project

If NEXO Brain is useful to you, consider:

  • Star this repo — it helps others discover the project and motivates continued development

  • Sponsor on GitHub — support ongoing development directly

  • Share your experience — tell others how you're using cognitive memory in your AI workflows

  • Contribute — see CONTRIBUTING.md for guidelines. Issues and PRs welcome

  • Client parity / shared-brain maintenance — see docs/client-parity-checklist.md

  • Writing a personal script that calls the automation backend — see docs/personal-scripts-guide.md

Star History Chart

Memory Benchmark Snapshot

The full harness is in benchmarks/README.md. The first checked-in micro-benchmark compares the NEXO runtime against a static CLAUDE.md-only baseline on five recall-heavy scenarios:

Scenario

NEXO full stack

Static CLAUDE.md

No memory

Decision rationale recall

Pass

Partial

Fail

User preference recall

Pass

Partial

Fail

Repeat-error avoidance

Pass

Partial

Fail

Resume interrupted task

Pass

Partial

Fail

Related-context stitching

Pass

Fail

Fail

See benchmarks/results/memory-recall-vs-static.md for the rubric, prompt shape, and first-run notes.

Changelog

v3.0.1 — Python 3.10 Compatibility Patch (2026-04-06)

  • Restored Python 3.10 compatibility by replacing Python 3.11-only datetime.UTC with timezone.utc.

  • Added tomllib → tomli fallback plus declared runtime dependency for Python < 3.11.

  • Boot doctor now validates all critical JSON config artifacts: schedule.json, optionals.json, crons/manifest.json.

v3.0.0 — Protocol Discipline, Durable Execution, Measured Runtime (2026-04-06)

  • Protocol discipline runtime: Enforceable nexo_task_open/nexo_task_close, persistent protocol_debt, Cortex gates with durable check_id, conditioned-file guardrails across Claude hooks and Codex transcript audits.

  • Durable workflow runtime: nexo_workflow_open/update/resume/replay/list with persistent runs, steps, checkpoints, replay history, retry bookkeeping, and idempotent open keys.

  • Durable goals: nexo_goal_open/update/get/list for long-running work that stays active/blocked/abandoned/completed.

  • Operational truth: Deep Sleep survives schema drift, keep_alive reports alive/degraded/duplicated honestly, warning storms no longer count as healthy.

  • Measured product surface: 5-minute quickstart, Python SDK, reference verticals, measured compare scorecard with LoCoMo baselines and cost_per_solved_task.

  • Skill lifecycle: Testing, promotion, retirement, and composition flows. Evolution public-core peer-review for opt-in PRs.

v2.7.0 — Shared Brain Baseline (2026-04-06)

  • Managed Claude Code + Codex bootstrap with explicit CORE/USER contract.

  • Codex config sync and transcript-aware Deep Sleep across both clients.

  • 60-day long-horizon analysis, weekly/monthly summary artifacts.

  • Retrieval auto-mode and first measured engineering loop.

  • nexo chat opens the configured client instead of assuming Claude Code.

v2.6.9 — Integration Sync, CI/CD Pipeline (2026-04-04)

  • Release artifact sync: Automated version synchronization across Claude Code plugin, OpenClaw package, and ClawHub skill before every publish.

  • CI/CD pipeline: Full GitHub Actions workflow for publish + verification of all integration channels.

  • OpenClaw plugin hardened: Contract tests, correct runtime path, synchronized version. Published as @wazionapps/openclaw-memory-nexo-brain@2.6.9.

  • ClawHub skill hardened: Version-synced metadata, correct server path, post-publish smoke verification.

  • Claude Code plugin packaging: Verified plugin.json, .mcp.json, hooks included in npm tarball. Marketplace-ready.

v2.6.5 — Power Helper Hardening, Recovery Contracts (2026-04-04)

  • Power helper semantics explicit and safer: always_on = platform helper for best-effort background availability.

  • Catch-up recovery suppresses duplicate relaunches for in-flight cron_runs.

  • Runtime update/startup reconciles declared personal schedules automatically.

v2.6.3 — Cron Sync Fix, Hook Migration (2026-04-04)

  • Runtime cron sync skips same-file copies, avoiding SameFileError on synced runtimes.

  • Core hook migration normalizes legacy flat entries into Claude Code's required matcher + hooks[] format.

v2.6.2 — Startup Preflight, Personal Recovery, Power Policy (2026-04-04)

  • Startup preflight before nexo chat and server — safe local migrations, deferred remote updates.

  • Personal managed schedules can declare recovery contracts (wake/boot/catchup).

  • Persisted runtime power policy (always_on/disabled/unset). Installer and nexo update prompt once.

  • Packaged installs resolve update root correctly (fixes vunknown).

v2.6.0 — Personal Scripts Registry, Plugin Marketplace, Managed Evolution (2026-04-03)

  • Personal scripts registry: Scripts in NEXO_HOME/scripts/ tracked in SQLite with metadata, categories, schedules. Full lifecycle: create, sync, reconcile, schedule, unschedule, remove.

  • Orchestrator removed from core (breaking): Was opt-in personal automation adding complexity for all users. Existing users keep their setup in NEXO_HOME/scripts/.

  • Claude Code plugin structure: plugin.json, entry point, packaging for marketplace submission.

  • nexo chat: Official command to launch a NEXO terminal client, asking when multiple supported terminal clients are available.

  • Managed Evolution hardening: Can modify core behavior modules with rollback followups.

  • Cron recovery hardened: TCC diagnostics, keepalive sync, personal schedule catchup.

v2.5.0 — Runtime CLI, Doctor, Skills v2, Day Orchestrator (2026-04-03)

  • Runtime CLI (nexo): New operational CLI separate from installer. nexo scripts list/run/doctor/call for personal scripts, nexo doctor for diagnostics, nexo skills apply for executable skills, nexo update for one-step sync.

  • Unified Doctor: Modular diagnostic system with boot/runtime/deep tiers. Report-only by default, deterministic --fix mode. MCP tool nexo_doctor. LaunchAgent schedule drift detection and reconciliation.

  • Skills v2: Executable skills with guide/execute/hybrid modes. Security levels (read-only/local/remote) with explicit approval. Core vs personal vs community directories. Deep Sleep auto-evolution integration.

  • Day Orchestrator: Autonomous NEXO cycles every 15 min (8:00-23:00). Launches Claude Code headless with full MCP. Checks followups, emails, infra — acts autonomously, emails user only when needed. Opt-in.

  • Dashboard always-on: Web UI at localhost:6174 as persistent LaunchAgent. 23 modules, Jinja2 templating, dark theme. Opt-in.

  • Personal Scripts Framework: Auto-discovery in NEXO_HOME/scripts/, inline metadata, runtime detection, forbidden-pattern validation, vendorable helper, template.

  • Configurable operator name (UserContext singleton), watchdog normalized to 30 min, LaunchAgent drift fix.

v2.4.0 — Skills, Cron Scheduler, Security, Full Audit (2026-04-03)

  • Skill Auto-Creation: Deep Sleep extracts reusable procedures from sessions. Content stored as markdown with steps and gotchas. Trust pipeline with autonomous quality control.

  • Cron Scheduler: execution tracking (cron_runs table), nexo_schedule_status and nexo_schedule_add MCP tools, universal cron wrapper for all processes.

  • Deep Sleep v2.4: watermark-based collection (late-night sessions included), per-session checkpointing (crash-safe), retry x3, JSON parsing fix, auto-calibration of personality settings.

  • Security: credential redaction in tool logs, transcript sanitization, command injection fix in dashboard, path traversal protection in plugin loader.

  • Diary filter: startup only shows human sessions, auto-closed cron sessions filtered out. Email sessions preserved as real interactions.

  • Preflight CI: 66 automated checks (py_compile, bash -n, manifest consistency, npm artifact, forbidden markers).

  • Python 3.9 compat: from __future__ import annotations across 18 files.

  • Linux: full systemd timer support, .bashrc alias for interactive shells.

  • Passed 5-phase automated audit: Product, Failure, Security, Packaging, UX.

v2.2.0 — Trust Score v2 (2026-04-01)

  • Trust Score: fair daily calibration from Deep Sleep analysis. Score 0-100 based on corrections, autonomy, proactivity.

  • Cognitive Quarantine: new memories go through quarantine before promotion to LTM.

v2.0.0 — Unified Architecture (2026-03-31)

  • Code/data separation: Code in repo (src/), personal data in NEXO_HOME (default ~/.nexo/). NEXO_HOME env var required.

  • Plugin loader dual-directory: Scans src/plugins/ (base) then NEXO_HOME/plugins/ (personal override by filename).

  • Auto-update on startup: Non-blocking (5s max), resilient, opt-out via schedule.json. Separate from manual nexo_update tool.

  • Auto-diary: 3-layer system — PostToolUse every 10 calls, PreCompact emergency save, heartbeat DIARY_OVERDUE signal.

  • CLAUDE.md version tracker: Section markers enable safe core updates without losing user customizations.

  • schedule.json: Customizable process schedules with timezone support and auto_update flag.

  • 15 autonomous processes: Added auto-close-sessions, synthesis, backup, tcc-approve, prevent-sleep (cross-platform).

  • 7 hooks: SessionStart (timestamp + briefing), Stop, PostToolUse (capture + inbox), PreCompact, PostCompact.

  • 150+ MCP tools: Added nexo_update tool for manual updates with rollback.

  • Lambda fix: Decay values were 24x too aggressive (STM: 7h to 7d, LTM: 2.4d to 60d).

  • Guard scoping: Was returning 35+ irrelevant blocking rules; now scoped to area and gated to high/critical.

  • 12 rounds of external audit: ~60 findings resolved.

v1.7.0 — Full Internationalization + Linux Support (2026-03-31)

  • Full i18n: All UI strings, error messages, DB status values in English. NLP detection patterns retain bilingual keywords (Spanish + English) for multilingual user support.

  • Linux support: systemd user timers (preferred) or crontab fallback for all automated cognitive processes.

  • Auto-resolve followups: Change log entries automatically cross-reference and complete matching open followups.

  • Free-form learning categories: No more hardcoded category validation — use any category name.

  • CLAUDE.md template rewrite: 494 to 127 lines, compact procedural format with full heartbeat signal reactions.

  • Complete sanitization: All hardcoded paths use NEXO_HOME env var. No credentials or personal data in the distributed package. Migration scripts and maintainer tooling use configurable paths.

v1.6.0 — Nervous System + Dashboard v2 (2026-03-30)

  • Nervous System: 11 autonomous scripts (decay, deep sleep, self-audit, catchup, evolution, followup hygiene, immune, watchdog, github monitor, learning validator)

  • Dashboard v2: 6 interactive pages at localhost:6174 (Overview, Graph, Memory, Somatic, Adaptive, Sessions)

  • LaunchAgent Templates: macOS automation templates included in the package for scheduling the nervous system

  • Hooks: 7 total — SessionStart, Stop, PostToolUse, PreCompact, PostCompact

  • Installer: Now configures dashboard LaunchAgent, nervous system scripts, and all templates automatically

v1.5.2 — Deep Sleep (2026-03-29)

  • Deep Sleep: Reads full session transcripts (not just diary) — finds uncaptured corrections, protocol violations, missed commitments

  • Uses Claude CLI in --bare mode (no hooks, no CLAUDE.md interference)

  • Catch-up system re-runs yesterday if the Mac was off

v1.5.0 — Modular Core + Knowledge Graph Search (2026-03-29)

  • Architecture: db.py refactored into db/ package (11 modules); cognitive.py into cognitive/ package (6 modules)

  • KG Boost: Knowledge Graph connection count influences search result ranking

  • HNSW Vector Index: Optional approximate nearest neighbor acceleration (auto-activates above 10,000 memories)

  • Claim Graph: Decomposes blob memories into atomic verifiable facts with provenance and contradiction detection

  • Inter-terminal Auto-inbox (D+): nexo_startup accepts claude_session_id for automatic inbox delivery between parallel terminals

  • Tests: 156 pytest tests across 3 suites (cognitive, knowledge graph, migrations)

v1.4.1 — Multi-AI Code Review (2026-03-29)

  • Fix: 3 bugs found by GPT-5.4 (Codex CLI) + Gemini 2.5 (Gemini CLI) reviewing full codebase

  • Security: Memory sanitization prevents prompt injection via stored content

  • Migration #13: Normalizes legacy status values on upgrade

v1.4.0 — The Brain Dreams (2026-03-29)

  • Major: All 9 nightly scripts migrated from Python word-overlap to CLI wrapper pattern

  • Stop Hook v8: Session-scoped tool counting, buffer fallback removed

  • Guard: Behavioral rules section surfaces most-violated rules at session start

v1.3.0 — Evolution System (2026-03-28)

  • New: Self-improvement cycle — NEXO proposes and applies improvements weekly

  • Dual-mode: auto (low-risk) and review (owner approval required)

  • Circuit breaker, snapshot/rollback, immutable file protection

v1.2.3 — AGPL-3.0 License (2026-03-27)

  • License changed from MIT to AGPL-3.0

v1.2.1 — Stop Hook Hotfix (2026-03-27)

  • Fix: v1.2.0 deleted the flag on approve, causing infinite block loops if session didn't close immediately

  • Fix: Removed TTL on flag — it persists until SessionStart cleans it up next session

  • New: Trivial sessions (<5 meaningful tool calls) skip post-mortem entirely and approve immediately

  • SessionStart hook now cleans up .postmortem-complete flag on session start

v1.2.0 — Blocking Stop Hook (2026-03-27)

  • Fix: Stop hook now uses "decision": "block" instead of "approve" to enforce post-mortem execution

  • Previous behavior: hook injected systemMessage but AI had already responded — instructions were never processed

  • New behavior: session close is blocked until AI completes self-critique, session diary, buffer entry, and followups

  • Flag-based mechanism (.postmortem-complete) allows second close attempt to succeed

  • Works for all NEXO users, not just specific setups

v1.1.1 — Multi-terminal fix (2026-03-27)

  • Fix: PostCompact now reads the correct session's checkpoint in multi-terminal setups

  • Changelog section added to README

v1.1.0 — Context Continuity (2026-03-27)

  • Context Continuity: PreCompact/PostCompact hooks preserve session state across compaction events

  • New session_checkpoints SQLite table + migration #12

  • New tools: nexo_checkpoint_save, nexo_checkpoint_read

  • Heartbeat automatically maintains checkpoint every interaction

  • Core Memory Block re-injected post-compaction with task, files, decisions, reasoning thread

  • 115+ total tools at the time, 20 categories

v1.0.0 — Cognitive Cortex + Stable Release (2026-03-26)

  • Cognitive Cortex: architectural inhibitory control (ASK/PROPOSE/ACT modes)

  • 30 Core Rules as immutable DNA in SQLite

  • Designed via 3-way AI debate (Claude Opus + GPT-5.4 + Gemini 3.1 Pro)

  • Artifact Registry for operational facts

  • Full benchmark suite (LoCoMo F1: 0.588)

v0.10.0 — Smart Context (2026-03-22)

  • Smart Startup: pre-loads memories from pending followups + diary

  • Context Packet: structured injection for subagents

  • Auto-Prime: keyword-triggered area learnings in heartbeat

  • Diary Archive: permanent subconscious memory (180d+ auto-archived)

v0.9.0 — Cognitive Memory (2026-03-15)

  • Atkinson-Shiffrin memory model (STM → LTM promotion)

  • Semantic RAG with fastembed (BAAI/bge-base-en-v1.5, 768 dims)

  • Trust scoring, sentiment detection, adaptive personality modes

  • Ebbinghaus decay, sister detection, quarantine system

License

AGPL-3.0 -- see LICENSE


Created by Francisco Cerdà Puigserver & NEXO (Claude Opus) · Built by WAzion

Available Tools

170 tools
nexo_answerB

Answer a pending question from another session.

ParametersJSON Schema
NameRequiredDescriptionDefault
qidYesThe question ID shown in heartbeat output.
answerYesYour response.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only states the action without detailing side effects, authentication requirements, or what happens after answering. For a mutation tool, more behavioral context is needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence (8 words) that immediately conveys the core purpose. No extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the sibling set and lack of annotations or output schema, the description is too minimal. It does not provide enough context about when it should be used or what the outcome is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters. The tool description adds no additional meaning beyond what the schema already provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (answer), resource (pending question), and context (from another session). It distinguishes this from sibling tools like nexo_ask and nexo_check_answer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, such as nexo_check_answer. The description implies a use case but does not provide explicit when/when-not instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_api_callA

Make an authenticated HTTP request to the NEXO Desktop backend (nexo-desktop.com).

The session bearer is auto-loaded from the OS keychain — the agent never sees or handles tokens. Use this for any /api/* endpoint the user has permission for: provider-proxy/, credits/, cards/*, auth/app-tokens, etc.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYespath starting with '/' (e.g. '/api/provider-proxy/call').
methodYesHTTP method (GET / POST / PUT / DELETE / PATCH).
base_urlNooverride default base (default: https://nexo-desktop.com).
body_jsonNoJSON string of the request body. Empty for GET.
headers_jsonNooptional extra headers as a JSON object. Authorization is ignored.
idempotency_keyNoUUID v4 to dedupe POST/PUT retries (avoids double-charge).

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the key behavioral trait that the session bearer is auto-loaded from the OS keychain and the agent never sees tokens. However, it does not mention error handling, rate limits, idempotency behavior, or response format. The high schema coverage mitigates some gaps, but more depth would improve transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences. The first sentence states the core function, and the second adds essential auth context and usage scope. Every word adds value with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers auth well but does not explain what the response looks like (e.g., raw JSON from the API) or error handling. Given no output schema, more completeness about return values would benefit the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3. The description does not add significant meaning beyond the schema; it focuses on the tool's purpose rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it makes authenticated HTTP requests to the NEXO Desktop backend, specifying the scope of endpoints (any /api/* endpoint user has permission for). It implicitly distinguishes from sibling tools by being the generic API call tool, but does not explicitly differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use this tool: 'Use this for any /api/* endpoint the user has permission for'. It does not explicitly list when not to use or alternatives, but the usage context is specific and helpful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_askA

Ask a question to another session (they see it on next heartbeat).

ParametersJSON Schema
NameRequiredDescriptionDefault
to_sidYesTarget session ID.
from_sidYesYour session ID.
questionYesThe question text.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Mentions timing behavior ('next heartbeat') but lacks details on persistence, limits, or whether the target must be active. With no annotations, more behavioral context would help.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that is front-loaded and efficient, with no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with 3 string params and no output schema, the description covers the core purpose and a key behavioral trait. Minor gaps like return value or errors are acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all 3 parameters with descriptions. The description (e.g., 'Your session ID') adds minimal value beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action 'ask a question to another session' and specifies the timing ('on next heartbeat'), distinguishing it from siblings like nexo_send or nexo_answer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies use for asynchronous inter-session communication but gives no explicit when-to-use or alternatives. Among many sibling tools, an agent might need more guidance to choose correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_automation_reconcileC

Build or apply the safe automation reconciliation plan.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description does not disclose side effects, safety implications, or what 'safe automation reconciliation plan' entails. The behavioral impact of building vs applying is unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single front-loaded sentence, no wasted words. Could be expanded slightly without losing conciseness, but it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks explanation of what a reconciliation plan is, return value, and how build/apply work. For a tool with 1 parameter and no output schema, the description is too minimal to be fully helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The boolean parameter 'apply' is implicitly tied to the two actions in the description (build vs apply), but no explicit mapping is given. With 0% schema coverage, the description partially compensates by naming both modes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses verbs 'Build' and 'apply' and a resource 'reconciliation plan', clearly indicating a dual-mode tool. Distinguishes from siblings as no other tool mentions reconciliation, though the concept remains somewhat vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no prerequisites or exclusions. The description does not help the agent decide between build or apply modes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_automation_supervisorB

Read-only supervisor report for automations, cron runs, cron spool and Evolution policy.

ParametersJSON Schema
NameRequiredDescriptionDefault
markdownNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly states 'read-only', which is a key behavioral trait. However, it does not disclose other behaviors such as authentication needs, rate limits, or output format details. The markdown parameter's effect is not explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence with no wasted words. The description is front-loaded with the key verb and subject, effectively communicating the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one optional param, no output schema), and the description covers scope and read-only nature. But it lacks details on return value or behavior when markdown is true/false, which would be helpful for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description should clarify the markdown parameter. It only mentions 'report' but does not explain how the boolean markdown parameter affects the output, leaving the agent uncertain about its semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it's a read-only supervisor report covering automations, cron runs, cron spool, and Evolution policy. This is specific and distinguishes it from siblings like nexo_automation_reconcile, but could be more precise about the report's contents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives (e.g., nexo_automation_reconcile). The description only implies usage via the 'read-only' qualifier, lacking context for when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_capability_explainC

Explain one NEXO product capability with source and safety context.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
localeNoes
capability_idNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavior. It only mentions 'with source and safety context,' hinting at output content but not revealing side effects, read-only nature, or required permissions. The tool likely reads data but that's not confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence (11 words), which is concise but sacrifices necessary detail. It front-loads 'explain' but lacks structure for multiple dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and 0% schema coverage, the description fails to cover essential aspects like return value, parameter usage, or constraints. It is severely incomplete for a tool with three parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description contains no information about the three parameters (capability_id, query, locale). The agent has no guidance on how to populate them, making selection and invocation guesswork.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Explain one NEXO product capability with source and safety context' clearly states a specific verb (explain) and resource (one NEXO product capability), and adds unique context about source and safety. However, it doesn't explicitly differentiate from sibling 'nexo_product_capabilities' which lists capabilities, but the 'one' suggests a single explain action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like nexo_product_capabilities or nexo_tool_explain. There is no indication of prerequisites, typical scenarios, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_card_matchC

Find official NEXO protocol cards for a user request.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
localeNoes
categoryNo
business_typeNo
include_protocolNo

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should fully disclose behavioral traits, but it only gives a generic purpose without mentioning side effects, auth requirements, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is concise but too brief given the tool's parameter count and lack of context; it sacrifices necessary detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, no output schema, and no annotations, the description fails to explain return values or parameter roles, leaving the tool's usage significantly underdefined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, yet the description adds no explanation for any of the six parameters (e.g., query, limit, category). The agent gains no insight beyond the schema definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Find official NEXO protocol cards for a user request,' which clearly identifies the verb and resource but lacks specificity to distinguish from similar sibling tools like nexo_skill_match or nexo_answer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor any exclusions or context for appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_check_answerC

Check if a question has been answered.

ParametersJSON Schema
NameRequiredDescriptionDefault
qidYesThe question ID from nexo_ask.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is the sole source of behavioral info since annotations are absent. It implies a read-only check but does not explicitly state its safety profile or whether it has side effects. Minimal disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence with no redundancy. For a simple tool, this is appropriately sized and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is provided, and the description does not explain what the tool returns (e.g., boolean or status). Given the tool's simplicity, some indication of the return value or behavior would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (qid described as 'The question ID from nexo_ask'). The description adds no additional meaning beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Check if a question has been answered' uses a specific verb and identifies the resource (question via qid). It clearly states the tool's operation but lacks differentiation from sibling tools like nexo_answer and nexo_ask, which might perform related actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. With siblings like nexo_answer and nexo_ask, the description should indicate that this tool is for checking status rather than performing actions, but it does not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_checkpoint_readA

Read the latest session checkpoint. Used by PostCompact hook and for manual recovery.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidNoSession ID. If empty, returns the most recent checkpoint from any session.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It states it's a read operation, indicating non-destructiveness, but does not disclose further behavioral traits like permissions, logging, or side effects. Adequate but minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences, front-loading the core action and adding context without unnecessary words. Every sentence serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with one optional parameter and no output schema, the description is fairly complete but lacks details on return format or structure. Adequate given the context, but could mention what a checkpoint entails.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage for the sole parameter, already describing the meaning and default behavior. The tool description does not add new semantic information beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads the latest session checkpoint, indicating a specific verb and resource. It provides context of usage in PostCompact hook and manual recovery, but does not explicitly distinguish from sibling tools like nexo_checkpoint_save.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage scenarios (PostCompact hook, manual recovery) but lacks explicit guidance on when not to use or alternatives. The schema parameter hint about empty session ID adds some context, but no direct comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_checkpoint_saveA

Save a session checkpoint for auto-compaction continuity.

Call this BEFORE context compaction to preserve session state. The PostCompact hook reads this checkpoint and re-injects it as a Core Memory Block, so the session continues seamlessly.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYesSession ID.
taskNoCurrent task description.
next_stepNoThe concrete next action to take.
task_statusNoOne of 'active', 'investigating', 'fixing', 'deploying', 'blocked'.active
active_filesNoJSON array of file paths currently being worked on.[]
current_goalNoWhat you're trying to achieve right now (1-2 sentences).
errors_foundNoErrors encountered and their status (resolved/open).
reasoning_threadNoYour current chain of thought (1-2 sentences).
decisions_summaryNoRecent decisions with brief reasoning (2-3 lines).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description fully handles transparency. It explains the checkpoint is re-injected by PostCompact hook, revealing the behavioral flow. Does not mention side effects like overwriting, but for a save operation this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each purposeful: first states action, second and third add usage and context. No fluff, well front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description sufficiently explains the tool within its context (auto-compaction). While no output schema is present, the description covers the main behavior and necessary sequence. Could mention return value, but not a major gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so description adds minimal extra parameter meaning beyond schema. However, it provides workflow context that indirectly helps understand parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool saves a session checkpoint for auto-compaction continuity, with a specific verb ('Save'), resource ('session checkpoint'), and distinguishes from sibling tools like nexo_checkpoint_read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call this tool 'BEFORE context compaction', providing clear when-to-use guidance. It does not list alternatives or when-not-to-use, but the context is strong enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_closeC

Close a verified closure item or reject/stale it explicitly.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNocompleted
item_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It only hints at mutation via 'close' and 'reject/stale', but fails to disclose side effects, required item state, permissions, or what happens to the item after. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, efficiently conveying the core action. However, it could be slightly expanded to include parameter hints without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a closure workflow tool with no output schema and multiple siblings, the description lacks state transition info, prerequisites, and outcome details. It is too brief to be fully complete in context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 0% coverage on parameters (item_id, reason). The description does not explain any parameter, such as the role of 'reason' or how to specify 'rejected' or 'stale'. No added meaning beyond schema, which itself provides none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Close a verified closure item or reject/stale it explicitly' clearly states the verb 'close' and the resource 'closure item'. It distinguishes two modes (closing verified or rejecting/staling), which helps differentiate from siblings like nexo_closure_verify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after verification or when needing to reject/stale, but it does not explicitly state when to use versus alternatives (e.g., nexo_closure_verify, nexo_closure_status). No when-not-to-use guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_item_getC

Return one closure item with sources and events.

ParametersJSON Schema
NameRequiredDescriptionDefault
item_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states the return content but does not reveal that this is a read-only operation, whether it has side effects, auth requirements, or error behavior. For a simple get, minimal disclosure but insufficient for full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short sentence with no wasted words. It is front-loaded with the key action. However, it lacks structure (e.g., bullet points) that could aid readability, but given the simplicity, this is acceptable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of many sibling closure tools and no output schema, the description is too terse. It does not explain what a 'closure item' is, how to use the returned data, or how this tool fits into the closure workflow. Agent might need more context to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (description mentions no parameters). The single required parameter 'item_id' is not explained—what it represents, how to obtain it, or format. The tool name implies it's a closure item ID, but the description adds no value beyond the schema's existence.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a single closure item, including its sources and events. This specifies the verb (return), resource (closure item), and scope (one with sources/events), effectively distinguishing it from sibling tools like nexo_closure_close or nexo_closure_status which perform other actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the many sibling closure tools (e.g., nexo_closure_triage, nexo_closure_verify). It does not mention prerequisites, alternatives, or conditions. The context implicitly suggests retrieving an item by ID, but explicit usage guidance is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_nextB

Return the next ranked closure items without executing source actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
areaNo
kindNo
limitNo
stateNo
sourceNo
max_riskNo
include_waitingNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full burden. It discloses that source actions are not executed, which is a key behavioral trait for safety. However, it does not address other aspects like read-only nature, idempotency, or potential side effects from the ranking logic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that immediately states the action. It is front-loaded and each word is meaningful. However, the extreme brevity sacrifices necessary detail for parameter semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, no output schema, and no annotation support, the description is far from complete. It only covers the high-level behavior, leaving the agent without the information needed to correctly invoke the tool with appropriate parameter values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 7 parameters with 0% schema description coverage, and the description provides no information about what each parameter does. The agent has no semantic guidance for using 'limit', 'include_waiting', 'source', 'kind', 'state', 'max_risk', or 'area'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'Return', the resource 'ranked closure items', and includes a unique qualifier 'without executing source actions', distinguishing it from other closure tools like nexo_closure_close or nexo_closure_triage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that this tool is for previewing next closure items without side effects, but provides no explicit guidance on when to use it versus alternatives like nexo_closure_status or nexo_closure_snapshot. The qualifier 'without executing source actions' is a hint but not sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_snapshotC

Write and return an Operational Closure Plane daily snapshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
refreshNo
snapshot_dateNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It mentions 'write and return' but does not explain side effects (e.g., whether previous snapshots are overwritten), authentication requirements, rate limits, or what 'Operational Closure Plane' entails. This lack of detail significantly limits transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence that is concise and not verbose. However, it could be more informative without sacrificing brevity, earning a middle score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and parameter descriptions, the description is too brief to be complete. It does not explain what a snapshot contains, how it is used, or how the tool fits into the broader workflow. More detail is needed for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 3 parameters (refresh, snapshot_date, limit) with defaults but no descriptions. The description provides no additional meaning, leaving the agent to infer their purpose from names alone. This is insufficient for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Write and return an Operational Closure Plane daily snapshot' clearly states the action (write and return) and the resource (daily snapshot), and it distinguishes from sibling tools like nexo_closure_close or nexo_closure_status by focusing on snapshot creation. However, it does not explicitly differentiate from other snapshot tools like nexo_continuity_snapshot_write, so it is not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, context, or situations where it should not be used. The description only states the function without any usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_statusC

Read the Operational Closure Plane status and ranked closure queue.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
refreshNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states 'Read' implying non-destructive operation, but lacks details on side effects, cost, or response characteristics. Behavioral traits beyond reading are absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but not optimally structured. It conveys purpose without waste, but misses opportunities to include additional useful information in a front-loaded manner.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 parameters, no output schema), the description covers purpose but lacks parameter explanations, usage context, and behavioral transparency. It feels incomplete for an agent to use effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the parameters 'refresh' or 'limit' at all. It adds no meaning beyond the schema, failing to compensate for low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads 'Operational Closure Plane status and ranked closure queue', which is a specific verb and resource. It distinguishes from sibling closure tools like nexo_closure_close or nexo_closure_snapshot by focusing on status and queue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as nexo_closure_snapshot or nexo_closure_item_get. The description does not mention prerequisites, exclusions, or context for appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_triageB

Triage a closure item without executing its source action.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
ownerNo
stateNo
item_idYes
next_actionNo
duplicate_ofNo
blocker_reasonNo
capability_statusNo
evidence_requiredNo
capability_requiredNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only states 'Triage a closure item without executing its source action,' but does not explain side effects, required permissions, or what happens to the item (e.g., update, categorization). Given the many optional parameters, the tool likely modifies state, but this is not confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence. No filler words; every part adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, no output schema, no annotations), the description is too minimal. It does not cover what the tool returns, the meaning of 'triage' in this context, or how parameters interact. An AI agent would struggle to use this correctly without additional knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Description does not explain any of the 10 parameters (e.g., state, blocker_reason, next_action). Schema description coverage is 0%, so the AI agent must infer meaning solely from parameter names, which is insufficient for a tool with abstract fields like 'evidence_required' and 'capability_status'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Triage') and target resource ('a closure item'), and distinguishes it by noting it does not execute the source action, which differentiates it from siblings like 'nexo_closure_close'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use: for triage without execution. It does not explicitly list when not to use or provide alternatives, but the context 'without executing its source action' gives sufficient guidance for an AI agent to understand appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_closure_verifyC

Record verification evidence for a closure item.

ParametersJSON Schema
NameRequiredDescriptionDefault
item_idYes
evidenceYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only says 'record evidence,' implying mutation, but does not mention idempotency, permission requirements, side effects on existing evidence, or error handling for invalid item_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no wasted words. However, it could benefit from additional structure or details to improve usefulness without sacrificing brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% schema parameter coverage, the description is insufficient for a mutation tool. It omits information about return values, side effects, and parameter specifics, making it incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not clarify the meaning of 'item_id' or 'evidence' beyond their types. It fails to explain what format evidence should be in or how item_id identifies the closure item.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Record verification evidence for a closure item' uses a specific verb ('record') and resource ('verification evidence for a closure item'), clearly distinguishing it from sibling closure tools like nexo_closure_close or nexo_closure_item_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool vs alternatives such as nexo_evidence_record or other closure tools. There is no mention of prerequisites, when not to use it, or typical usage scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_cognitive_control_observatoryC

Read-only metrics for Local Context, learnings, followups and intraday memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
window_secondsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates it is 'read-only', implying no side effects, but does not explain what the metrics represent, how they are computed, or any rate limits. With no annotations, this is insufficient for an agent to fully understand behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but lacks structure. It front-loads the read-only nature but omits parameter details or output format, making it minimally adequate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one parameter with no explanation, no output schema, and the description only lists domains. An agent would lack critical information on how to invoke it correctly and what to expect in return.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'window_seconds' has no schema description (0% coverage) and is not mentioned in the tool description. Agents have no guidance on its purpose or valid values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides 'Read-only metrics' for specific domains (Local Context, learnings, followups, intraday memory), identifying the resource and action. However, it does not differentiate from siblings like nexo_local_context or nexo_intraday_memory_cycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives, such as nexo_local_context or nexo_memory_health. No exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_confidence_checkC

Decide whether an answer should proceed directly or be verified first.

ParametersJSON Schema
NameRequiredDescriptionDefault
areaNo
goalYes
stakesNo
unknownsNo[]
task_typeNoanswer
constraintsNo[]
context_hintNo
evidence_refsNo[]
verification_stepNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description is the sole source for behavioral traits. It only states a decision-making function without disclosing how the decision is made, what side effects exist, or any return value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that efficiently conveys the core purpose without unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters, no output schema, and no annotations, the description is too brief to be complete. It does not cover inputs, outputs, or behavior beyond a high-level decision.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for 9 parameters, and the description does not mention any parameters or their roles, leaving the agent with no semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb 'Decide' and resource 'whether an answer should proceed directly or be verified first'. However, it does not differentiate from sibling tools, which lack a similar function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidelines are provided about when to use this tool versus alternatives. There is no mention of context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_context_packetA

Build a context packet for subagent injection. Returns learnings + changes + followups + preferences + cognitive memories for a specific area.

MUST call before delegating ANY task to a subagent. Inject the result into the subagent's prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
areaYesProject/area name (e.g., 'ecommerce', 'shopify', 'backend', 'mobile-app', 'nexo', 'infrastructure').
filesNoOptional comma-separated file paths for additional context.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It discloses the return content (learnings, changes, etc.) and the mandatory before-delegation requirement. However, it does not explicitly state whether the tool has side effects or is read-only, leaving some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences cover purpose, output, and usage. Every sentence adds value without redundancy, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema, the description compensates by listing the types of information returned. It provides necessary usage context and aligns with the sibling tool set. However, it omits potential error conditions or variations in output based on parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents parameters well. The description adds context by mentioning 'for a specific area' and linking to usage, but it does not elaborate on how the 'files' parameter modifies the output beyond the schema's description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool builds a context packet for subagent injection and lists the types of information returned (learnings, changes, followups, preferences, cognitive memories). It effectively distinguishes this from sibling tools like nexo_context_router by specifying the exact output and use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs that this tool MUST be called before delegating any task to a subagent and directs how to use the result (inject into the subagent's prompt). This provides precise when-to-use guidance and clear usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_context_routerC

Return compact local context evidence suitable for injection before a reply.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
intentNoanswer
max_charsNo
current_contextNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states the output type ('compact local context evidence') but discloses no behavioral traits such as side effects, permissions required, or how evidence is selected. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 9 words, which is extremely terse. While brevity is good, it omits necessary detail for a tool with 5 parameters and many siblings, making it insufficiently informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is provided, and the description fails to explain what 'local context evidence' entails, how it differs from other context tools, or any constraints. The tool's complexity (5 params, many siblings) requires a richer description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage and 5 parameters (limit, query, intent, max_chars, current_context), the description adds zero information about any parameter. An AI agent cannot infer parameter usage from the description alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it returns 'compact local context evidence suitable for injection before a reply', which gives a clear verb and resource. However, it does not differentiate from sibling tools like nexo_local_context, nexo_context_packet, or nexo_recent_context, which likely serve similar purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. Lacks any context for selection, no when-not-to-use, and no mention of prerequisites or scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_continuity_auditC

Return the forensic continuity timeline for a conversation.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
conversation_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavioral traits, but it only describes the output ('timeline') without mentioning read-only nature, side effects, authorization needs, or error behavior. The short description fails to provide transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (one sentence, 7 words). While there is no wasted text, the brevity sacrifices clarity and completeness. It is appropriately front-loaded but under-informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's apparent complexity (forensic timeline), the lack of annotations, output schema, and parameter descriptions makes the description woefully incomplete. It fails to explain return format, edge cases, or usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no meaning to the parameters. It does not explain what 'conversation_id' is, expected format, or how 'limit' affects results. The description offers no semantic value beyond the parameter names themselves.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return') and the resource ('forensic continuity timeline for a conversation'). It effectively communicates the tool's primary function, distinguishing it from siblings like snapshot read/write and compaction events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus its siblings (e.g., nexo_continuity_snapshot_read, nexo_continuity_resume_bundle). The description lacks any context about prerequisites, alternatives, or conditions for use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_continuity_compaction_eventC

Persist a compaction-related continuity event into the canonical snapshot stream.

ParametersJSON Schema
NameRequiredDescriptionDefault
payloadNo
trace_idNo
event_typeNopost_compact
session_idNo
conversation_idYes

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It indicates a persist (write) operation but omits side effects, idempotency, conflict behavior, or required permissions. There is no contradiction with annotations because none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but at the cost of providing necessary detail. It front-loads the verb but lacks structure to aid quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, 0% schema coverage, and a complex tool with 5 parameters and many siblings, the description is severely incomplete. It fails to explain what a compaction event is, how to construct the event, or what the result of persisting it entails.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 5 parameters with 0% coverage (no descriptions). The description does not explain any parameter's meaning or usage, such as what 'payload' or 'event_type' should contain. It adds no value beyond the parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a clear verb ('Persist') and resource ('compaction-related continuity event'), and indicates the target is the 'canonical snapshot stream'. This distinguishes it from sibling tools like nexo_continuity_audit or nexo_continuity_snapshot_write by emphasizing compaction, though it relies on domain jargon.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus other continuity or snapshot tools. There is no mention of prerequisites, alternatives, or when not to use it, leaving the agent without decision context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_continuity_resume_bundleC

Build the small continuity bundle Desktop injects after restore or stale resume loss.

ParametersJSON Schema
NameRequiredDescriptionDefault
clientNo
session_idNo
token_budgetNo
conversation_idNo
external_session_idNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations available, the description carries the full burden of disclosing behavioral traits. It fails to mention any side effects, permissions required, or output characteristics, leaving the agent uncertain about the tool's impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it omits critical details that would make it more effective. It earns its place but could be restructured to include essential parameter or behavioral information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five undocumented parameters, no output schema, and no annotations, the description is far too sparse. An AI agent lacks sufficient context to understand what the tool produces or how to supply the parameters correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has five parameters, none documented in the schema (0% coverage), and the description provides no explanations for any of them. Parameters like 'client', 'session_id', and 'token_budget' are left entirely ambiguous, making correct invocation difficult.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a clear action ('build') and resource ('small continuity bundle') with a well-defined context ('after restore or stale resume loss'). It distinguishes itself among many sibling tools by focusing on a specific small bundle for Desktop injection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after restore or stale resume loss, offering some use-case context. However, it provides no explicit when-not-to-use guidance or alternative tool recommendations, which are essential for an AI agent deciding among many similar tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_continuity_snapshot_readB

Read recent continuity snapshots by conversation_id or session_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
session_idNo
conversation_idNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It states 'Read' implying read-only, but does not mention permissions, rate limits, or any side effects. This is insufficient for a tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front- loads the action and resource. It is efficient but could include additional useful details without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description should explain what is returned (e.g., fields, format, pagination). It only mentions the filtering input, leaving the return value unspecified. This is incomplete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description adds some meaning by naming 'conversation_id' and 'session_id' as filters. However, it does not explain the 'limit' parameter or default behaviors, leaving gaps for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and resource 'recent continuity snapshots' with filtering criteria by conversation_id or session_id, distinguishing it from sibling tools like nexo_continuity_snapshot_write.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as nexo_continuity_resume_bundle or nexo_continuity_audit. The description does not include when-not-to-use or context-specific scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_continuity_snapshot_writeC

Write a durable continuity snapshot for Desktop/Brain handoff.

ParametersJSON Schema
NameRequiredDescriptionDefault
clientNo
payloadNo
trace_idNo
event_typeNoturn_end
session_idNo
conversation_idYes
idempotency_keyNo
external_session_idNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. The description only mentions 'durable' which implies persistence, but fails to disclose side effects, idempotency, failure behavior, or permissions. Vague for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no wasted words. However, it could be improved by adding structure (e.g., bullet points) for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters, no output schema, and no annotations, the description is severely incomplete. It does not explain what a continuity snapshot contains, how handoff works, or what the return value is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description adds no information about any of the 8 parameters. It does not explain the meaning of 'conversation_id' or any other field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Write', resource 'durable continuity snapshot', and purpose 'for Desktop/Brain handoff'. It distinguishes from sibling 'nexo_continuity_snapshot_read' by specifying write operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context via 'Desktop/Brain handoff' but provides no explicit guidance on when to use versus alternatives, nor any conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_cortex_checkC

Cognitive pre-action check. Call before significant actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYes
planNo[]
unknownsNo[]
task_typeNoanswer
constraintsNo[]
known_factsNo[]
evidence_refsNo[]
verification_stepNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, and the description does not disclose any behavioral traits (e.g., read-only, destructive, required permissions). The agent cannot infer side effects or safety from the description alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (two clauses) but fails to pack useful information. It is under-informative rather than concise, making it insufficient for a tool with many parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (8 parameters, no output schema, many siblings), the description lacks essential details about behavior, return values, and parameter usage, making it severely incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 8 parameters and 0% schema description coverage, the description adds no value by explaining what parameters like 'goal', 'plan', 'unknowns' mean or how to use them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it is a 'Cognitive pre-action check' and should be called 'before significant actions,' giving a general purpose but lacking specificity on what exactly is checked or how it differs from similar sibling tools like nexo_confidence_check or nexo_guard_check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Call before significant actions' provides a usage hint, but it does not clarify when not to use it or suggest alternatives, leaving the agent with limited guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_cortex_decideC

Evaluate concrete alternatives for a high-impact task and persist the recommendation.

ParametersJSON Schema
NameRequiredDescriptionDefault
areaNo
goalYes
goal_idNo
task_idNo
task_typeNoexecute
session_idNo
constraintsNo[]
alternativesYes
context_hintNo
impact_levelNohigh
evidence_refsNo[]
goal_profile_idNo
linked_outcome_idNo
auto_create_outcomeNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behaviors. It mentions persistence ('persist the recommendation') but omits side effects, permissions, idempotency, or what happens on repeated calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 12 words, which is efficient but insufficient for a tool with 14 parameters. It is front-loaded but lacks necessary detail to earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 14 parameters, no annotations, no output schema, and no parameter descriptions, the description is severely incomplete. It does not provide enough context for the agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description only adds meaning for 'goal' and 'alternatives' implicitly. The other 12 parameters (task_type, impact_level, etc.) are not explained, leaving the agent unable to use them correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates alternatives and persists recommendations, with a specific verb ('Evaluate') and resource ('concrete alternatives'). It distinguishes from siblings like nexo_closure_* or nexo_goal_* which handle different tasks, but lacks explicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like nexo_cortex_check or nexo_goal_open. The description implies high-impact tasks but doesn't specify context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_create_app_tokenA

Create a persistent AppToken for the current user via POST /api/auth/app-tokens.

Use this when a card needs to mint a token that will live inside a snippet the user pastes on their own website (chatbot widget, embed, public API autoresponder). The plain-text token is returned ONCE — embed it in the generated snippet and never store it elsewhere.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYeshuman label for the token (e.g. 'chatbot-mitienda-com').
abilitiesNocomma-separated abilities. Allowed: provider-proxy:call, provider-proxy:estimate, credits:read. Defaults to 'provider-proxy:call' if empty.
expires_atNoISO 8601 future date, empty for non-expiring token.
allowed_platformsNocomma-separated platform keys (openai, anthropic, gemini, ...). Empty = all platforms the user has access to.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses critical behavior: 'The plain-text token is returned ONCE — embed it in the generated snippet and never store it elsewhere.' This warns about the token's one-time availability and storage instruction, which is essential for correct usage. Could be improved by mentioning permissions or idempotency, but current info is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise: two sentences in the main description (one for action, one for usage) plus the one-time return warning. Every sentence adds value. No fluff. Front-loaded with the API endpoint and purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple creation tool with 4 parameters (all documented in schema), the description covers the purpose, usage context, and a critical behavior (one-time return). It lacks output format details, but no output schema exists. Could mention error cases, but the description is reasonably complete for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (every parameter described in schema). Description adds minimal extra value beyond the schema, e.g., it does not further explain 'abilities' or 'expires_at' beyond what's in the schema. Baseline of 3 applies because schema already does the heavy lifting; description doesn't add significant new semantic info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states: 'Create a persistent AppToken for the current user via POST /api/auth/app-tokens.' It specifies the verb (create), resource (AppToken), and the HTTP endpoint. The use case (embedding in a user's website snippet) further clarifies its purpose, distinguishing it from other tools in the large sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description gives explicit context: 'Use this when a card needs to mint a token that will live inside a snippet the user pastes on their own website.' This provides clear usage guidance. It does not explicitly state when not to use or name alternatives, but the context is sufficient for an AI agent to select this tool appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_credential_createC

Store a new credential.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesKey name (e.g., api_key, token, ssh).
notesNoDescription or context (optional).
valueYesThe secret value.
serviceYesService name (e.g., google-ads, cloudflare).

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description minimally indicates a write operation ('store') but does not disclose important behavior such as whether existing credentials with the same service/key are overwritten, error handling, or response format. With no annotations, the description bears the full burden of transparency and falls short.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (4 words) but lacks structure. It is front-loaded but too brief to be optimally useful; slightly longer description could improve clarity without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description should provide more context about the credential creation process, such as security implications, idempotency, or integration with other credential tools. It is incomplete for a sensitive write operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are well-documented in the schema. The description adds no additional meaning beyond the schema descriptions, thus meeting the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Store a new credential', specifying a verb and resource. It distinguishes itself from sibling tools like delete, get, list, and update, though the differentiation is implicit via the action 'create'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, no prerequisites or conditions are mentioned, and there is no advice on avoiding misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_credential_deleteA

Delete credential(s). If no key, deletes all for the service.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoSpecific key (optional). If empty, deletes ALL for service.
serviceYesService name.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must carry the burden. It mentions destructive behavior and the conditional bulk delete, but lacks details on consequences like irreversibility, permissions, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences only, front-loaded with the action and key behavior. No unnecessary words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool with 2 parameters and no output schema, the description adequately covers the main behavior and conditional logic. Slightly more context about permanence would be ideal, but still complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and parameter descriptions already explain the key behavior. The tool description repeats this information without adding new meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Delete), resource (credential(s)), and scope (by service, optionally by key). It distinguishes from siblings like nexo_credential_create, get, list, update by focusing on deletion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies deletion context but does not explicitly state when to use this tool versus alternatives like update or create. No guidance on prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_credential_getB

Get credential value(s) for a service.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoSpecific key (optional). If empty, returns all for the service.
serviceYesService name (e.g., google-ads, meta-ads, shopify).

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description carries full burden. It only states 'Get,' which implies a read operation, but does not disclose behavioral traits such as idempotency, whether credentials are cached, authentication requirements, or error behavior (e.g., missing service or key). Insufficient for safe agent use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise: one sentence with no superfluous words. All content is directly relevant despite limited detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema is absent; the description does not explain return values (e.g., credential values format, behavior when key is empty vs specified). For a retrieval tool, this leaves significant gaps in agent understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%; both parameters (service, key) are described in the schema. The description adds no additional meaning beyond what the schema provides. Baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb-object structure: 'Get credential value(s) for a service.' The description specifically identifies the resource (credentials) and distinguishes this tool from sibling CRUD tools (e.g., nexo_credential_create, nexo_credential_delete).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Minimal guidance. The description states the function but does not explicitly indicate when to use this tool versus alternatives like nexo_credential_list (which might list credentials without retrieving values) or when read access is appropriate. Usage is implied by the name but lacks explicit context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_credential_listA

List credentials (names and notes only, no values).

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceNoFilter by service (optional). If empty, shows all.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description clearly states that it returns only names and notes, not values, which is a key behavioral trait beyond the schema. With no annotations, this provides important context. However, it could mention other traits like pagination or rate limits, though not essential for a simple list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one short sentence, front-loaded with the verb and resource, and includes the critical limitation on values. Every word adds value with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description does not specify the return format (e.g., list of objects, fields). While it states names and notes only, an agent might benefit from knowing the structure. Still, it is adequate for a straightforward list tool with well-named sibling functions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the description adds no extra meaning to the 'service' parameter beyond what the schema already provides. The description's mention of 'names and notes only' relates to output, not parameter semantics. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the verb 'list', the resource 'credentials', and the scope: 'names and notes only, no values'. It effectively distinguishes from sibling credential tools (create, delete, get, update) that handle values or mutations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for listing credential metadata, but does not explicitly state when to use this tool vs alternatives like nexo_credential_get. No guidance on when not to use or which tool to choose for retrieving full credentials.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_credential_updateC

Update a credential's value and/or notes.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesKey name.
notesNoNew notes (optional).
valueNoNew value (optional).
serviceYesService name.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must cover behavior. It only says 'Update', which implies mutation, but doesn't disclose side effects, permissions, idempotency, or what happens with default values. Insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence of 8 words, very concise. It communicates the core action efficiently, though it could benefit from slightly more detail without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description fails to explain what the tool returns (e.g., updated credential, success message). For a simple update tool, this is a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already describes all parameters. The description adds no extra meaning beyond stating 'value and/or notes', which is already in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool updates a credential's value and/or notes, matching the verb 'update' and resource 'credential'. It is specific but doesn't explicitly differentiate from sibling tools like create or delete, though the name helps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., create, delete, get, list). No prerequisites, context, or exclusions provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_drive_actA

Mark a drive signal as investigated with an outcome.

Call this after NEXO has autonomously investigated a READY signal.

ParametersJSON Schema
NameRequiredDescriptionDefault
outcomeYesWhat was found during investigation.
signal_idYesSignal ID that was investigated.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It adds behavioral context by indicating the tool is a follow-up to NEXO's autonomous investigation. However, it does not disclose potential side effects, required permissions, or what happens if called out of sequence (e.g., if the signal is not in READY state).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences: the first states the purpose, and the second provides usage context. Every sentence contributes meaning, and the information is front-loaded. No redundant or vague phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple nature of the tool (recording an outcome) and the rich schema, the description is largely complete. It explains when to use the tool and what it does. It could potentially mention that the signal must be in 'READY' state or that the outcome string format is free-form, but overall it's sufficient for an agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with clear descriptions for both parameters ('Signal ID that was investigated' and 'What was found during investigation'). The description adds no further semantic value beyond what the schema already provides. With high coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Mark a drive signal as investigated') and the resource ('drive signal') with a verb+resource structure. It also distinguishes itself by specifying the context: 'Call this after NEXO has autonomously investigated a READY signal,' which differentiates it from siblings like nexo_drive_dismiss or nexo_drive_reinforce.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Call this after NEXO has autonomously investigated a READY signal.' It implies that this tool should not be called before investigation or for non-READY signals. However, it does not explicitly mention alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_drive_dismissA

Dismiss a drive signal (archived, not deleted).

Call this when a signal is not worth investigating.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesWhy this signal was dismissed.
signal_idYesSignal ID to dismiss.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description clarifies that the tool archives rather than deletes, which is helpful. However, with no annotations, it fails to disclose other behavioral traits like reversibility, permissions, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and an immediate usage guideline. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two required params and no output schema, the description covers the essential aspects: what it does and when to use it. It lacks info about return value, but that is not critical given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both parameters described adequately in the input schema. The description adds no additional meaning beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Dismiss' and the resource 'drive signal', and adds the crucial clarification that it archives rather than deletes. This is specific enough to distinguish it from potential deletion tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use the tool ('when a signal is not worth investigating'), but does not mention when not to use it or alternatives (e.g., nexo_drive_act or nexo_drive_reinforce).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_drive_reinforceA

Reinforce a drive signal with a new observation.

Increases tension and may promote the signal status (latent → rising → ready).

ParametersJSON Schema
NameRequiredDescriptionDefault
signal_idYesSignal ID to reinforce.
observationYesNew observation that supports this signal.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses 'increases tension' and 'may promote signal status', which are behavioral effects. However, it lacks details on what 'tension' means, potential side effects, or what happens on failure. The transparency is adequate but not thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two concise sentences. The first provides the core action and resource, while the second explains the effect. No redundant information, and the purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two simple parameters and no output schema. The description explains the purpose and effect but does not mention what the tool returns or how success/failure is indicated. For a mutation tool, return info is useful for agents. This gap lowers completeness from ideal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the parameter descriptions already explain the fields. The description's mention of 'New observation that supports this signal' and 'Signal ID to reinforce' adds no new semantic value beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reinforces a drive signal with a new observation and explains the effect on signal status (latent → rising → ready). It distinguishes itself from sibling drive tools like 'nexo_drive_act' and 'nexo_drive_dismiss' by specifying the reinforcement action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use (when you have a new observation to strengthen a signal) but provides no explicit guidance on when not to use it or contrasts with alternatives. No sibling differentiation is present, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_drive_signalsA

List autonomous drive/curiosity signals.

Drive signals are observations NEXO accumulates during normal work. When tension crosses threshold, NEXO investigates silently.

ParametersJSON Schema
NameRequiredDescriptionDefault
areaNoFilter by operational area (shopify, google-ads, wazion, nexo, etc.).
limitNoMax signals to return (default 20).
statusNoFilter by status (latent, rising, ready, acted, dismissed). Default: active only.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description gives context about drive signals and tension, but does not disclose whether listing is read-only, any side effects, or required permissions. With no annotations, more behavioral detail would be beneficial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two short sentences that front-load the key action and then provide context. No extraneous words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description explains the concept of drive signals, it does not mention the output format or what the returned list contains. Given the absence of an output schema, a richer description would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides descriptions for all three parameters (status, area, limit). The description does not add further meaning to the parameters, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the verb 'List' and the resource 'autonomous drive/curiosity signals'. It also explains the nature of drive signals, which helps distinguish from action-oriented siblings like nexo_drive_act.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like nexo_drive_act or nexo_drive_dismiss. It does not specify conditions for listing signals.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_embedding_migration_statusA

Return read-only status for the cognitive embedding migration.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates the tool is 'read-only', implying no side effects, but doesn't disclose other behavioral traits such as possible error states or return format. With no annotations, the description carries full burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence with no wasted words. It is front-loaded and efficiently communicates the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (no parameters, no output schema), the description is minimally adequate but lacks details on what the status values are or how to interpret the response. More context would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters, so schema coverage is 100%. The description adds no parameter information, which is appropriate. Baseline score of 3 is applied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return') and the resource ('read-only status for the cognitive embedding migration'). This distinguishes it from sibling tools, as no other tool mentions embedding migration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool or alternatives. It doesn't specify prerequisites or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_entity_dossierC

Build a full local dossier for one entity with aggregates and evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
max_charsNo
max_factsNo
max_assetsNo
max_chunksNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It says 'full local dossier' suggesting local data aggregation, but it does not mention side effects, performance implications, or whether the operation is read-only. This is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

While one sentence is brief, it is under-specified for a tool with 5 parameters and no schema descriptions. Conciseness should not sacrifice necessary information; here it fails to provide any parameter guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and complex parameters, the description is far from complete. It does not define what a dossier includes, how to interpret results, or how the parameters affect the output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% coverage (no parameter descriptions). The description only mentions 'aggregates and evidence' but does not explain any of the five parameters (query, max_assets, max_chunks, max_facts, max_chars). The agent cannot infer how to use these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Build a full local dossier for one entity with aggregates and evidence' clearly states the action (build), resource (dossier), scope (one entity), and content (aggregates and evidence). It distinguishes from sibling tools like nexo_evidence_search or nexo_memory_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for compiling a comprehensive view of a single entity, but it does not explicitly state when to use it versus alternatives, nor does it provide prerequisites or scenarios. For a tool with many siblings, this is a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_evidence_recordC

Record a compact evidence pointer without storing raw command output.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNorecorded
summaryYes
source_idNo
object_refNo
session_idNo
object_typeNoartifact
source_typeNoevidence_ledger
verificationNo
conversation_idNo
idempotency_keyNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral transparency burden. It only says 'without storing raw command output', omitting other important aspects like idempotency, overwrite behavior, or authorization needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. However, it is overly terse and could benefit from additional context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 10 parameters, no output schema, and no annotations, the description is severely incomplete. It fails to explain the core concept of 'compact evidence pointer' and leaves the agent without guidance on how to populate the fields correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no information about the 10 parameters. It does not explain what 'summary', 'source_id', 'object_ref', etc., mean or how they relate to the evidence pointer.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: recording a compact evidence pointer without storing raw output. It uses specific verb 'record' and noun 'evidence pointer', and contrasts with storing raw output, distinguishing it from likely siblings like evidence search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you need to record a reference to evidence without full data, but it does not provide explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_filesB

Show all tracked files across all active sessions with conflict detection.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must convey behavioral traits. Mentions 'conflict detection' but does not explain what that entails, nor whether the operation is read-only or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no extraneous words. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description is minimally adequate. However, it leaves 'conflict detection' undefined and does not clarify the scope or nature of the returned data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, so schema coverage is 100%. Description adds no parameter details, but baseline is 4 for zero-parameter tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool shows all tracked files across active sessions with conflict detection. It distinguishes from many sibling tools by its specific verb and resource, though not explicitly differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, such as nexo_files vs nexo_context_packet or nexo_memory_search. Missing usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_completeB

Mark a followup as completed. Appends result to verification field.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFollowup ID (e.g., NF45).
resultNoWhat was found/done (optional).

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It mentions marking as completed and appending the result, but does not state prerequisites (e.g., followup exists and is incomplete), reversibility, error conditions, or permissions. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no wasted words. The action and key behavior are front-loaded. Every sentence serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and a mutation operation, the description is incomplete. It lacks information about return values, error behavior, state assumptions, and prerequisites. For an agent to use it safely, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters (id and result) with 100% coverage. The description adds value by stating that the result is appended to a verification field, which provides context for its usage. However, additional details like format or constraints are missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (mark a followup as completed) and the side effect (appends result to verification field). However, it does not explicitly differentiate from sibling tools like nexo_followup_update or nexo_followup_note, though the verb 'complete' is specific enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as nexo_followup_update or nexo_followup_delete. The description lacks any context for when it is appropriate to complete a followup.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_createB

Create a new agent followup (autonomous task).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesUnique ID starting with 'NF' (e.g., NF-MCP2).
dateNoTarget date YYYY-MM-DD (optional).
ownerNo'user' | 'waiting' | 'agent' | 'shared'. Leave empty for auto-classification.
internalNo'1'/'true' hides from default user views (agent bookkeeping, protocol, audit). Leave empty to auto-classify by ID prefix.
priorityNocritical, high, medium, low (default: medium).medium
exceptionNoReason this followup should be allowed even under an active autonomy mandate (NF-DS-45569A27). Valid only for the three pre-approved cases: >1GB download, credential the operator must physically enter, or a presence-dependent session with María/Nora.
reasoningNoWHY this followup exists — what decision/context led to it (optional).
recurrenceNoAuto-regenerate pattern (optional). Formats: 'weekly:monday', 'monthly:1', 'monthly:15', 'quarterly'. When completed, a new followup is auto-created with the next date. The completed one is archived with date suffix.
descriptionYesWhat to verify/do.
verificationNoHow to verify completion (optional).

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits such as side effects, permissions, or idempotency. The term 'autonomous task' hints at background execution but lacks detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (8 words) and front-loaded with the verb and resource. However, it may be too terse but contains no unnecessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters and no output schema, the brief description lacks completeness. It does not explain what 'autonomous task' entails or provide context for optional fields like recurrence and exception.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Since schema description coverage is 100%, the baseline is 3. The description does not add additional meaning beyond the schema, but this is acceptable given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Create' and the resource 'agent followup (autonomous task)', distinguishing it from sibling tools like nexo_followup_complete, nexo_followup_delete, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use (to create new followups) but does not provide explicit guidance on when not to use or alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_deleteA

Soft-delete a followup.

IMPORTANT: call nexo_followup_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFollowup ID (e.g., NF45).
read_tokenNoToken returned by `nexo_followup_get`.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must cover behavioral traits. It states 'soft-delete' but does not elaborate on what that entails—whether reversible, hidden, or affecting other data. The agent lacks detail on side effects or resumption.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences: the first nails the purpose, the second delivers crucial usage instruction. No filler, front-loaded. Ideal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is minimally viable for invoking the tool. It instructs on the prerequisite but omits behavioral outcomes (e.g., if deletion is successful, what the response looks like). Adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both parameters described. The description adds no new semantic details beyond the schema, but the IMPORTANT note reinforces the read_token requirement. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Soft-delete a followup' with a specific verb and resource. The term 'soft-delete' distinguishes it from other followup sibling tools like complete or restore, though it could be more explicit about the reversibility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit prerequisite: 'call nexo_followup_get first and pass its READ_TOKEN.' This gives clear sequential guidance, though it does not exclude alternative tools or scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_getA

Read a followup with its history and usage rules.

IMPORTANT: before update/delete/restore/note, call this tool first and use the returned READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description discloses that the tool returns not only the followup data but also history and usage rules, and emphasizes the need for a READ_TOKEN for mutations. This is helpful, though it does not detail token expiration or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise: two sentences plus an important note. Every sentence adds value, and the critical usage instruction is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple parameter set and no output schema, the description covers the essential functional context and usage pattern. It could be improved by hinting at what the response includes beyond the READ_TOKEN.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter 'id' with no description. The tool description does not elaborate on the meaning or format of 'id', leaving it ambiguous despite 0% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads a followup including its history and usage rules. It distinguishes from sibling mutation tools by specifying it is a prerequisite before update/delete/restore/note.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to use the tool: before update/delete/restore/note operations, and that the returned READ_TOKEN is needed. This provides strong usage context and no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_lifecycleC

Return followups grouped by lifecycle lane for runner, dashboard and startup parity.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but does not disclose behavioral traits such as whether the tool is read-only, destructive, or how it handles pagination or limits. It only states what it returns, lacking important context for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but lacks structure. It could benefit from additional sentences to provide context and parameter details without being overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (1 optional param, no output schema, no annotations), the description is incomplete. It does not explain what 'lifecycle lane' means or how the grouping is structured, missing critical context for proper tool usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% parameter description coverage, and the tool description does not explain the 'limit' parameter. The description adds no value beyond the schema, leaving the agent without guidance on how to use the parameter effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns followups grouped by lifecycle lane, using a specific verb and resource. However, it does not differentiate from sibling followup tools like nexo_followup_get or nexo_followup_complete, and the phrase 'for runner, dashboard and startup parity' is vague, reducing clarity slightly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus other followup tools. The phrase 'for runner, dashboard and startup parity' hints at use cases but is not explicit, and there is no mention of when not to use or alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_noteA

Append a note to followup history.

IMPORTANT: call nexo_followup_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFollowup ID (e.g., NF45).
noteYesOperational note to append to history.
actorNoActor label for the history note.nexo
read_tokenNoToken returned by `nexo_followup_get`.

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description should fully disclose behavioral traits. It only states the append action and the token requirement. It does not mention permissions, idempotency, whether the operation is reversible, or any side effects on the followup history.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with a line break, very concise. The important usage note is immediately highlighted. Every sentence is necessary and efficiently communicates the purpose and a key prerequisite.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple nature of appending a note and the 100% schema coverage, the description is adequate. However, it lacks context about how the note is incorporated (e.g., appended to a sequential list) and the expected return value (no output schema). It could be more complete by explaining the effect on history.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already well-documented. The description adds critical semantic context for the 'read_token' parameter by specifying it must come from nexo_followup_get. This is valuable beyond the schema's description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Append a note to followup history' uses a specific verb ('append') and identifies the resource ('followup history'). It clearly distinguishes from sibling tools like nexo_followup_create and nexo_followup_update by specifying an append action rather than create or update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the prerequisite: 'call nexo_followup_get first and pass its READ_TOKEN.' This guides the agent on the required prior step. However, it does not explicitly mention when not to use the tool or provide alternatives for different use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_restoreA

Restore a soft-deleted followup back to PENDING.

IMPORTANT: call nexo_followup_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFollowup ID (e.g., NF45).
read_tokenNoToken returned by `nexo_followup_get`.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It states the tool restores to PENDING (implying mutation) but doesn't disclose side effects, error conditions, or behavior if already PENDING. Adequate for a simple operation but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences. First describes the action, second provides critical prerequisite. No waste, front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and lack of output schema, the description covers the essential behavior and prerequisite. Could mention what happens if read_token is missing or invalid, but overall sufficient for the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already describes both parameters (id, read_token) with 100% coverage. Description does not add extra meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action: restore a soft-deleted followup to PENDING. The verb 'restore' and resource 'followup' are specific. Distinguishes from siblings like nexo_followup_delete (which deletes) and nexo_followup_create (creates new).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call nexo_followup_get first and pass its READ_TOKEN, providing a clear prerequisite. Does not exclude other scenarios but the guidance is strong and context-aware.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_followup_updateA

Update fields of an existing followup. Only non-empty fields are changed.

IMPORTANT: call nexo_followup_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFollowup ID (e.g., NF45).
dateNoNew date YYYY-MM-DD (optional).
ownerNoNew 'user'|'waiting'|'agent'|'shared' (optional).
statusNoNew status (optional).
internalNo'1'/'0' to re-classify visibility (optional).
priorityNocritical, high, medium, low (optional).
read_tokenNoToken returned by `nexo_followup_get`.
descriptionNoNew description (optional).
verificationNoNew verification text (optional).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that only non-empty fields are changed and requires a read_token, but lacks details on side effects, error handling, or what happens if the ID is invalid. Adequate but not thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with two sentences and a clear imperative for the prerequisite. It is front-loaded and every sentence adds value, though it could be slightly more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 parameters, no output schema, no annotations), the description covers the core update behavior and a critical prerequisite. However, it lacks information about return values, validation, and error conditions, making it minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds a general behavioral note ('Only non-empty fields are changed') but does not add per-parameter meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool updates an existing followup and specifies that only non-empty fields are changed. The verb 'Update' and resource 'followup' are explicit, and among sibling tools (e.g., create, delete, get), its role is well-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an important prerequisite: calling nexo_followup_get first and passing its read_token. However, it does not explain when to prefer update over alternatives like create or delete, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_goal_getB

Read one durable goal and optionally include linked workflow runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
goal_idYes
include_runsNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It states the tool reads a goal, but omits details about authentication needs, error responses, idempotency, or what 'include linked workflow runs' entails (e.g., performance impact).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence of 11 words, front-loaded with the purpose, with no superfluous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with two parameters, the description is minimally adequate. However, it does not describe output structure or error conditions, which would be helpful given the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clarifies that include_runs controls whether linked workflow runs are included, but does not elaborate on goal_id beyond the implicit 'one durable goal'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and the resource 'one durable goal', and distinguishes from sibling tools like nexo_goal_list and nexo_goal_update by implying a single, read-only operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like nexo_goal_list or nexo_goal_open. The description only states what it does, not when it is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_goal_listC

List durable goals.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNo
include_closedNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only says 'list', which implies a read-only operation, but it fails to disclose pagination, filtering behavior, rate limits, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (3 words), which is concise but undersized for a tool with 3 parameters. It sacrifices necessary details for brevity, making it less helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, no output schema, and no parameter descriptions, the description is severely incomplete. The agent receives no information about return values, filtering behavior, or what constitutes a 'durable goal', leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the input schema lacks parameter descriptions. The tool description does not explain the meaning of 'limit', 'status', or 'include_closed', relying solely on parameter names which are only partially self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action 'list' and the resource 'durable goals', distinguishing it from siblings like 'nexo_goal_get' (single goal) and 'nexo_goal_update' (modify). However, it lacks specificity on scope or filtering, making it only marginally clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor does it mention any prerequisites or exclusions. The usage context is only implied by the verb 'list'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_goal_openC

Open a durable goal so objectives survive sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
ownerNo
titleYes
priorityNonormal
objectiveNo
next_actionNo
shared_stateNo{}
parent_goal_idNo
success_signalNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose side effects, idempotency, or state changes. It only states 'open a durable goal' without explaining what happens to existing goals, required permissions, or return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence of 9 words is concise but under-specifies. It lacks structure and omits critical information given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having 9 parameters and no output schema or annotations, the description provides only a minimal purpose statement. It fails to cover return values, behavior, or parameter details, making it inadequate for safe invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 9 parameters with 0% description coverage, and the tool description does not explain any parameter meaning. It adds no value beyond the raw schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool opens a durable goal for persisting objectives across sessions, clearly distinguishing it from sibling tools like nexo_goal_get or nexo_goal_list. However, it could be more specific about what 'open' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like nexo_goal_update or nexo_task_open. The description implies use for creating persistent goals but does not discuss prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_goal_updateC

Update a durable goal with active/blocked/completed state.

ParametersJSON Schema
NameRequiredDescriptionDefault
ownerNo
titleNo
statusNo
goal_idYes
objectiveNo
next_actionNo
shared_stateNo
blocker_reasonNo
parent_goal_idNo
success_signalNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It only states 'Update' (mutation) and mentions states, but lacks details on side effects, permissions, or response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence is very concise, but it could be slightly restructured to include more key details in a compact form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters, no output schema, and no annotations, the description is far too brief. It doesn't specify partial update behavior, field interactions, or return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for 10 parameters. The description only adds meaning to the 'status' parameter (states), leaving all other parameters (owner, title, objective, etc.) unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool updates a durable goal and specifies three possible states (active/blocked/completed), distinguishing it from sibling tools like nexo_goal_get, nexo_goal_list, and nexo_goal_open.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor does it specify prerequisites or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_guard_checkC

Check learnings relevant to files/area before reading or editing code.

ParametersJSON Schema
NameRequiredDescriptionDefault
areaNo
filesNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states 'check learnings' without indicating whether it modifies state, what the expected side effects are, or how results are presented. Minimal disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence of 9 words is efficient and front-loaded, but risks being too terse for complex understanding. No unnecessary words, but structure could be improved with more detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given two parameters, no output schema, and no annotations, the description is under-specified. It fails to explain what 'learnings' are, how the result is returned, or any constraints, leaving the agent with significant ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain parameters. It mentions 'files/area' loosely mapping to the two parameters, but does not define their format, allowed values, or behavior. The description adds some meaning but is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks learnings relevant to files or area, with a specific purpose of preparing for code reading/editing. It differentiates from siblings like nexo_cortex_check and nexo_checkpoint_read by focusing on 'learnings' and a pre-action context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use before reading or editing code, but does not provide explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives among the many nexo tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_guardian_rule_overrideA

Temporarily override a Guardian rule's mode (Plan Consolidado 0.17).

Writes to ~/.nexo/config/guardian-runtime-overrides.json which guardian_config.rule_mode already honours at read time. Useful when a rule is noisy during an incident and needs to drop to shadow for an hour without a server restart.

ParametersJSON Schema
NameRequiredDescriptionDefault
ttlNoWindow for the override. One of ``1h`` / ``24h`` / ``session`` (session = 12 h best-effort cap). Default ``24h``.24h
modeNoOne of ``off`` / ``shadow`` / ``soft`` / ``hard``. Pass empty string together with a rule_id to clear an existing override. Core rules (R13/R14/R16/R25/R30) reject ``off``.
rule_idYesThe full rule identifier (e.g. ``R13_pre_edit_guard``).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full transparency burden. It discloses that the tool writes to a specific JSON file and that the override is temporary (no server restart needed). It does not cover all potential side effects (e.g., interactions with multiple overrides) but provides essential behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loading the core purpose and then adding a concrete use case. Every word is purposeful; no extraneous information. This is an ideal length for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description could mention what the tool returns. However, for a simple override tool, the provided information is largely sufficient. The description omits the ability to clear an override (which is documented in the schema), but the schema covers that gap, so overall completeness is good.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed parameter descriptions. The tool's description adds no additional parameter semantics beyond the schema; it only reiterates the use case. Therefore, the description provides minimal added value for parameter understanding beyond what the schema already offers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: temporarily override a Guardian rule's mode. It specifies the verb ('override'), resource ('Guardian rule'), and scope ('temporarily'). Among the sibling tools, no other tool does this, so it is well-distinguished.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a concrete use case ('when a rule is noisy during an incident and needs to drop to shadow for an hour without a server restart'), which helps the agent understand when to use it. However, it does not explicitly mention when not to use it or list alternative approaches.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_heartbeatA

Update session task, check inbox and pending questions. Auto-detects trust events.

Call this at the START of every user interaction (before doing work). May surface silent runtime signals DIARY_OVERDUE, GUARD_REMINDER, LEARNING_REMINDER, and PROTOCOL_DEBT; clients must treat those as internal obligations, not user-visible content. Output always begins with a NOW_UTC line (ISO-8601, UTC) — use it as the authoritative wall-clock time for any artifact (emails, diaries, followups) to avoid date/day-of-week drift across long sessions. Args: sid: Your session ID from nexo_startup. task: Brief description of current work (5-10 words). context_hint: Last 2-3 sentences from the user or current topic. Used for sentiment detection, trust auto-scoring, and mid-session RAG. ALWAYS provide this for best results.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
taskYes
context_hintNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behaviors: updates session task, checks inbox/pending questions, auto-detects trust events, may surface silent signals (DIARY_OVERDUE, GUARD_REMINDER, etc.), and output always begins with a NOW_UTC line. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured and front-loaded with summary and usage instruction. Composed of a few sentences without redundancy, though the parameter documentation could be slightly more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, no output schema, and no annotations, the description covers all necessary aspects: purpose, when to call, side effects, output format (NOW_UTC), and parameter details. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description must provide full parameter meaning. It does so clearly: sid as session ID from nexo_startup, task as brief 5-10 word description, context_hint as last 2-3 sentences for sentiment detection and trust auto-scoring, with advice to always provide for best results.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource combination: 'Update session task, check inbox and pending questions. Auto-detects trust events.' This clearly distinguishes it from sibling tools that focus on other operations like reconciliation, memory, or workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call at the START of every user interaction before doing work, and notes that silent runtime signals must be treated as internal obligations. Lacks explicit instructions on when not to use or alternatives, but the mandatory nature is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_hook_runsA

List recent hook lifecycle runs and per-hook health summary.

Closes Fase 3 item 7 of NEXO-AUDIT-2026-04-11. Each NEXO hook (session-start, post-compact, pre-compact, inbox-hook, etc.) writes a row to hook_runs when it finishes via scripts/nexo-hook-record.py. This tool reads them back so the agent can answer "is the hook pipeline healthy?" without needing the dashboard or grepping log files.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNoHow far back to look (default 24).
limitNoMax raw rows to return when summary_only=False (default 50).
statusNoOptional exact status filter (ok|error|skipped|timeout|blocked).
hook_nameNoOptional substring filter (LIKE %name%).
summary_onlyNoIf True, return only the per-hook health summary (success rate, p50/p95 duration, unhealthy hooks) and skip the raw row list.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It discloses the read-only nature implicitly ('List') and adds context about data source, but does not explicitly state it's non-destructive or mention auth/rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences) and front-loaded with the purpose. It includes necessary background without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 optional params, no output schema, no annotations), the description provides context and purpose but lacks details about return format, pagination, or behavior beyond the schema parameter descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add meaning beyond the schema's parameter descriptions; it only provides background context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List' and the resource 'hook lifecycle runs' and 'per-hook health summary'. It provides specific context (audit item, script source) but does not explicitly distinguish from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear use case ('answer is the hook pipeline healthy?') and context (no dashboard or log grepping), but does not provide explicit when-not-to-use instructions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_hot_context_listC

List hot-context items currently alive in the recent continuity window.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
limitNo
stateNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates the items are 'currently alive' (temporary) but does not specify side effects, read-only nature, or permissions. This is insufficient for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single 10-word sentence, which is concise but too brief to provide necessary context. It sacrifices clarity for brevity, earning a middle score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 parameters, no output schema, many siblings), the description is incomplete. It fails to define key terms like 'hot-context items' or 'continuity window', and does not cover return values or filter behavior, leaving substantial gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 3 parameters with 0% description coverage, and the description adds no explanations. The agent has no guidance on the meaning of 'state' or the role of 'hours' and 'limit' beyond their names and defaults. This makes correct parameter usage unlikely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool lists 'hot-context items' that are 'alive in the recent continuity window', providing a specific verb and resource. However, it does not differentiate from similar sibling tools like nexo_recent_context or nexo_context_router, leaving ambiguity about when to use this specific tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. Given numerous sibling tools for context retrieval (e.g., nexo_recent_context, nexo_context_packet), the absence of usage context or exclusions significantly hinders correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_index_add_dirB

Register a new directory for FTS5 search indexing. Survives restarts.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to directory (supports ~).
notesNoDescription of what this directory contains.
dir_typeNo'code' for source files, 'md' for markdown docs.code
patternsNoComma-separated glob patterns (only for code type).*.php,*.js,*.json,*.py,*.ts,*.tsx

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Only mentions persistence ('survives restarts'), but does not disclose other important behaviors such as overwriting existing registrations, permission requirements, or immediate indexing effects. With no annotations, the description carries full responsibility for transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise: two sentences that efficiently convey the core purpose and a key behavioral trait, with no superfluous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks information about return values (no output schema), error scenarios (e.g., duplicate directory, invalid path), or side effects. For a registration tool with multiple parameters, this omission significantly impairs completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and schema descriptions are adequate. The tool description adds no extra meaning beyond what the schema already provides, so the baseline score applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('register a new directory') and the purpose ('for FTS5 search indexing'), effectively distinguishing it from sibling tools like nexo_index_dirs and nexo_index_remove_dir.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like listing or removing directories, nor any prerequisites or conditions for use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_index_dirsA

List all directories being indexed by FTS5 (builtin + dynamic).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states that it lists directories being indexed, but does not disclose permissions, read-only nature, or any other behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is too brief for a tool with no output schema and no annotations. It does not explain FTS5 or what 'builtin + dynamic' means, nor does it indicate the format of the returned list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters and schema coverage is 100%. The description adds meaning by stating the tool's purpose, which is sufficient given zero parameters (baseline 4).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'List' and resource 'directories being indexed by FTS5', clearly distinguishing from sibling tools like nexo_index_add_dir and nexo_index_remove_dir.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies it is for inspection, but does not explicitly state when to use it versus alternatives or provide any when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_index_remove_dirB

Remove a directory from FTS5 indexing and clean up its entries.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to directory to remove.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must disclose behavioral traits. It mentions 'clean up its entries' but does not clarify if the operation is reversible, what permissions are needed, or if there are any side effects (e.g., performance impact).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that conveys the essential action and cleanup without redundancy or unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema, no annotations), the description covers the core operation. However, it could mention whether the operation is permanent or if there is a confirmation step, but for a straightforward removal tool it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (one parameter 'path' described). The description repeats the schema's parameter description without adding new semantic meaning, so baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Remove'), the resource ('directory from FTS5 indexing'), and the side effect ('clean up its entries'). It distinguishes itself from sibling 'nexo_index_add_dir' which performs the opposite operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'nexo_index_add_dir' or 'nexo_index_dirs', nor any prerequisites or conditions for removal. The description is silent on usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_intraday_memory_cycleC

Run a low-limit daytime memory observation cycle that only publishes evidence-backed intraday facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
backfill_limitNo
pending_sla_secondsNo

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavior. It mentions 'low-limit' and 'daytime' but does not explain side effects, idempotency, or potential impacts on state. Critical behavioral traits are omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise. However, it lacks structure (e.g., no separation of purpose, usage, or behavior) and does not front-load the most critical information for tool selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has three parameters, no output schema, and no annotations, the description is severely incomplete. It does not cover return values, expected behavior, or error conditions, leaving the agent without enough context to use it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description does not mention any of the three parameters (limit, backfill_limit, pending_sla_seconds). Thus, it adds no value beyond the schema's bare definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it runs a 'low-limit daytime memory observation cycle' and publishes 'evidence-backed intraday facts', giving some sense of purpose. However, it does not differentiate from sibling tools like nexo_memory_observation_process or nexo_memory_backfill, leaving ambiguity about when this specific tool should be used.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor any context about prerequisites, limitations, or exclusions. It fails to help the agent decide between this and other memory-related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_addC

Add a new learning (resolved error, pattern, gotcha).

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesShort title for the learning.
contentYesFull description with context and solution.
categoryYesFree-form category name (e.g., 'backend', 'frontend', 'devops', 'infrastructure', 'security'). Use consistent names across learnings.
priorityNocritical, high, medium, low (default: medium). Critical/high never decay below floor.medium
reasoningNoWHY this matters — what led to discovering this (optional).
applies_toNoFiles, systems, or areas this learning applies to (optional).
preventionNoConcrete rule/check that prevents repeating this mistake (optional).
review_daysNoDays until this learning should be reviewed again (default 30).
supersedes_idNoExisting learning ID this new canonical rule replaces (optional).
source_authorityNoAuthority tier for conflict resolution: francisco_correction, explicit_instruction, code_test_evidence, deep_sleep, inference.explicit_instruction

TDQS

C2.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. However, it only says 'Add a new learning' with no information on idempotency, side effects (e.g., does it trigger retroactive application?), conflict handling, auth requirements, or rate limits. This is severely lacking for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no unnecessary words. However, it lacks structure (e.g., no bullet points or clear breakdown) and could be slightly more informative without becoming verbose. It earns its place but could be improved.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 10 parameters, full schema documentation, and no output schema, the description is too brief. It does not mention what the tool returns (e.g., the ID of the created learning), any confirmation, or how the 'supersedes_id' or 'source_authority' parameters affect behavior. The description is insufficient for complete understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all 10 parameters having detailed descriptions in the input schema. The tool description adds no additional meaning beyond what schema provides (e.g., the parenthetical 'resolved error, pattern, gotcha' is implicit from the name and schema). Baseline of 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Add a new learning (resolved error, pattern, gotcha).' It uses a specific verb ('Add') and resource ('learning'), and differentiates from sibling tools like nexo_learning_update, nexo_learning_delete, and nexo_learning_list by focusing on creation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. Among many sibling learning tools (e.g., nexo_learning_update, nexo_learning_search, nexo_learning_delete), there is no context for when to choose 'add' over other operations, nor any prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_apply_retroactivelyA

Scan recent decisions and surface those that conflict with a learning's prevention rule.

Closes Fase 2 item 3 of NEXO-AUDIT-2026-04-11. Use this when you add a new rule and want to retroactively check whether past decisions still hold. Creates deterministic NF-RETRO-L-D followups so the helper is idempotent across reruns. nexo_learning_add invokes this automatically when the new learning has a prevention field — call this tool manually only when you want to re-scan with a longer window or a different threshold.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoIf True, scores matches but does not create followups.
min_scoreNoMatch threshold in [0.0, 1.0] (default 0.4).
learning_idYesID of the learning to apply.
max_matchesNoCap on followups created per call (default 5).
lookback_daysNoHow many days back to scan decisions (default 14).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that the tool 'creates deterministic NF-RETRO-L<learning>-D<decision> followups so the helper is idempotent across reruns' and explains the dry_run parameter behavior. It does not mention permissions, side effects, or safety, but the core behavioral traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a single purpose sentence, followed by a brief implementation detail (Fase 2 item 3), usage guidance, and behavioral note on idempotency. Every sentence adds value, and the total is under four lines. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 5 parameters and no output schema. The description covers the main intent, usage scenario, and key behavioral traits (idempotency, followup naming). It does not specify the return format (e.g., count of matches or followups created), which would be helpful but is not critical. Overall, it provides sufficient context for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add significant meaning beyond the schema for individual parameters. It provides overall context but lacks additional detail for each parameter. The schema already describes each parameter with clear names and defaults, so the description's contribution is minimal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Scan recent decisions and surface those that conflict with a learning's prevention rule.' It uses a specific verb (scan, surface) and resource (decisions related to a learning's prevention rule). It differentiates from sibling tool nexo_learning_add by noting that the latter automatically invokes this tool when a prevention field is present.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use: 'Use this when you add a new rule and want to retroactively check whether past decisions still hold.' It also explains when not to use manually: 'nexo_learning_add invokes this automatically when the new learning has a prevention field — call this tool manually only when you want to re-scan with a longer window or a different threshold.' This provides clear guidance on alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_deleteC

Delete a learning entry.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesLearning ID number.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description implies a destructive action ('delete') but does not disclose permanence, cascading effects, required permissions, or any side effects. With no annotations, the description carries the full burden, and it is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently conveys the basic purpose, but it lacks any additional context or structure. It is minimal but not informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (one parameter, no output schema), the description is minimally adequate. However, it could benefit from behavioral details like irreversibility or confirmation requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, providing the parameter 'id' with a basic description. The tool description adds no additional meaning beyond 'delete', so a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Delete a learning entry', specifying the verb and resource. However, it does not differentiate from siblings like nexo_learning_update or nexo_learning_apply_retroactively, which also modify learning entries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as nexo_learning_update for modifying or nexo_learning_apply_retroactively for retroactive application. There is no mention of prerequisites or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_listB

List all learnings, grouped by category.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoAccepted for Desktop/client compatibility; grouping is owned by the Brain handler.
categoryNoFilter by category (optional). If empty, shows all grouped.
created_afterNoFilter to learnings created at or after this date/time (optional).
created_beforeNoFilter to learnings created at or before this date/time (optional).

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It only states the basic function without disclosing behavioral traits such as read-only nature, side effects, pagination, or how grouping actually works. The agent lacks critical context about the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, very concise. It conveys the core purpose without extra words. However, it could include more information without becoming bloated, so it slightly misses the balance between conciseness and completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and many sibling tools, the description is too sparse. It does not explain the output format, grouping details, or how filters interact with grouping. An agent would need to infer significant behavior, making it incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and descriptions are provided for all parameters. The description adds minimal value beyond the schema: it mentions grouping by category but does not elaborate on created_after, created_before, or limit behaviors. Baseline 3 is appropriate as schema already details parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'List all learnings, grouped by category,' which provides a specific verb and resource. It distinguishes this tool from siblings like nexo_learning_search (search) and nexo_learning_add (add) by focusing on listing with grouping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for listing learnings with optional category filtering, but it does not provide explicit guidance on when to use this tool versus alternatives like nexo_learning_search. No when-not or prerequisite information is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_qualityB

Score learning quality so fragile rules can be strengthened before they mislead guard or retrieval.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoSpecific learning ID to inspect (optional).
limitNoMax learnings to score when listing (default 20).
statusNoFilter by lifecycle status such as active/superseded (default active).active
categoryNoFilter by category (optional).

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It fails to specify whether the tool modifies data (it only 'scores'), what permissions are required, or what the output looks like. The phrase 'strengthened' suggests reasoning but not the action of the tool itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise. However, it could be more structured by adding a second sentence to clarify usage or output. It is front-loaded but somewhat cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 optional parameters, no output schema, no annotations), the description should provide more context about return values, default behavior, and how scoring works. It only explains the 'why' but not the 'how' or 'what', leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for its 4 parameters, so the schema already documents usage. The description does not add meaningful semantic information beyond what the schema provides, thus baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Score learning quality') and its purpose ('so fragile rules can be strengthened'). It differentiates from sibling learning tools by focusing on quality scoring rather than CRUD operations. However, the term 'learning' is somewhat ambiguous without further context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (before fragile rules mislead guard or retrieval), providing some context. However, it does not explicitly state when not to use it or mention alternatives like nexo_learning_search or nexo_learning_update for different needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_resolve_candidateB

Dry-run the canonical learning resolver without creating or updating learnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
contentYes
categoryYes
priorityNomedium
reasoningNo
applies_toNo
preventionNo
supersedes_idNo
source_authorityNoinference

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses the dry-run behavior and lack of creation/update, but omits details about return values, error handling, or any other side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence and very concise. However, it sacrifices necessary detail for brevity, resulting in under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters, no output schema, and no annotations, the description is far too minimal. It fails to explain parameter usage, return values, or behavior under different inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% parameter description coverage, and the description adds no explanation of any of the 9 parameters. The agent must infer meaning from names alone, which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a 'dry-run' of the 'canonical learning resolver' and explicitly says it does not create or update learnings. This precisely distinguishes it from siblings like nexo_learning_add and nexo_learning_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for testing or previewing resolutions without side effects. However, it does not explicitly state when to use this tool versus alternatives, nor does it provide exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_learning_updateB

Update a learning entry. Only non-empty fields are changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesLearning ID number.
titleNoNew title (optional).
statusNoNew status such as active/superseded (optional).
contentNoNew content (optional).
categoryNoNew category (optional).
priorityNocritical, high, medium, low (optional).
reasoningNoNew reasoning/context (optional).
applies_toNoNew applies_to target(s) (optional).
preventionNoNew prevention rule (optional).
review_daysNoNew review interval in days (optional).
supersedes_idNoExisting learning ID this updated canonical rule replaces (optional).

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the key trait 'Only non-empty fields are changed' (partial update semantics), but omits important details such as error behavior on invalid id, return value, side effects, and permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently conveys the core purpose and a critical behavioral rule. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 11 parameters, no output schema, and no annotations, the description is too minimal. It lacks usage context, result format, error handling, and any guidance on how the 'only non-empty fields' behavior interacts with defaults. The agent would need additional information to use it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all 11 parameters, so the baseline is 3. The description does not add any further meaning to the parameters beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update a learning entry' provides a clear verb-resource pair, and the added detail 'Only non-empty fields are changed' distinguishes it from a full replacement update. However, it does not contrast it explicitly with sibling tools like nexo_learning_add or nexo_learning_delete, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives (e.g., nexo_learning_add for creation, nexo_learning_delete for removal). The agent is not told prerequisites, when not to use it, or how it fits in a workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_asset_getB

Return one indexed local asset by asset id.

ParametersJSON Schema
NameRequiredDescriptionDefault
asset_idYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It indicates a read operation ('Return') but does not disclose error behavior, existence requirements, or any side effects. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence that efficiently conveys the tool's core function without extraneous words. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description does not mention return values (no output schema) or error conditions. For a simple get operation it is adequate but could benefit from a hint about what is returned. Sibling tools are numerous but the purpose is clear enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter (asset_id) with 0% description coverage. The description adds 'by asset id', clarifying it's the lookup key, but does not elaborate on format, validation, or constraints. Minimal added value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Return' and the resource 'one indexed local asset' with the key parameter 'asset id'. It distinguishes this tool from siblings like nexo_local_asset_neighbors, which retrieves neighbors, by specifying singular retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no exclusions or prerequisites. The description simply states the function without context for selection among many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_asset_neighborsC

Return graph relations around one indexed local asset.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
asset_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and a minimal description, the tool's behavioral traits are undisclosed. It does not state whether the operation is read-only, destructive, or requires authentication. The description only says 'Return', implying read-only, but this is not explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at one sentence, but it lacks structure and front-loads no critical information. It earns its place but is too brief to be truly helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is insufficient. It fails to explain what the returned graph relations look like, their format, or any constraints (e.g., max limit). The agent is left guessing about the tool's full behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the description adds no information about the two parameters (asset_id, limit). The agent receives no additional meaning beyond the schema's basic types and default, making it hard to use correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns 'graph relations around one indexed local asset,' using a specific verb and resource. However, it does not elaborate on what 'graph relations' entails, which could be ambiguous, and it fails to distinguish from siblings like nexo_entity_dossier or nexo_local_asset_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., nexo_evidence_search, nexo_local_asset_get). The description lacks context for appropriate usage, such as specifying that it is for graph traversal or neighbor discovery.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_contextB

Retrieve local evidence before answering or acting.

Use mode='compact' for normal answers. Use mode='full' only for deep debugging, ideally with a higher max_chars and a specific query.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNocompact
limitNo
queryYes
intentNoanswer
max_charsNo
current_contextNo
include_entitiesNo
evidence_requiredNo
include_relationsNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility. It mentions retrieving local evidence but does not disclose whether the operation is read-only, side effects, rate limits, or what 'local evidence' constitutes. Minimal behavioral insight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two succinct sentences with a clear purpose and usage advice. No redundant text, well front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 9 parameters, no schema descriptions, no output schema, and no annotations, the description is insufficient. It fails to explain key parameters or the return format, leaving the agent underinformed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only addresses two parameters (mode and max_chars) out of nine. The query, limit, intent, and other parameters are left unexplained, adding little meaning beyond the schema's names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Retrieve local evidence before answering or acting,' specifying the verb and resource. It differentiates modes but does not explicitly distinguish from sibling tools like nexo_evidence_search or nexo_memory_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises using mode='compact' for normal answers and mode='full' for deep debugging with higher max_chars and specific query, providing clear context for when to use each mode. However, no guidance on when to avoid this tool in favor of siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_controlC

Control the local memory index.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNooptional folder to add and scan when action=run_once.
limitNooptional per-root scan limit for cooperative background cycles.
actionNoone of run_once, pause, resume, clear_index, set_performance.run_once
process_limitNomax pending jobs to process in this cycle.
performance_profileNolow, medium, high or extreme when action=set_performance.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should disclose behavioral traits. It only says 'control,' which implies mutating actions (e.g., clear_index, pause) but does not clarify side effects, permissions, or whether operations are reversible. The agent is left guessing about consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, but it sacrifices informativeness for brevity. Every sentence should earn its place; this one sentence fails to provide necessary context. It is under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters, no output schema, and no annotations, the description is incomplete. It does not explain the different actions (run_once, pause, resume, etc.) or their effects, leaving the agent with insufficient information to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already documented in the input schema. The description adds no extra semantic value beyond what the schema provides. Baseline 3 is appropriate, as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Control the local memory index' uses a vague verb 'control' and does not specify what operations are possible or the resource being affected. It barely adds value over the tool name, and among sibling tools like nexo_local_index_status and nexo_local_index_purge, the purpose is unclear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus sibling tools such as nexo_local_index_status or nexo_local_index_purge. The description lacks context about prerequisites or appropriate scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_diagnostics_tailC

Return recent local memory diagnostic log entries.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral disclosure, but it only states a read operation. It fails to mention any side effects, authentication needs, rate limits, or behavior details (e.g., whether the output is truncated, if it's a blocking call, etc.).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no redundant information. It is appropriately sized for a simple tool and front-loads the key action and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (1 parameter, no output schema, no annotations), the description is incomplete. It does not explain what 'diagnostic log entries' contain, the format of the output, or any prerequisites. A more complete description would include details on the log content or usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no meaning beyond the input schema. The sole parameter 'limit' is not mentioned or explained despite 0% schema description coverage. The agent gets no insight into how this parameter affects results (e.g., number of entries, default of 100).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the action (Return) and resource (recent local memory diagnostic log entries), making the tool's purpose immediately understandable. It distinguishes this tool from siblings like nexo_local_index_status or nexo_local_index_control, which focus on status or control rather than log retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as nexo_local_index_status or other diagnostic tools. There are no explicit when-to-use, when-not-to-use, or alternative suggestions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_exclusionsC

List, add or remove local memory exclusions.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
actionNolist
reasonNouser

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must convey behavioral traits. It indicates the tool performs both read (list) and write (add, remove) operations, but it does not disclose side effects (e.g., persistence, impact on system), required permissions, or aspects like idempotency. This is insufficient for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise (5 words), which is efficient but comes at the cost of essential detail. The single sentence front-loads the action but omits context. It is not a model of good conciseness as it sacrifices completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters with no descriptions, no output schema, and involves both reading and writing state, the description is severely incomplete. An agent cannot reliably decide how to format actions or interpret results without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention any of the three parameters (path, action, reason). Schema coverage is 0%, meaning the description adds no value beyond the raw schema. It fails to clarify default behaviors or valid inputs, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool can 'List, add or remove local memory exclusions', providing a specific set of actions and resource. This distinguishes it from other nexo_local_index_* siblings which focus on different aspects of local index management. However, it lacks detail on what constitutes a 'local memory exclusion'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like nexo_local_index_control. The description does not specify contexts appropriate for listing, adding, or removing exclusions, nor any conditions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_filetypesC

List, include, exclude or reset local memory file extension rules.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoextract
actionNolist
reasonNouser
extensionNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits, but it only mentions actions. It does not explain side effects, permissions needed, or what 'reset' entails, leaving significant ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with one sentence; however, it is appropriately structured for a simple tool. It could benefit from a brief note on each parameter, but it avoids unnecessary verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 0% schema coverage, 4 undocumented parameters, no output schema, and numerous sibling tools, the description is insufficient. It does not explain how the tool fits into the local index workflow or what each action does in detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should explain parameters. It provides high-level actions but does not document the meaning or valid values of 'action', 'extension', 'mode', or 'reason'. The agent cannot infer how to use these parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's actions: list, include, exclude, or reset local memory file extension rules. However, it does not differentiate from sibling tools like nexo_local_index_exclusions or nexo_local_index_control, which may have overlapping functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description only lists actions without context or prerequisites, leaving the agent unsure of the appropriate scenario for each action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_migrate_roots_v2C

Plan or apply Local Memory roots v2 cleanup.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It merely states 'Plan or apply... cleanup' without disclosing whether operations are destructive, require preconditions, or what side effects occur. The boolean parameter implies two modes but no behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with one sentence conveying the core purpose. No unnecessary words, but it could be slightly expanded to cover key behavioral aspects without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (1 param, no output schema, 0% schema coverage), the description is too bare. It lacks any context about outcomes, prerequisites, or risks, which is especially important for a potentially destructive cleanup operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter 'apply' with default false. The description hints that false means plan and true means apply, adding value beyond the schema which has no description. However, it does not explain what 'plan' outputs or what 'apply' does concretely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to plan or apply cleanup for Local Memory roots v2. It uses a specific verb-resource combination, distinguishing it from generic index management tools, though it could be more explicit about what 'cleanup' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus siblings like nexo_local_index_roots or nexo_local_index_control. Does not explain the difference between 'plan' and 'apply' modes or when to choose one over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_modelsC

Return or warm local model state for the Local Context Layer.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNostatus
local_files_onlyNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It mentions 'Return or warm' but does not explain side effects, network activity, or state changes. The description is insufficient for agent understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short (one sentence), which is concise, but it sacrifices clarity. It could be expanded with critical details while remaining concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two parameters with no schema descriptions and no output schema, the description should explain parameter behavior and return values. It fails to do so, leaving the tool underdocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no meaning to the parameters 'action' and 'local_files_only'. The default values are noted in the schema but their behavior and valid values (e.g., possible actions) are unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb ('Return or warm') and identifies the resource ('local model state') and context ('Local Context Layer'). However, it lacks differentiation from sibling tools like nexo_local_index_status or nexo_local_index_control, and the term 'warm' is informal and undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, context, or exclusions, leaving the agent without decision support.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_purgeB

Purge one indexed asset or clear the full local index.

ParametersJSON Schema
NameRequiredDescriptionDefault
asset_idNo
clear_allNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, and description does not disclose side effects, irreversibility, or impact on other tools. Only states the action without behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, to the point, no redundant words. Could benefit from slight expansion for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks details on return value, error conditions, and parameter interaction. Incomplete given no output schema and low schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Description implies usage of asset_id for single purge and clear_all for full clear, but does not clarify behavior when both are set or neither. With 0% schema coverage, description adds partial meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it purges one asset or clears the entire index, using specific verb and resource. Differentiates from siblings like nexo_local_index_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use asset_id vs clear_all, no prerequisites or context. Does not mention alternatives among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_rootsC

List, add or remove local memory roots.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNonormal
pathNo
depthNo
actionNolist

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description fails to disclose behavioral traits such as destructive nature of add/remove, side effects, or authorization requirements. The burden falls entirely on the description, which offers minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence. However, the brevity sacrifices necessary detail for a tool with multiple parameters and actions. It is not appropriately sized for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, no output schema, no annotations, and many sibling tools, the description is insufficient. It omits return behavior, parameter effects, and operation selection mechanism.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description does not explain any parameter (mode, path, depth, action) or how they relate to listing, adding, or removing. Without annotation or schema descriptions, the agent receives no added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool can 'list, add or remove local memory roots.' It identifies the resource and possible operations, but does not clarify what constitutes a 'root' or how these operations are invoked, leaving purpose vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like nexo_local_index_control or nexo_local_index_exclusions. No context on relevant scenarios or prerequisites for list/add/remove.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_service_configC

Render service configuration metadata for macOS, Windows or Linux.

ParametersJSON Schema
NameRequiredDescriptionDefault
platform_nameNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must convey behavioral traits. It only states 'Render service configuration metadata' which implies a read operation, but it does not disclose whether any changes are made, authorization needs, or side effects. The behavior is under-specified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that is front-loaded with the purpose. There is no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one parameter and likely returns configuration data, the description is too minimal. It does not explain what 'service configuration metadata' includes, how the platform_name parameter affects output, or any prerequisites. Completeness is low.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one optional parameter 'platform_name' with a default empty string, but the description does not mention or explain this parameter at all. Since schema coverage is 0%, the description adds no value for parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Render' and the resource 'service configuration metadata' for three specific platforms (macOS, Windows, Linux). However, it does not differentiate from sibling tools like 'nexo_local_index_status' or 'nexo_local_index_control'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor are there any usage restrictions or prerequisites mentioned. The description is purely a statement of function.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_local_index_statusA

Return local memory index status for Desktop settings and support diagnostics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states that the tool returns status, with no disclosure of behavior traits such as latency, staleness, safety (read-only), or authentication needs. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently conveys purpose and context without extraneous information. Every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no parameters and a simple purpose, the description lacks detail about the return value (structure, fields, format) and any potential errors or prerequisites. With no output schema, this omission reduces completeness for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and schema coverage is trivially 100%. The description does not need to add parameter details. Baseline score of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return') and the resource ('local memory index status'), and specifies its intended use ('for Desktop settings and support diagnostics'). It effectively distinguishes from sibling tools like nexo_local_index_control or nexo_local_index_diagnostics_tail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a diagnostic context but does not explicitly state when to use this tool versus alternatives, nor does it provide exclusions or comparative guidance. It leaves the agent to infer usage from the tool name and purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_managed_mcp_statusC

Read or apply the Brain-owned managed MCP catalog reconciliation plan.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It mentions 'read or apply' but provides no details on side effects, required permissions, or what applying entails. Minimal behavioral insight beyond the binary operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no extraneous words. It is front-loaded with the action and resource, but could be slightly more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is insufficient. It does not cover return values, error conditions, or behavior when 'apply' is true vs false. For a tool with one parameter, more detail is needed to ensure correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one boolean parameter 'apply' with 0% description coverage. The description does not mention or explain this parameter, leaving agents to guess the effect based on the tool's name and brief description. The parameter's role is inferred but not clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads or applies a specific resource (the Brain-owned managed MCP catalog reconciliation plan). The verb 'read or apply' and the specific plan reference provide clear purpose, though it does not explicitly differentiate from sibling tools like nexo_automation_reconcile.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. The description lacks any when-to-use or when-not-to-use context, leaving the agent to infer applicability.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_mcp_write_queue_statusC

Return durable MCP write queue counts and recent records.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states the return type but does not disclose any behavioral traits like read-only nature, potential side effects, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but lacks detail. It is appropriately front-loaded but too minimal to be fully informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple schema and no output schema, the description still lacks important context: what counts represent (pending, succeeded, failed?), what recent records contain, and default behavior. The tool is incomplete for an agent to use confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter ('limit') with 0% description coverage, and the tool description adds no explanation for it. The agent receives no guidance on how to use the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns 'durable MCP write queue counts and recent records', specifying the exact resource and outputs. It distinguishes itself from sibling tools, as no other tool focuses on the write queue status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, nor any exclusions or context for usage. With many sibling tools, such guidance would be beneficial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_answerC

Answer a memory question only when evidence exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
time_rangeNo
project_hintNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully convey behavioral traits. It mentions the evidence condition but does not explain what happens when evidence is missing, access requirements, or side effects. This is insufficient transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short (one sentence) but this brevity sacrifices needed detail. It is under-specified rather than concise, as important context is missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, no output schema, and no annotations, the description is grossly incomplete. It does not explain how to use the parameters, what the output looks like, or the behavior in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description adds no meaning to the parameters. The schema provides names and types but no explanations. The description should clarify the role of 'query', 'limit', 'time_range', and 'project_hint' but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it answers a memory question, distinguishing it from other memory tools like nexo_memory_search which searches for evidence rather than answers. However, it could be more specific about the nature of the answer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide guidance on when to use this tool versus siblings like nexo_memory_search. It only notes a precondition ('only when evidence exists') but lacks alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_backfillC

Backfill Memory Observations v2 from existing Brain tables.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sourcesNo

TDQS

C2.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description should disclose behavioral traits like side effects, permissions, or overwrite behavior. It only states the action without any caveats or details about what happens during backfill.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise at one sentence, which is efficient but under-specified. It is not verbose, but lacks necessary structural elements like parameter explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given two parameters, no output schema, and no annotations, the description is incomplete. It fails to explain what the parameters do, the backfill process, or any return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention or explain the parameters (limit and sources). With 0% schema description coverage, the agent receives no guidance on parameter usage beyond defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the action (backfill), the resource (Memory Observations v2), and the source (existing Brain tables). It is specific and distinguishes from sibling tools like nexo_memory_answer or nexo_memory_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool, when not to, or any alternatives. The description is a single sentence with no contextual usage hints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_event_listC

List raw Memory Observations v2 events captured by hooks/tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
source_idNo
event_typeNo
session_idNo
project_keyNo
source_typeNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It implies a read operation ('List') but fails to discuss pagination, performance impact, or any side effects. The description is too brief to provide adequate transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but lacks any structured breakdown. It is not verbose, but it also fails to pack meaningful detail into the limited space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, no output schema, and no annotations, the description is severely incomplete. An agent cannot understand the tool's capabilities, parameters, or expected output without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the 7 parameters (limit, query, source_id, etc.). The agent receives no semantic help beyond the parameter names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists 'raw Memory Observations v2 events' and specifies the source as 'hooks/tasks'. It is specific about verb and resource, though it could more explicitly differentiate from similar tools like nexo_memory_observation_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like nexo_memory_observation_list or nexo_memory_event_stats. There are no exclusions or context for appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_event_statsC

Summarize raw Memory Observations v2 event counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must convey behavioral traits. It only states it summarizes counts, but does not explicitly state it is read-only, side-effect-free, or other behavioral details beyond the implied read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (one sentence) and front-loaded with the core action. However, it sacrifices necessary detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even for a simple tool with one parameter and no output schema, the description lacks essential context such as what the summary includes, the output format, or any caveats about data freshness or filtering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage and one parameter 'days', the description adds no meaning beyond the schema. It does not explain the parameter's role (e.g., the time window for counts) or any constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool summarizes 'raw Memory Observations v2 event counts,' using a specific verb and resource. It distinguishes from siblings like 'nexo_memory_event_list' and 'nexo_memory_observation_stats' by specifying the summary nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives, nor any conditions or prerequisites. It fails to provide usage context or exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_forgetA

SELECTIVE-FORGET: verifiably erase a revoked secret, or correct a fact.

Two modes, never mixed:

  • mode='secret' (HARD-FORGET): scrubs the secret from EVERY table of ALL LIVE DBs the agent retrieves from — nexo, cognitive, local-context, local-context-usage, and email — discovered by live introspection (sqlite_master + PRAGMA) over each subsystem's own canonical path resolver, not a curated list — plus every FTS index (incl. local_chunks_fts), on-disk transcripts, and legacy shadow DBs (cognitive

    • local-context). The row is DELETED where it IS the secret (or carries an embedding/FTS copy), and REDACTED in place where the secret is embedded in an otherwise-useful record (diary, item_history.note, local_chunks.text, entity_facts.value, change_log...). It then RE-ENUMERATES and RE-SCANS everything and only returns complete=True at total zero; any survivor (even in a table not anticipated) is reported as complete=False with its . location. Use ONLY for revoked credentials/secrets or toxic data. Destructive run requires dry_run=False AND confirm=True (a dry-run count is returned otherwise). HNSW persisted indices are invalidated. SCOPE: point-in-time backups / snapshots are NOT swept (retention) — rotate the secret; the report's backup_scope field states this so complete=True never overclaims.

  • mode='fact' (CORRECT-FACT): does NOT delete. Useful memory is preserved via the existing reversible supersede; correct it with nexo_learning_update / supersede instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'secret' (default, hard) or 'fact' (soft, reversible).secret
valueYesthe literal secret value to forget.
reasonNofree-text audit note recorded in the ledger.
confirmNorequired (with dry_run=False) to actually delete in secret mode.
dry_runNowhen True (default) only count matches; never mutate.
use_regexNoalso match generic secret-shaped tokens (secret mode only).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully details the destructive behavior of secret mode, including the comprehensive scope (multiple DBs, FTS indices, transcripts, shadow DBs), the redaction process, the confirmation requirement (dry_run=False + confirm=True), and the return conditions (complete=True only if zero survivors). It also notes HNSW index invalidation and backup scope limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat lengthy but well-organized with bullet points and mode headings. It is front-loaded with the core purpose and mode distinction. The verbosity is justified by the tool's complexity, though a slight reduction could improve conciseness without losing essential detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool, the description comprehensively addresses all necessary aspects: two modes, parameter usage, safety mechanisms, return values (complete=True/False with survivor reporting), and edge cases like backup scope. Even without an output schema, the description adequately describes the tool's behavior and results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the input schema already covers all 6 parameters with descriptions, the tool description adds significant context: it explains the role of 'reason' as an audit note, clarifies that 'dry_run' and 'confirm' control destructive execution, and specifies that 'use_regex' is limited to secret mode. This goes beyond the schema's basic definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a selective forget mechanism with two distinct modes: 'secret' for hard deletion of revoked secrets and 'fact' for correcting facts without deletion. The verb 'erase' and 'scrubs' are specific, and the tool is well-distinguished from any likely siblings due to its unique purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use each mode: secret mode for revoked credentials/secrets or toxic data, and fact mode for corrections, directing the agent to use nexo_learning_update for fact corrections instead. This provides clear guidance on alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_healthB

Return Memory Observations v2 health and table status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must disclose behavioral traits. It only states the return of health and table status, omitting any information about side effects, authorization needs, rate limits, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, concise and front-loaded. It is minimally sufficient but could be slightly expanded for clarity without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema), the description covers the basic purpose. However, it lacks information on the expected output format or any contextual details that would help an agent interpret the results. For a health-check tool, this is minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100% (trivially). No parameter documentation is needed, and the description does not need to add meaning beyond the schema. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action ('Return') and resource ('Memory Observations v2 health and table status'). It is specific enough to distinguish from sibling tools like nexo_memory_observation_list or nexo_status, but could be more explicit about what 'table status' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor any conditions or prerequisites. The description is purely declarative without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_maintenanceD

Run safe Memory Observations v2 maintenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
retry_failedNo
process_limitNo
backfill_limitNo
backfill_sourcesNo

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description must disclose behavior. It only says 'safe', which is ambiguous and does not clarify whether the tool is read-only or modifies data, nor does it mention side effects, rate limits, or dependencies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (one sentence), but it omits crucial information. While brevity is valued, here it comes at the cost of clarity and completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has four optional parameters, no output schema, and many related sibling tools, the description is severely incomplete. It fails to explain what maintenance entails, what the parameters control, or what the outcome is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no explanation of the four parameters. Parameter names like retry_failed and process_limit are somewhat self-explanatory but lack context; the description does not clarify their role or expected values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource as 'Memory Observations v2' and action as 'run maintenance', but 'maintenance' is vague and does not differentiate the tool from similar memory tools like nexo_memory_observation_process or nexo_memory_backfill.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. There is no indication of when to use this tool versus siblings, no prerequisites, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_observation_listC

List passive Memory Observations v2 rows.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
statusNo
session_idNo
project_keyNo
observation_typeNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It only states 'List passive Memory Observations v2 rows' without explaining side effects, read-only nature, rate limits, or what 'passive' implies. The behavior beyond listing is opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but overly terse. It saves space but sacrifices clarity and completeness, failing to provide necessary context for a tool with multiple parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, no output schema, no annotations), the description is severely lacking. It does not specify return format, filtering behavior, or how parameters affect results, leaving the agent with insufficient information to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 6 parameters with 0% description coverage, and the description does not mention or explain any parameter. The agent receives no additional meaning beyond parameter names, which is insufficient for correct use.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies the action 'List' and resource 'passive Memory Observations v2 rows', making the basic purpose clear. However, it does not differentiate from similar sibling tools like nexo_memory_event_list or nexo_memory_search, lacking specificity about what makes observations 'passive'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There are no mentions of prerequisites, context, or filters that might distinguish it from other memory listing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_observation_processC

Process pending raw memory events into passive observations.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
backfill_limitNo
pending_sla_secondsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description lacks behavioral disclosure beyond the basic action of 'process'. There are no annotations to help, and the description does not mention side effects, idempotency, or whether it's destructive, leaving the agent uninformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence, which is efficient, but it is too brief and omits important details that could be included without significant expansion.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given three parameters, no output schema, and no parameter descriptions, the description is severely incomplete. It does not convey what the tool returns, how parameters affect behavior, or any constraints, leaving the agent with substantial ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description should explain the three parameters (limit, backfill_limit, pending_sla_seconds), but it provides no information about them. The agent must infer from parameter names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool processes pending raw memory events into passive observations, providing a specific verb and resource. However, it doesn't differentiate from sibling tools like nexo_memory_event_list or nexo_memory_observation_list, which have similar purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor are there any prerequisites, conditions, or exclusions mentioned. The agent has no context for invocation decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_observation_statsC

Summarize passive Memory Observations v2 rows and queue status.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states it 'summarizes', implying a read-only operation, but does not confirm lack of side effects, required permissions, or rate limits. The description is too sparse to ensure safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no fluff, achieving brevity. However, it omits critical details such as return format or parameter meaning, making it under-informative for its length. Conciseness should not come at the expense of completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter descriptions, the description fails to provide sufficient context for an agent to understand the tool's purpose, inputs, and outputs. Critical gaps include what 'queue status' entails and the structure of the summary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the description does not explain the 'days' parameter (integer, default 7). The agent must infer it represents the time range for summarization. The description adds no meaning beyond the schema field name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Summarize') and the specific resources ('passive Memory Observations v2 rows' and 'queue status'). It distinguishes from siblings like nexo_memory_observation_list (which lists) and nexo_memory_health (which checks health). However, it could be more explicit about the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like nexo_memory_observation_list or nexo_memory_event_stats. There is no mention of prerequisites, suitable contexts, or exclusions. The agent lacks decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_memory_timelineC

Return a chronological Memory Observations v2 timeline.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
time_rangeNo
project_hintNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries full burden. It only states it returns a chronological timeline but omits behavioral details such as pagination, ordering, return format, or side effects. With no annotations, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but overly sparse. It front-loads the main purpose but lacks necessary details about parameters and usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, no output schema, and no annotations, the description is incomplete. It fails to explain how to use the tool effectively, leaving the agent with unsupported assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the 4 parameters (limit, query, time_range, project_hint). The description adds no semantic value beyond what the schema's type hints offer.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns a chronological timeline of Memory Observations v2, giving a clear verb and resource. However, it does not differentiate from siblings like nexo_memory_observation_list or nexo_memory_search, which could cause ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the many related memory tools (e.g., nexo_memory_observation_list, nexo_memory_event_list). There is no mention of when-not-to-use or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_menuA

Generate the NEXO operations center menu with alerts and active sessions.

Shows: date, due alerts, all menu items by category, active sessions. Uses box-drawing characters for formatting.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the use of box-drawing characters for formatting, but does not explicitly state that the tool is read-only or has no side effects. For a menu generation tool, this is adequate but could be more transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each adding value: main purpose, content listing, formatting detail. No redundant or unnecessary information. Efficiently communicates the tool's behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (no parameters, no output schema), the description covers the key aspects: what it generates, what it shows, and formatting. It could mention the return format (e.g., plain text with box-drawing characters), but overall provides sufficient context for an agent to understand its use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters and schema coverage is 100%, so the baseline is 4. The description adds no parameter information because none exist, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a NEXO operations center menu and explicitly lists its contents (date, alerts, menu items by category, active sessions). This specific verb+resource combination uniquely distinguishes it from sibling tools like nexo_status or nexo_system_catalog.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when an overview menu is needed, but provides no explicit guidance on when to use this tool versus alternatives or any exclusions. Sibling tools like nexo_status or nexo_startup might serve related purposes, but no comparison is offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_opportunity_feedbackC

Record proposal feedback and apply suppression/snooze when requested.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
feedbackYes
proposal_idYes
snooze_untilNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavior. It mentions recording and suppression/snooze but omits details like side effects, required permissions, or the effect on opportunity state. The behavior of snooze_until is not clarified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at one sentence, but lacks structure such as bullet points or front-loading of key details. It is minimally viable but not optimally organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and four parameters, the description is incomplete. It does not cover return values, error scenarios, or detailed usage, leaving significant gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no meaning to the parameters. For example, it does not explain the expected format of snooze_until or the nature of feedback. The description provides no value beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool records proposal feedback and applies suppression/snooze, which is a clear verb-resource combination. However, it does not distinguish this from sibling tools like nexo_opportunity_suppress, and 'proposal feedback' is somewhat ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as nexo_opportunity_suppress or nexo_opportunity_get. The description lacks context for appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_opportunity_getC

Return one opportunity with evidence and read-only preparations.

ParametersJSON Schema
NameRequiredDescriptionDefault
opportunity_idYes
include_evidenceNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. 'Read-only preparations' hints at non-modification but does not disclose side effects, permissions, rate limits, or error handling. Minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise (8 words) and front-loaded. However, brevity sacrifices detail, making it less informative than it could be. Efficient but not maximally helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-param tool, the description is partial. It covers purpose but lacks usage guidance, parameter descriptions, and behavioral details. Without annotations, more completeness is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%. The description only hints at 'include_evidence' via 'with evidence,' but does not explain 'opportunity_id' or the effect of the boolean. Incomplete compensation for low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Return one opportunity with evidence and read-only preparations,' specifying the verb (Return), resource (opportunity), and key feature (evidence). It distinguishes from siblings like nexo_opportunity_queue (list) and nexo_opportunity_refresh (update). However, 'read-only preparations' is somewhat vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus siblings. It does not mention when to use, when not to use, or alternatives. The description lacks context for effective tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_opportunity_queueC

Return at most three evidence-backed proposals for a user surface.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
refreshNo
surfaceNohome
include_snoozedNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It mentions 'evidence-backed' and 'at most three' but does not disclose whether the tool is read-only, requires authentication, or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence. While it lacks detail, it is front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, no output schema, and no annotations, the description is severely incomplete. It fails to explain key concepts like 'user surface' or how parameters affect results, leaving agents with significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description adds no parameter information. The four parameters (surface, limit, refresh, include_snoozed) are unexplained, leaving the agent to infer meaning from names and defaults alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns evidence-backed proposals for a user surface and imposes an 'at most three' limit. However, it does not differentiate from sibling opportunity tools like nexo_opportunity_get or nexo_opportunity_refresh.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives or prerequisites. The description implies it is for retrieving proposals but does not specify scenarios or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_opportunity_refreshC

Generate Opportunity Orchestrator candidates from existing evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
sourcesNo
write_reportNo
limit_per_sourceNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must fully disclose behavioral traits. It mentions generating candidates but does not specify whether this is a read-only or mutation operation, what side effects occur, or if any data is modified. The phrase 'from existing evidence' is ambiguous about data sources and consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence, but this brevity sacrifices critical information. While it is front-loaded, it does not earn its place by providing sufficient detail for effective tool selection and invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, no output schema, no annotations, and a large set of sibling tools, the description is far from complete. It lacks details on return values, side effects, parameter meanings, and appropriate usage contexts, making it inadequate for an agent to reliably use this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with no parameter descriptions, and the description does not explain any of the four parameters (dry_run, sources, limit_per_source, write_report). The description adds no meaning beyond the schema's basic type/default information, leaving the agent unclear on how to use these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates Opportunity Orchestrator candidates from existing evidence, specifying a distinct verb and resource. It distinguishes from sibling opportunity tools like nexo_opportunity_get or nexo_opportunity_suppress, though it could be more precise about what 'existing evidence' encompasses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as nexo_opportunity_queue or nexo_opportunity_feedback. No contextual conditions, prerequisites, or exclusion criteria are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_opportunity_suppressC

Suppress repeated opportunity suggestions by scope.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNo
scope_keyYes
expires_atNo
scope_typeYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description implies a mutating action ('suppress'), but does not disclose behavioral traits such as reversibility, side effects, or whether the suppression is permanent or temporary. With no annotations, the description carries the full burden, which it fails to meet.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise. However, it is under-specified and could benefit from more structure. It is not verbose, but its brevity comes at the cost of clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, 4 parameters, and no annotations, the description does not cover basic contextual needs like what the tool returns, error conditions, or the overall effect. It is incomplete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It only adds 'by scope' which loosely relates to scope_type and scope_key, but provides no meaning for 'reason' or 'expires_at'. The description adds minimal value beyond parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb 'Suppress' and the resource 'opportunity suggestions', with the qualifier 'by scope'. This clearly indicates the tool's purpose but does not differentiate it from sibling tools like nexo_opportunity_feedback or nexo_opportunity_refresh.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as nexo_opportunity_feedback or nexo_opportunity_refresh. The description lacks context about prerequisites, typical use cases, or when suppression is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_plugin_listA

List all loaded plugins and their tools, showing source (repo/personal).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations; description mentions output content but not side effects or performance. Adequate for a simple read-only list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with action, no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no parameters and no output schema, the description fully conveys what it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters; schema coverage is 100%. Description adds no param info but baseline is 4 per rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it lists all loaded plugins and their tools, with source info. Distinguished from plugin load/remove siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or alternatives, but the purpose is self-evident for a parameterless list tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_plugin_loadA

Load or reload a plugin. Searches repo plugins/ first, then NEXO_HOME/personal/plugins/.

ParametersJSON Schema
NameRequiredDescriptionDefault
filenameYesPlugin filename (e.g., 'entities.py').

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses search order but fails to explain side effects (e.g., overwrite behavior, error cases, or if a reload updates existing plugins).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose, followed by search order. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a simple load operation but lacking details on output (no output schema), error handling, and concurrency (e.g., what happens if already loaded).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds value by specifying the search locations for the filename parameter. It goes beyond the schema's basic type definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Load or reload') and resource ('a plugin'), with additional detail on search order. This distinguishes it from sibling tools like nexo_plugin_list and nexo_plugin_remove.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (to load/reload a plugin) but lacks explicit guidance on when not to use it or alternatives. No mention of prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_plugin_removeA

Unregister a plugin's tools from MCP (does not delete files).

ParametersJSON Schema
NameRequiredDescriptionDefault
filenameYesPlugin filename (e.g., 'entities.py').

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description discloses one key behavioral trait (no file deletion), but without annotations, more details are needed about side effects, reversibility, prerequisites, or impact on other plugins. The disclosure is helpful but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise single sentence (11 words) that front-loads the action and includes a clarifying disclaimer. Every word is necessary and contributes to understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool, the description is largely complete. It explains the core function and a key limitation. However, it does not mention how to obtain the filename (e.g., via list) or describe the return value, which would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a clear description for the only parameter. The tool description adds no additional semantic value beyond what the schema already provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action 'Unregister a plugin's tools from MCP' with a specific verb and resource, and explicitly distinguishes what it does not do ('does not delete files'), setting it apart from potentially similar operations like deletion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like nexo_plugin_list or nexo_plugin_load. The action is implied by the name but lacks contextual instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_pre_action_contextC

Build the 24h recent-context bundle that should be reviewed before acting.

Especially useful for emails, orchestrators, and any work where the same topic may reappear hours later.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
limitNo
queryNo
session_idNo
context_keyNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior. It only says 'build ... bundle that should be reviewed,' suggesting a read-only operation, but fails to state side effects, authorization needs, or whether it modifies state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and front-loads the main purpose. It is concise, though it could be better structured to include parameter context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and 0% schema coverage, the description is far too brief. It does not explain the bundle's content, how to use parameters, or how it differs from similar sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the 5 parameters (hours, limit, query, session_id, context_key). The description adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Build the 24h recent-context bundle that should be reviewed before acting' with a specific verb and resource. However, it does not differentiate from sibling tools like nexo_recent_context or nexo_hot_context_list, which appear similar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It mentions 'especially useful for emails, orchestrators, and any work where the same topic may reappear hours later,' implying when to use, but lacks explicit when-not-to-use or alternative tool guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_pre_answer_routeC

Route a user turn through pre-answer continuity sources before responding.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidNo
areaNo
filesNo
queryNo
clientNomcp
intentNoauto
surfaceNopre_answer
budget_msNo
token_budgetNo
conversation_idNo
current_contextNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but only mentions routing 'before responding,' without disclosing side effects, error conditions, or other behavioral traits like required permissions or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it sacrifices completeness. It provides the basic action but no additional structure or details that would aid an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (11 params, no output schema, no annotations, many siblings), the description is far too minimal. An agent cannot effectively determine how to invoke or interpret results from this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 11 parameters with 0% description coverage, and the description adds no meaning beyond the parameter names. This is insufficient for an agent to understand how to properly populate the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('route') and the resource ('user turn through pre-answer continuity sources'), distinguishing it from many sibling tools that focus on answering or other tasks. However, it does not explicitly differentiate from similar routing tools like nexo_context_router.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as nexo_answer or nexo_context_router. The description lacks context for when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_product_answerC

Answer a NEXO product question using the structured product catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
localeNoes
questionYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral disclosure. It only states the tool 'answers' but does not reveal if it is read-only, requires authentication, has side effects, or returns cached or fresh data. The agent lacks critical safety and behavior information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but excessively terse. While it conveys the core purpose, it omits essential details that could be included without significant bloat, making it borderline under-informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, no output schema, and no annotations, the description is severely incomplete. It lacks information on return values, error handling, invocation constraints, and integration with other nexo tools. The agent cannot reliably use this tool based solely on the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the tool description adds no additional parameter meaning. The 'question' parameter is implicit but not detailed (e.g., format, query language). 'locale' and 'limit' are not described at all. The description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: answering NEXO product questions using a structured catalog. It distinguishes from sibling tools like nexo_answer and nexo_memory_answer by specifying 'product catalog' as the knowledge source. However, it does not elaborate on the scope or types of questions handled.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidelines are provided. The description does not indicate when to use this tool versus alternatives like nexo_answer or nexo_check_answer. There is no mention of prerequisites, limitations, or context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_product_capabilitiesC

Search the structured NEXO product capability catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
statusNo
categoryNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states 'Search', implying a read-only operation, but does not disclose behavioral traits such as case sensitivity, ordering, pagination, or side effects. With no annotations, the description carries the full burden and is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it lacks structure and necessary details. It could be improved with a brief example or mention of parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 optional parameters and no output schema, the description is incomplete. It fails to explain query format, category/status values, or limit behavior. Without annotations, more context is needed for effective use among many siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 4 parameters (query, category, status, limit) with no descriptions. The tool description does not explain the meaning, valid values, or format of any parameter. For example, what are valid values for category and status? This is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses 'Search the structured NEXO product capability catalog', clearly indicating the action and target. It is specific enough to differentiate from general search tools, but could be more distinct from siblings like 'nexo_capability_explain' or 'nexo_product_answer' that also deal with capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description does not indicate when to search rather than use other product-related tools like 'nexo_capability_explain' or 'nexo_product_knowledge_validate'. The agent has no context for selecting this tool over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_product_knowledge_validateB

Validate the structured NEXO product knowledge catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully convey behavioral traits. It states 'validate' but does not disclose whether the tool is read-only, what side effects (if any) occur, or what the output or impact is. The agent cannot assess safety or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words. It is highly concise and front-loaded, wasting no space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and no output schema, the description should provide more context about what 'validate' entails (e.g., returns success/failure, generates a report, or modifies state). The minimal description leaves the agent uncertain about the tool's outcomes and usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, and schema description coverage is 100%. The description does not need to add parameter semantics. A score of 4 is appropriate as the description does not detract and is sufficient for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool validates the structured NEXO product knowledge catalog, which is a specific and distinct purpose among siblings. However, 'validate' is somewhat generic and could be more specific about the nature of validation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor does it mention any prerequisites or context. It is a single sentence with no usage scenarios or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_product_surface_statusC

Show which NEXO product capabilities are exposed by a product surface.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
surfaceNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior. It only says 'Show', implying read-only, but lacks details on what 'product surface' means, parameter effects, or return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but it lacks necessary detail. It is under-specified rather than efficiently concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and abundant sibling tools, the description is insufficient. It does not explain what a product surface is, expected output, or how to effectively use the parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no information about the 'surface' and 'limit' parameters. Their purpose, default behavior, and constraints are unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool shows NEXO product capabilities exposed by a product surface. However, it does not distinguish from siblings like nexo_product_capabilities, which likely lists all capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. No prerequisites, context, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_protocol_debt_resolveC

Resolve protocol debt records by id or filters.

ParametersJSON Schema
NameRequiredDescriptionDefault
debt_idNo
task_idNo
debt_idsNo
debt_typesNo
resolutionNo
session_idNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description only says 'resolve', implying a mutation. Lacks details on side effects, reversibility, or idempotence. Minimal behavioral disclosure beyond the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise (one sentence), but lacks structure (e.g., bullet points or examples). Not wasteful, but could be better organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

6 parameters, no output schema, no annotations. Description omits return values, side effects, required context (e.g., session_id usage), and does not clarify resolution types. Severely incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. Description mentions only 'id' and 'filters' vaguely, failing to map to any of the 6 parameters (debt_id, task_id, debt_ids, debt_types, resolution, session_id). No parameter explanations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'resolve' and resource 'protocol debt records', and indicates two modes ('by id or filters'). It clearly states the tool's action and scope, though it does not differentiate from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description mentions id/filters but provides no context for choosing between them or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_recent_contextC

Read recent hot context and continuity events from the last N hours.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
limitNo
queryNo
context_keyNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must carry the full burden. It states 'read' implying non-destructive behavior, but lacks details on output format, pagination, authentication needs, or side effects. The minimal descriptor adds little beyond the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 10 words, front-loading the key action and resource. It is appropriately concise, though it sacrifices completeness for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, no output schema, and no annotations, the description is too sparse. It fails to explain return values, parameter interactions, default behaviors, or how this tool fits into the broader system.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description only clarifies the 'hours' parameter (by mentioning 'last N hours'). The other three parameters (limit, query, context_key) are completely unexplained, leaving the agent without necessary semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads 'recent hot context and continuity events' with a temporal scope of 'last N hours'. It uses a specific verb and resource, but lacks explicit differentiation from siblings like nexo_hot_context_list or nexo_context_packet.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. Among many context-related siblings, there is no mention of prerequisites, exclusions, or recommended contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_recent_context_captureB

Capture/update a recent 24h context item and append an event.

Use this for important ongoing threads that should stay mentally fresh across sessions/clients.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorNonexo
ownerNo
stateNoactive
titleYes
topicNo
detailsNo
summaryNo
metadataNo
source_idNo
ttl_hoursNo
session_idNo
context_keyNo
source_typeNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must bear the full burden. It mentions 'capture/update' and 'append an event' but does not disclose whether the operation is idempotent, what happens if a title already exists, any required permissions, or side effects. The behavioral traits are insufficiently described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description consists of two sentences, with the action front-loaded. It is reasonably concise, but could potentially incorporate more information without becoming verbose. No redundant or irrelevant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 13 parameters, no output schema, and no annotations, the description is far from complete. It covers only the basic purpose and a usage hint, leaving out parameter details, return value, error scenarios, and behavioral nuances. The tool's complexity demands a much richer description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the description adds no meaning beyond the parameter names. The description only implicitly references 'title' (required) but does not explain the purpose of the other 12 parameters (actor, owner, state, topic, details, summary, metadata, source_id, ttl_hours, session_id, context_key, source_type). This is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Capture/update a recent 24h context item and append an event.' The verb 'capture/update' combined with 'recent 24h context item' and 'append an event' distinguishes it from sibling tools like nexo_recent_context (likely read-only) and nexo_recent_context_resolve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a usage context: 'Use this for important ongoing threads that should stay mentally fresh across sessions/clients.' This implies the tool is for maintaining continuity on significant topics. However, it does not explicitly state when not to use it or mention alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_recent_context_resolveC

Resolve a recent hot-context item and append a resolution event.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorNonexo
topicNo
source_idNo
ttl_hoursNo
resolutionNo
session_idNo
context_keyNo
source_typeNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only says 'resolve' and 'append a resolution event', but does not explain what resolution entails (e.g., state changes, side effects, permissions). This is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no fluff. However, given the complexity of 8 parameters, it is too brief and lacks necessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fails to cover required aspects given the high parameter count, no output schema, and no annotations. It does not explain return values, outcomes, or parameter usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description provides no explanation of any of the 8 parameters. The schema only gives names and defaults; the description adds no meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Resolve' and the resource 'recent hot-context item', and mentions appending a resolution event. However, it does not differentiate from sibling tools like nexo_recent_context or nexo_recent_context_capture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. No context about preconditions or typical use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reindexA

Force full rebuild of the FTS5 search index. Use after bulk changes or if search seems stale.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. While it says 'Force full rebuild,' it does not explain side effects (e.g., temporary unavailability, resource consumption, or impact on ongoing searches). This lack of transparency is a significant gap for a potentially destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the action, and contains no extraneous words. Every sentence provides essential value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given tool simplicity (no params), the description covers purpose and usage. However, it omits return value or success/failure indication. Since there is no output schema, the description should hint at what the agent can expect (e.g., confirmation or error). This gap prevents full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema coverage is 100% (empty). Baseline for 0 parameters is 4. The description effectively conveys that no input is needed, so it meets the baseline without adding unnecessary detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Force full rebuild') and the resource ('FTS5 search index'). Among sibling tools like nexo_index_add_dir and nexo_local_index_control, this tool's purpose is distinct as a full index rebuild, making it easy to select.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage scenarios: 'after bulk changes or if search seems stale.' This gives clear context for when to invoke the tool. It lacks explicit when-not-to-use or alternatives, but the guidance is sufficient for a simple tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_completeA

Mark a reminder as completed with today's date.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReminder ID (e.g., R87).

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description alone must disclose behavior. It states the date is set to today, but does not mention side effects (e.g., whether the reminder must be active, whether it can be undone, or what happens to recurrence). Basic transparency but some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no extraneous words. It efficiently communicates the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description covers the essential context. However, it does not specify the return value or confirm action, which could improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a clear description of the 'id' parameter. The description adds no parameter-level details beyond the schema. Baseline score applies due to high schema quality.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Mark a reminder as completed with today's date' clearly states the action (mark as completed), the resource (reminder), and the specific behavior (using today's date). It distinguishes from sibling tools like nexo_reminder_create, nexo_reminder_delete, and nexo_reminder_update, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (when you want to complete a reminder) but does not explicitly state when to use this tool vs. alternatives like nexo_reminder_update, which might also change completion status. No when-not-to-use or prerequisite info is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_createC

Create a new reminder for the user.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesUnique ID starting with 'R' (e.g., R90).
dateNoTarget date YYYY-MM-DD (optional).
ownerNo'user' | 'waiting' | 'agent' | 'shared'. Leave empty to auto-classify by description heuristic.
categoryNoOne of: decisions, tasks, waiting, ideas, general.general
internalNo'1'/'true' to mark as agent bookkeeping (hidden from default user views). Leave empty to auto-classify.
descriptionYesWhat needs to be done.

TDQS

C2.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states 'Create a new reminder' without revealing any side effects, error conditions, permission requirements, or what happens if an ID already exists. This is insufficient for safe tool invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but it sacrifices informativeness for brevity. It does not earn its place by providing useful details beyond the tool name. Could be expanded slightly to include key behavioral notes without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, no output schema, and no annotations, the description is critically incomplete. It lacks any mention of return format, error handling, or success/failure conditions. The agent cannot determine what happens after creation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already described in the input schema. The description adds no additional meaning beyond the literal wording. Baseline score of 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create a new reminder for the user' clearly states the core action and distinguishes it from sibling tools like nexo_reminder_update, nexo_reminder_delete, etc. However, it lacks specificity about the required unique ID format starting with 'R' and the auto-classification behavior for owner and internal fields.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus sibling tools. For example, it does not mention that nexo_reminder_update is for modifying existing reminders or that nexo_reminder_complete marks them as done. The agent must infer usage solely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_deleteA

Soft-delete a reminder.

IMPORTANT: call nexo_reminder_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReminder ID (e.g., R87).
read_tokenNoToken returned by `nexo_reminder_get`.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description bears full burden. It states 'soft-delete' implying reversibility but does not explain soft-delete behavior (e.g., can be restored, hidden from lists). No contradictions; description is adequate but lacks depth on side effects or auth requirements beyond the token.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise: one sentence plus an imperative prerequisite. No wasted words, front-loaded with the action. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool with two params and no output schema, the description covers the essential purpose and prerequisite. It omits return value or confirmation, but that is acceptable given simplicity. Context is sufficient for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions. The description repeats that read_token comes from nexo_reminder_get, adding no new info beyond schema. Baseline 3 is appropriate as schema already documents parameters well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Soft-delete a reminder.' with a specific verb and resource, and distinguishes from sibling tools like nexo_reminder_complete or nexo_reminder_restore by the prerequisite of calling nexo_reminder_get first.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call nexo_reminder_get first and pass its read_token, providing a clear usage prerequisite. Could further differentiate from alternatives like nexo_reminder_complete or nexo_reminder_update, but the given guidance is helpful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_getA

Read a reminder with its history and usage rules.

IMPORTANT: before update/delete/restore/note, call this tool first and use the returned READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses that the tool reads a reminder and returns a READ_TOKEN needed for mutations. This indicates a safe read operation with a side effect of returning a token, which is useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences. The first sentence states the purpose, and the second provides an important usage instruction. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains what is retrieved (reminder with history and usage rules) and that it returns a READ_TOKEN. This is sufficient for an agent to understand the tool's behavior and usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one required parameter 'id' with 0% schema description coverage. The description does not further explain the parameter, but since it is a simple string identifier and the tool name implies what it represents, the meaning is clear. A higher score would require explicit mention of the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads a reminder with its history and usage rules. It uses a specific verb ('read') and resource ('reminder'), and it is distinct from sibling tools like nexo_reminder_create or nexo_reminder_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises to call this tool before update/delete/restore/note operations and to use the returned READ_TOKEN. This provides clear context for when to use it, though it does not mention when not to use it or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_noteA

Append a note to reminder history.

IMPORTANT: call nexo_reminder_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReminder ID (e.g., R87).
noteYesOperational note to append to history.
actorNoActor label for the history note.nexo
read_tokenNoToken returned by `nexo_reminder_get`.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states the action (append a note) and a prerequisite, but does not mention side effects, reversibility, authorization beyond the token, or error conditions. Minimal disclosure for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences front-load the purpose and critical prerequisite. Every word earns its place; no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward append operation with no output schema, the description covers the action and prerequisite. It lacks mention of return format or error scenarios, but is mostly complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by emphasizing the importance of the read_token parameter and its source, but does not elaborate on other parameters beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool appends a note to reminder history, using specific verb and resource. It distinguishes from sibling reminder tools (create, complete, delete, etc.) by focusing on appending notes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit prerequisite to call nexo_reminder_get first and pass its token. While it lacks explicit when-not-to-use or alternatives, the prerequisite provides clear context for appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_restoreA

Restore a soft-deleted reminder back to PENDING.

IMPORTANT: call nexo_reminder_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReminder ID (e.g., R87).
read_tokenNoToken returned by `nexo_reminder_get`.

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. Only mentions restore to PENDING and token requirement. Lacks details on permissions, side effects, or what happens if reminder is not soft-deleted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no extraneous words. Critical precondition emphasized in all caps. Highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema; description does not explain return values or what 'restore' entails. Misses clarifying 'soft-deleted' concept. Incomplete for a standalone tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions. Description adds context about read_token origin from nexo_reminder_get, but does not provide additional meaning beyond schema for the id parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Restore a soft-deleted reminder back to PENDING' with a specific verb, resource, and target state. Distinguishes from siblings like nexo_reminder_delete and nexo_reminder_complete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call nexo_reminder_get first and pass its READ_TOKEN, providing a clear prerequisite. Does not explicitly state when not to use, but the context implies it is only for soft-deleted reminders.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_remindersC

Check reminders and followups.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNo'due', 'all', 'followups', 'completed', 'deleted', 'history', or 'any'due

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must bear the burden of disclosing behavioral traits. It implies a read-only 'check' operation, but does not explicitly state whether it has side effects, requires authentication, or imposes rate limits. The description is insufficient for an agent to understand the tool's safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is efficient. However, it is overly brief and lacks structure (e.g., no separation of purpose and usage). Every word earns its place, but additional context could be added without becoming wordy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one optional parameter and no output schema, the description is minimally complete but fails to explain the output format (e.g., list of reminders? with what fields?) or the behavior when no filter is specified. An agent using this tool would need to infer too much, especially with sibling tools like nexo_reminder_get that have clearer scopes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (one parameter 'filter' with description). The description adds no additional meaning beyond the schema—it does not mention that the default is 'due' or explain the effect of different filter values. Since schema already documents the parameter adequately, no penalty, but no added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Check reminders and followups' clearly indicates the tool is for viewing/listing reminders and followups, but it does not differentiate from sibling tools like nexo_reminder_get (which retrieves a specific reminder) or nexo_reminder_complete (which marks complete). The verb 'check' is vague, and the resource is identified but not scoped.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. With numerous sibling tools for reminder operations (e.g., create, update, delete), the description lacks context for when 'check' is appropriate. No exclusions, prerequisites, or references to other tools are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_reminder_updateA

Update fields of an existing reminder. Only non-empty fields are changed.

IMPORTANT: call nexo_reminder_get first and pass its READ_TOKEN.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReminder ID (e.g., R87).
dateNoNew date YYYY-MM-DD (optional).
ownerNoNew 'user'|'waiting'|'agent'|'shared' (optional).
statusNoNew status (optional).
categoryNoNew category (optional).
internalNo'1'/'0' to re-classify visibility (optional).
read_tokenNoToken returned by `nexo_reminder_get`.
descriptionNoNew description (optional).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

States 'Only non-empty fields are changed', revealing how the tool handles optional parameters. No annotations are provided, so the description carries the burden; it adequately discloses the partial update behavior and the token requirement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences plus a highlighted IMPORTANT note. Every sentence is informative and necessary. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers core behavior, precondition, and optional parameter handling. Lacks output explanation, but no output schema exists. Sufficient for a mutation tool with clear prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the description adds meaning by explaining that empty fields are ignored. It also reinforces the read_token parameter's origin. This goes beyond what the schema alone provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Update fields of an existing reminder', which is a specific verb+resource. It distinguishes from sibling tools like nexo_reminder_create and nexo_reminder_delete by focusing on updates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call nexo_reminder_get first and pass its read_token, providing a clear prerequisite. Does not mention when not to use or compare with alternatives, but the directive is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_runtime_service_statusB

Return the resident NEXO Runtime Service status for diagnostics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description is responsible for disclosing behavioral traits. It only states it returns status, but does not mention whether it is read-only, any permissions required, or side effects. The agent lacks crucial information about the tool's safety and operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently conveys the core purpose. However, it sacrifices completeness for brevity, which is acceptable for such a simple tool but prevents a higher score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema, no annotations), the description is the sole source of information. It fails to specify what the status output contains (e.g., fields like running state) or how it differs from similar status tools, leaving the agent with incomplete understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is trivially 100%. The description does not need to explain parameters, and it meets the baseline of 4 by providing context that the schema cannot, albeit minimally.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the 'resident NEXO Runtime Service status' with a specific verb 'Return' and a defined purpose 'for diagnostics.' It is not a tautology and distinguishes somewhat from general status tools, though it could be more explicit about what makes it unique among siblings like nexo_status or nexo_heartbeat.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides minimal usage guidance beyond stating the purpose is 'for diagnostics.' It does not specify when to use this tool over alternatives, nor are there explicit when-not-to-use instructions. For a tool with many sibling status tools, more context is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_saved_not_used_auditC

Audit stores that write data without verified later consumption.

ParametersJSON Schema
NameRequiredDescriptionDefault
markdownNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description is expected to disclose behavior. It only states 'audit' which implies a read-only check, but potential side effects, permissions, or output details are omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise at one sentence, but it lacks necessary details such as what the output includes, making it too brief to be fully useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple schema and no output schema, the description should explain what 'stores' are and what the audit result looks like; it fails to provide this context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention the 'markdown' parameter, and the input schema has 0% description coverage, leaving the parameter's purpose entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool audits stores that write data without verified later consumption, making the specific resource and action identifiable. However, it does not differentiate from other audit-like sibling tools such as nexo_continuity_audit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool compared to alternatives like nexo_continuity_audit or nexo_confidence_check.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_sendB

Send a fire-and-forget message to another session or broadcast.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesMessage content.
to_sidYesTarget session ID, or 'all' for broadcast.
from_sidYesYour session ID.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided. The description says 'fire-and-forget' implying no delivery confirmation, but does not explain failure behavior, idempotency, or queuing. Minimal behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise single sentence with no redundant information. Front-loaded with the verb 'Send' and key attributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter tool with no output schema, the description is adequate but lacks details on delivery guarantees, error handling, or idempotency. Meets minimum viability.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds 'broadcast' context for to_sid='all', but that is already implied by the schema description. No additional semantics beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool sends a message (fire-and-forget) to a session or broadcast. It distinguishes from siblings like nexo_answer and nexo_ask by indicating no response is expected.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., nexo_answer, nexo_ask). The description only states what it does, not when to choose it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_compliance_stateC

Return Brain-verifiable heartbeat, diary, learning, and close compliance state.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidNo
diary_window_minutesNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It implies a read-only operation but does not disclose side effects, authentication needs, or rate limits. The term 'brain-verifiable' is ambiguous without further explanation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but the jargon 'brain-verifiable' may confuse, and the structure front-loads all information adequately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and minimal parameter explanation. For a tool returning compliance state, an agent needs to know what the output looks like and how parameters influence it. The current description is insufficient for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description does not explain the 'sid' parameter (session ID?) or how 'diary_window_minutes' affects the returned state. This leaves the agent guessing despite having default values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Return' and identifies the resource as 'brain-verifiable heartbeat, diary, learning, and close compliance state'. This clearly states the tool's output and scope, distinguishing it from sibling tools that don't mention compliance state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus the many sibling tools (e.g., nexo_session_diary_read, nexo_heartbeat). No exclusions, prerequisites, or context for appropriate invocation are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_diary_readC

Read recent session diaries for context continuity.

ParametersJSON Schema
NameRequiredDescriptionDefault
briefNo
domainNo
last_nNo
last_dayNo
session_idNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must imply behavioral traits. 'Read' suggests a safe, non-destructive operation, but details like default ranges (last_n, last_day) are not explained. The description lacks explicit statements about side effects or constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is front-loaded and concise. However, it could include slightly more detail without being overly verbose, such as clarifying default behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, no output schema, and no annotations, the description is too minimal. It does not cover return values, parameter interactions, or how it differs from similar tools, leaving significant gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with no parameter descriptions. The description does not explain any parameters, forcing inference from names and defaults. This adds minimal value beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Read) and resource (session diaries) with a purpose (context continuity). It distinguishes the tool from its write counterpart and other read tools, but does not explicitly differentiate from similar context-reading tools like nexo_recent_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives among siblings. It does not provide context for when it is appropriate or when it should be avoided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_diary_writeB

Write end-of-session diary with decisions, pending context, and self-critique.

Exposed as an essential MCP tool because Desktop close/archive/app-exit enforcement depends on it even when dynamic plugin loading is disabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNo
sourceNoclaude
pendingNo
summaryNo
decisionsNo
discardedNo
session_idNo
context_nextNo
mental_stateNo
payload_jsonNo
user_signalsNo
self_critiqueNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description reveals critical role in session shutdown but lacks details on side effects (e.g., overwrite, append), permissions, or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with front-loaded purpose. The second sentence adds contextual criticality but is slightly redundant. Efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 12 parameters, no output schema, and no annotations, the description is insufficient. Missing details on behavior, parameter formats, return values, and prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description partially compensates by naming a few parameters (decisions, pending, self-critique) but leaves most of the 12 parameters unexplained, requiring assumptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: writing an end-of-session diary with decisions, pending context, and self-critique. It distinguishes from sibling tool 'nexo_session_diary_read' and other session tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description mentions it's essential for enforcement but does not provide explicit use cases or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_export_bundleC

Export a machine-readable session bundle for cross-client handoff or archival.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidNo
pathNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral traits. It does not disclose whether the operation is destructive, modifies session state, or requires specific permissions. The mention 'export' implies a read operation but lacks explicit safety or side-effect details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words, achieving high conciseness. However, it leans toward being overly brief, sacrificing information for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter descriptions, the description is incomplete. It does not explain the bundle format, how the export works, or what to expect in response. Baseline complexity is moderate, but the description fails to cover essential context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no meaning to the two parameters (sid, path). The agent is left guessing what values are valid or how to use them. The defaults are also undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool exports a machine-readable session bundle for cross-client handoff or archival, providing a specific verb and resource. However, it does not differentiate from sibling tools like nexo_session_portable_context, which also deal with session transfers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., nexo_continuity_resume_bundle or nexo_session_compliance_state). The agent receives no context on prerequisites or scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_log_closeB

Close an automation_runs row opened by nexo_session_log_create.

ParametersJSON Schema
NameRequiredDescriptionDefault
errorNoshort error message if the session failed.
returncodeNochild exit code (0 = ok).
session_idYesid returned by the create call.
cost_sourceNoshort label for cost provenance.
duration_msNowall-clock duration in milliseconds.
input_tokensNo
output_tokensNo
total_cost_usdNocost in USD as a string (parsed to float).
telemetry_sourceNoshort label identifying where the counts came from ("desktop_stream", "codex_json", ...).
cached_input_tokensNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. The phrase 'Close an automation_runs row' implies a mutation, but the description does not explain what happens on closure (e.g., data finalization, side effects, or state changes). Given the lack of annotations, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that directly states the purpose. It is concise and front-loaded, with no wasted words. However, it is perhaps too brief given the tool's complexity, but it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, no output schema, and no annotations, the description is very short and does not cover essential context such as what closing entails, expected return values, or how to use parameters effectively. It is incomplete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 70% (7 of 10 parameters have descriptions), so the baseline is 3. The description adds no parameter-specific information; it only refers to the session_id indirectly. While the schema already explains most parameters, the description does not enhance understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'close' and the resource 'automation_runs row', directly referencing its counterpart 'nexo_session_log_create' to establish a clear pair. This effectively distinguishes it from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly indicates that this tool should be called after nexo_session_log_create by referencing 'opened by nexo_session_log_create', but it provides no explicit guidance on when to use vs. alternatives, no caution about required parameters (e.g., session_id is required but not emphasized), and no mention of prerequisites or side effects.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_log_createA

Open an automation_runs row for an interactive Claude/Codex session.

Designed for clients that spawn Claude/Codex directly (notably NEXO Desktop, which runs a TypeScript process that shells out to the CLI without going through run_automation_prompt). Call this BEFORE spawning the child, store the returned session_id, then call nexo_session_log_close when the session ends.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory the session is anchored to.
pidNoChild process PID if available.
modelNoConcrete model the client resolved, e.g. "claude-opus-4-7[1m]".
callerYesRegistered caller id (see src/resonance_map.py). For Desktop's "new conversation" button, use "desktop_new_session".
backendYes"claude_code" or "codex".
session_typeNo"interactive_chat" | "interactive_desktop" — how the session is shaped. Default "interactive_desktop".interactive_desktop
resonance_tierNoTier label ("maximo"/"alto"/"medio"/"bajo"). If left empty the Brain resolves it from caller.
context_excerptNoOptional first-prompt preview (used to size prompt_chars in telemetry).
reasoning_effortNoConcrete effort string, e.g. "xhigh".

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It states the tool 'opens' a row and returns a session_id, indicating creation. However, it does not disclose potential side effects, idempotency, error behaviors, or what happens on duplicate calls. This is adequate but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph with a clear front-loaded statement of purpose, followed by contextual details and a specific usage sequence. Every sentence contributes value, and there is no redundancy or unnecessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the purpose, context, and required sequence of calls (open, store id, close). It references the paired tool nexo_session_log_close. It does not detail the return format beyond 'session_id', but given no output schema, this is sufficient. It could be improved by mentioning behavior on duplicate calls or error conditions, but it is mostly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so the schema already documents each parameter well. The description itself does not add any additional semantics or usage notes for the parameters beyond stating that the session_id is returned. Thus, baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool opens an automation_runs row for an interactive Claude/Codex session. It specifies the use case for direct spawning (e.g., NEXO Desktop) and distinguishes it from the alternative run_automation_prompt, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs to call this BEFORE spawning the child, store the returned session_id, and call nexo_session_log_close when the session ends. It also provides context about when this tool is appropriate (direct spawning) and mentions the alternative path (run_automation_prompt).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_session_portable_contextB

Build a portable handoff packet for another client/runtime.

Use this when another client should continue the same work with explicit task/checkpoint/goal/workflow context instead of relying on memory alone.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It only states it builds a packet but does not disclose any behavioral traits like whether it is read-only, destructive, or requires authentication.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences, front-loading the purpose and usage. Every sentence adds value, though the parameter is not mentioned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and no parameter documentation, the description is incomplete. It does not explain what the packet contains, how it is used, or the role of the 'sid' parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter 'sid' with 0% description coverage, and the description does not explain its meaning or usage, adding no value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool builds a portable handoff packet for another client/runtime, specifying the action and resource. It differentiates from relying on memory alone, though not explicitly from sibling tools like nexo_context_packet.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: when another client should continue the same work with explicit context. It does not mention when not to use or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_skill_matchC

Find reusable NEXO skills for the current task.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
levelNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits such as read-only vs destructive, permissions, side effects, or error conditions. 'Find' implies reading, but this is not explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise but severely underspecified for a tool with two parameters. It lacks structure and does not front-load critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and parameter descriptions, the description is highly incomplete. It does not explain return values, parameter usage, or when to use the tool, leaving significant gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description should explain the parameters but does not. It does not mention 'task' or 'level' or provide any meaning beyond their raw schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool finds reusable NEXO skills for the current task. The verb 'Find' and resource 'skills' are specific, and the name distinguishes it from siblings like nexo_card_match or nexo_evidence_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks any context about prerequisites, exclusions, or preferred scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_smart_startupA

Pre-load relevant cognitive memories based on pending followups, due reminders, and last session topics.

Call during startup (after nexo_startup) to ensure the session starts with the right context loaded. Returns up to 10 memories matching the current operational state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that it pre-loads memories and returns up to 10, but it does not detail potential side effects (e.g., whether it modifies state) or performance impact. Without annotations, the burden is higher, and more detail would be beneficial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences), front-loaded with the core purpose, and contains no fluff. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers the essential aspects: purpose, timing, and result limit. It is complete enough for a simple parameterless tool, though a note on idempotency could add completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters, so the description correctly provides no parameter information. With 100% coverage from the schema, this is adequate; no additional parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: pre-loading cognitive memories based on specific criteria (pending followups, due reminders, last session topics). It distinguishes itself from the sibling tool nexo_startup by specifying it should be called after that, thus avoiding confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to call during startup after nexo_startup, providing clear timing. However, it does not mention when not to use this tool or suggest alternatives, which prevents a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_startupA

Register new session, clean stale ones, return active sessions + alerts.

Call this ONCE at the start of every conversation. Returns the session ID (SID) — store it for use in all other nexo_ tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoInitial task description.Startup
session_tokenNoExternal client session token. Claude Code passes its UUID via hooks; other clients may pass a synthetic durable token when useful. Pass this to enable automatic inter-terminal inbox detection when available.
session_clientNoOptional client label such as `claude_code` or `codex`.
conversation_idNoStable client-side conversation identifier when available.
session_providerNoOptional provider label such as `anthropic` or `openai`.
claude_session_idNoLegacy alias for the external client session token.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full burden. It discloses that it registers sessions and cleans stale ones (mutation), but doesn't detail authentication requirements, rate limits, or the exact nature of 'alerts.' Some behavioral aspects are implied but not fully explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences that front-load the core function and then provide critical usage instruction. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of output schema, the description adequately covers key return values (active sessions, alerts, SID). However, it could specify the format or structure of 'active sessions' and 'alerts' for clarity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining the importance of the session ID (SID) for other tools, but doesn't significantly expand on parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it registers a new session, cleans stale ones, and returns active sessions plus alerts. The verb 'register' and resource 'session' are specific. Distinguishes itself from sibling tools like nexo_smart_startup by being the primary initialization tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call it once at the start of every conversation. Tells the agent to store the returned session ID for use in all other nexo_ tools, providing clear when-to-use guidance and action items.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_statusB

List active sessions. Filter by keyword if provided.

ParametersJSON Schema
NameRequiredDescriptionDefault
keywordNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It only states it lists sessions and filters by keyword, omitting details like read-only nature, pagination, or output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence with no wasted words. While concise, it could include more key details without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one optional parameter, no output schema), the description is minimally adequate. It fails to explain what 'active sessions' entails or how results are structured, but sibling tools fill some gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explains the keyword parameter is used for filtering, adding meaning beyond the schema (which has 0% description coverage). However, it doesn't specify which fields are searched, so improvement is possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists active sessions, providing a specific verb and resource. However, it does not differentiate from sibling tools like nexo_session_compliance_state or nexo_session_diary_read, which may also involve session data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions optional keyword filtering, implying a basic usage context. But it does not specify when to use this tool versus sibling tools (e.g., for aggregated session view vs. detailed session logs).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_stopB

Cleanly close a session. Removes it from active sessions immediately.

Call this when ending a conversation to avoid ghost sessions. Args: sid: Session ID to close.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states the immediate effect (removes from active sessions) but lacks details about side effects, required permissions, or reversibility. Since no annotations are provided, the description should carry more weight, but it remains minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (two sentences plus an args line) and front-loaded with the core purpose. Every sentence serves a function, though the args line could be integrated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple stop tool with one parameter and no output schema, the description covers the basics. However, it does not address post-close behavior or potential triggers, which would be helpful given the large number of sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain the parameter. It only says 'sid: Session ID to close,' which is nearly a restatement of the schema. No additional meaning or context is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Cleanly close a session. Removes it from active sessions immediately.' This provides a specific verb (close/remove) and resource (session), making the tool's purpose unambiguous. It distinguishes from sibling session tools by focusing on termination rather than reading or writing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises calling this tool 'when ending a conversation to avoid ghost sessions.' This gives clear guidance on when to use it, though it does not mention when not to use it or suggest alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_support_ticket_closeB

Close a real NEXO support ticket after evidence has been recorded.

ParametersJSON Schema
NameRequiredDescriptionDefault
ticket_idYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It indicates a state-changing (closing) action but fails to mention potential side effects, required permissions, irreversibility, or error conditions that an agent would need to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no redundant information, efficiently conveying the action and precondition. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter descriptions, the tool description is insufficient. It omits return values, error handling, and prerequisites beyond evidence recording, leaving significant gaps for an agent to safely use the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the main description does not mention the 'ticket_id' parameter or its format. The agent receives no guidance on what value to provide or how to obtain it, severely hindering correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('close'), the resource ('real NEXO support ticket'), and a precondition ('after evidence has been recorded'), effectively distinguishing it from sibling tools like create, read, reopen, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit context by specifying the precondition 'after evidence has been recorded', which guides when to use the tool. However, it does not explicitly state when not to use it or mention alternatives like reopening.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_support_ticket_createC

Create a real NEXO support ticket for a product bug/setup issue.

ParametersJSON Schema
NameRequiredDescriptionDefault
originNodesktop
messageYes
subjectYes
priorityNonormal
client_message_idNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It only states that it creates a ticket, without disclosing side effects, authentication needs, rate limits, or that it writes to an external system.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is short but lacks structure. It could be improved by front-loading key parameter details or usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, 5 parameters, and no annotations, the description is incomplete. It does not explain return values, parameter meanings, or how it differs from other support ticket tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no information about the 5 parameters (subject, message, priority, client_message_id, origin). No hints on allowed values, defaults, or formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Create' and the resource 'real NEXO support ticket', with specific scope 'for a product bug/setup issue'. This distinguishes it from sibling tools like nexo_support_ticket_close or nexo_support_ticket_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when needing to create a support ticket for bugs or setup issues, but does not provide explicit guidance on when not to use it or what alternatives exist (e.g., nexo_support_ticket_message for adding messages).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_support_ticket_listA

List real support tickets from the NEXO backend for the signed-in Desktop user.

Use this when the user asks about support tickets or bug reports. Do not substitute a private followup for an actual product support ticket.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNo

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states that tickets are 'real' and for the 'signed-in Desktop user', implying authentication and real data. It lacks details on pagination, ordering, error states, or what happens when no tickets exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences that front-load the core purpose and then add a usage caution. No extraneous words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of schema descriptions and no output schema, the description is insufficiently complete. It does not cover filtering, sorting, pagination, or what the response contains, leaving significant gaps for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the description does not explain the parameters (status, limit) or their allowed values. The agent receives no guidance on how to use these parameters beyond their schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list', the resource 'real support tickets', and the context 'from the NEXO backend for the signed-in Desktop user'. It effectively distinguishes this tool from siblings like nexo_support_ticket_read or nexo_support_ticket_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool ('when the user asks about support tickets or bug reports') and provides a caution ('Do not substitute a private followup for an actual product support ticket'). While it doesn't list alternative tools, the guidance is clear and helpful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_support_ticket_messageA

Append an evidence note to a real NEXO support ticket before status changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
ticket_idYes
client_message_idNo

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It only states 'append an evidence note,' lacking details on error cases, idempotency, side effects, or prerequisites (e.g., ticket must exist). Minimal disclosure beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no fluff. Front-loaded with verb and resource, efficiently conveying the core purpose. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with 3 parameters, no output schema, and no annotations, the description is too brief. It omits return value, error handling, and prerequisites, leaving significant gaps for an AI agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description must compensate. It vaguely hints that 'body' is the note content and 'ticket_id' identifies the ticket, but does not explain 'client_message_id' at all. Inadequate given the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it appends an evidence note to a NEXO support ticket, specifying the verb and resource. It also differentiates from siblings by noting it is for adding notes before status changes, distinguishing from creation, closing, or reading tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context by stating 'before status changes,' implying when to use this tool in a workflow. However, it does not explicitly mention when not to use it or list alternative tools for different ticket operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_support_ticket_readC

Read one real support ticket from the NEXO backend by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
ticket_idYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. Description mentions 'real support ticket' implying authenticity but does not disclose read-only nature or other behavioral traits. Average for absence of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single short sentence, front-loaded with verb. No unnecessary words. Could be improved by adding more detail without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given only one parameter, no output schema, and no annotations, the description is too minimal. It doesn't describe return value, error cases, or any confirmation message.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%. Description only mentions 'by id' without elaborating on the ticket_id parameter or its format. Minimal value added beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states verb (Read), resource (support ticket), and source (NEXO backend). Distinguishes from sibling tools like list, create, close. Could be slightly improved by explicitly referencing ticket_id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives (e.g., nexo_support_ticket_list). No prerequisites or exclusions mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_support_ticket_reopenC

Reopen a real NEXO support ticket if fresh evidence shows it is active.

ParametersJSON Schema
NameRequiredDescriptionDefault
ticket_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions 'reopen' which implies a state change, but does not disclose side effects, authorization needs, or what happens to the ticket's history. The description lacks transparency about the tool's behavior beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence with no fluff, starting with the action verb. It is concise but could benefit from a brief mention of the parameter; nonetheless, it earns its place by providing the core purpose without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and annotations, the description does not provide sufficient context for an agent to use the tool confidently. It does not describe return values, success/failure conditions, or prerequisites (e.g., the ticket must exist and be closed). The tool is more complex than the description suggests.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter 'ticket_id' with 0% description coverage. The description does not mention this parameter at all, leaving the agent to infer its purpose solely from the name. With low coverage, the description should compensate by explaining the parameter's role, but it fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reopens a NEXO support ticket, using a specific verb and resource. It adds a condition ('if fresh evidence shows it is active') that slightly distinguishes it from similar tools like create or close, though the exact meaning of 'reopen' could be more explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'if fresh evidence shows it is active' gives some guidance on when to use the tool, implying it should only be used when new activity justifies reopening. However, it does not explicitly compare to siblings like nexo_support_ticket_close or nexo_support_ticket_create, nor does it state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_system_catalogC

Read NEXO's live system catalog built from core tools, plugins, skills, scripts, crons, projects, and artifacts.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
sectionNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits but only states it is a 'Read' operation. It does not mention any side effects, pagination behavior, limit semantics, or authorization requirements. The agent cannot predict key behavioral aspects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, concise and to the point. It is easily readable, but could benefit from a structured format like bullet points to separate usage hints or parameter explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given three parameters with no documentation, no output schema, and no annotations, the description is insufficient. It fails to explain what the output contains, how parameters influence results, or any constraints. The agent cannot reliably use this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no explanation for the three parameters (limit, query, section). Their meaning, format, and interaction are completely opaque, forcing the agent to guess or fail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads a live system catalog built from various NEXO components, distinguishing it from sibling tools that handle automation, memory, checkpoints, etc. The verb 'Read' and resource 'system catalog' are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like nexo_local_index_status or nexo_memory_search. The agent has no context to decide whether this catalog is appropriate for a given query.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_task_acknowledge_guardC

Acknowledge blocking guard rules on an open protocol task.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
noteNo
task_idYes
learning_idsNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must fully explain behavior. It only states 'acknowledge' but does not disclose side effects, permissions, or whether the action is destructive or safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but it is too brief for a tool with four parameters. Conciseness does not compensate for the lack of necessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (4 parameters, no output schema, no annotations), the description is extremely incomplete. It provides no context on how to use the tool or interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain parameters. It fails to mention any of the four parameters (sid, note, task_id, learning_ids) or their meanings and usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('acknowledge'), resource ('blocking guard rules'), and context ('on an open protocol task'), effectively differentiating it from sibling tools like 'nexo_guard_check' and 'nexo_guardian_rule_override'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidelines are provided. The description does not indicate when to use this tool versus alternatives, nor does it specify prerequisites (e.g., task must be open, blocking guard rules must exist).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_task_closeC

Close a protocol task with evidence and optional artifacts.

For high-stakes/irreversible closures (publish stable, broadcast, payment, force-push, revoke) pass work_type/stakes plus artifact_hash and last_human_validation_of_artifact_hash (both must match) so the close gates can be satisfied instead of dead-ending.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
resultNo
stakesNo
outcomeNo
summaryNo
task_idYes
evidenceNo
work_typeNo
change_whyNo
followup_idNo
change_risksNo
triggered_byNo
verificationNo
artifact_hashNo
change_verifyNo
evidence_refsNo
files_changedNo
followup_dateNo
outcome_notesNo
change_summaryNo
learning_titleNo
followup_neededNo
learning_contentNo
learning_categoryNo
followup_reasoningNo
learning_reasoningNo
correction_happenedNo
followup_descriptionNo
followup_verificationNo
verification_evidenceNo
partial_verification_reasonNo
partial_verification_acknowledgedNo
last_human_validation_of_artifact_hashNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior. It hints at 'dead-ending' and 'close gates' but does not clearly state side effects, required permissions, or whether the action is destructive. The behavioral context is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two short paragraphs. The first sentence clearly states the purpose, and the second adds specialized guidance without unnecessary fluff. Could benefit from bullet points for parameter lists, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high parameter count (33), lack of schema descriptions, no output schema, and no annotations, the description is incomplete. It does not explain return values, prerequisites, or side effects, and fails to cover most parameters. The agent would need additional context to use this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description adds meaning for only a few parameters (work_type, stakes, artifact_hash, last_human_validation_of_artifact_hash) but leaves the remaining 29 parameters, including required ones like sid and task_id, undocumented. Meaningful guidance is limited.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Close a protocol task') and resource ('task'), providing a specific verb+resource. However, it does not explicitly distinguish from sibling tools like nexo_closure_close, though the 'protocol task' context helps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives guidance for high-stakes closures by specifying required parameters, but it does not explain when to use this tool versus alternatives like nexo_task_open or nexo_closure_verify. No explicit when-not-to-use or alternative mentions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_task_frequencyA

Check which operational tasks are overdue based on their frequency.

Compares last execution date vs configured frequency. Returns overdue tasks or 'all tasks up to date'.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It explains the behavior (comparing dates, returning list or status message). It does not mention side effects, but the operation appears read-only. Additional details on permissions or limitations would be beneficial but not strictly required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two short sentences giving the purpose and then the logic. No wasted words, front-loaded with the main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description completely explains what the tool does and what it returns. It is fully sufficient for an agent to understand and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters (schema coverage 100%), so the baseline is 4. The description does not need to add parameter information since none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Check which operational tasks are overdue based on their frequency', specifying a verb ('check') and a resource ('operational tasks'). It explains the logic (comparing last execution date vs frequency), which distinguishes it from sibling task tools like nexo_task_list or nexo_task_log.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: to see overdue tasks, and states what it returns (overdue tasks or 'all tasks up to date'). It does not explicitly state when not to use or list alternatives, but the context is clear enough for a simple check tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_task_listB

Show execution history for operational tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoHow many days back to show (default 30).
task_numNoFilter by task number (optional).

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. However, the description only states the basic function without revealing safety aspects (e.g., read-only nature), permissions required, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of a single sentence that conveys the core purpose. It is front-loaded and efficient, though it could benefit from slightly more structure to improve readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool with two optional parameters and no output schema, the description is adequate but minimal. It lacks details about output format, ordering, pagination, or context for the execution history, which could help the agent use it more effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes both parameters with clear descriptions (task_num for filtering, days for time range). The tool description adds no additional meaning beyond the schema, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool shows execution history for operational tasks. It is specific about the resource (operational tasks) and action (show history), but does not differentiate from similar sibling tools like nexo_task_log, which might also show history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for viewing task execution history but provides no guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_task_logB

Record that an operational task was executed.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoExecution summary (optional).
task_numYesTask number from the checklist (e.g., '7', '7b').
reasoningNoWHY this task was executed now — what triggered it (optional).
task_nameYesTask name (e.g., 'Google Ads').

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must bear the burden of behavioral disclosure. It declares a write operation but omits details on idempotency, side effects, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that states the purpose without waste, though slightly more context could be added.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple logging tool with fully described parameters in the schema, the description is adequate but minimal; it lacks notes on return behavior or typical usage patterns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description adds no new information about parameters beyond what is already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('record') and resource ('operational task execution'), clearly distinguishing it from sibling tools like nexo_task_open or nexo_task_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use vs alternatives; usage is implied by the verb 'record' but no exclusions or context are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_task_openC

Open a protocol task for non-trivial work.

ack_rules accepts "#95,#156" / "95,156" / "[95, 156]" and, when the guard surfaces blocking rules, acknowledges them inline instead of requiring a separate nexo_task_acknowledge_guard call.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
areaNo
goalNo
planNo[]
filesNo
stakesNo
unknownsNo[]
ack_rulesNo
task_typeNoanswer
constraintsNo[]
descriptionNo
known_factsNo[]
context_hintNo
project_hintNo
evidence_refsNo[]
verification_stepNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It only mentions that ack_rules can acknowledge guards inline. It does not disclose other behavioral traits such as idempotency, side effects (e.g., whether opening a task modifies state), required permissions, or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two sentences. The first sentence states the core purpose, and the second provides useful detail on ack_rules. It is front-loaded but could be improved by adding more critical information without increasing length significantly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (16 parameters, no output schema, no annotations), the description is severely incomplete. It fails to explain most parameters, the return value, or how to use the tool effectively. It is far from adequate for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the format of the ack_rules parameter in detail but does not describe the other 15 parameters (sid, area, goal, etc.). The explanation for ack_rules is helpful but insufficient for the overall parameter set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Open a protocol task for non-trivial work', which clearly indicates the action (open) and resource (protocol task). It distinguishes from siblings like nexo_goal_open and nexo_workflow_open by specifying 'protocol task', but does not explicitly contrast them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description hints at usage by noting that ack_rules can acknowledge guards inline, avoiding a separate nexo_task_acknowledge_guard call. However, it does not provide guidance on when to use this tool versus alternatives like nexo_task_close or others, nor does it state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_tool_explainB

Explain a live NEXO tool/capability from the generated system catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, has side effects, requires authentication, or what the output format is. The agent lacks essential safety and response expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no extraneous words. It efficiently conveys the core functionality without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and annotations, the description should provide more context about what the explanation includes (e.g., description, input schema, usage examples) and what 'live' implies. The current text is too sparse for an agent to fully understand the tool's behavior and response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'name' parameter. The description implies 'name' refers to the tool or capability identifier in the catalog, adding some meaning beyond the raw schema. However, it does not explain constraints (e.g., exact spelling, case sensitivity, required format), so compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a clear verb ('Explain') and resource ('a live NEXO tool/capability from the generated system catalog'), distinguishing it from sibling action-oriented tools like nexo_goal_get or nexo_workflow_open. It effectively communicates the tool's meta-catalog purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this tool when the agent needs to understand another tool. However, there are no explicit when-to-use or when-not-to-use instructions, nor any mention of alternatives among the many sibling tools. The guidance is minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_trackA

Track files being edited. Detects conflicts with other sessions.

MUST call before editing any shared file. Args: sid: Your session ID. paths: List of absolute file paths to track.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
pathsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the full burden of behavioral disclosure. It explains conflict detection and the prerequisite call. However, it does not describe the return value, error handling, or exact side effects (e.g., locking).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences plus args), front-loaded with purpose, and every sentence adds value. The 'MUST' directive emphasizes importance without extra wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description lacks details about the return value (e.g., conflict info, success confirmation) and does not clarify batch behavior or error states, leaving the agent partially uninformed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates by explaining 'sid: Your session ID' and 'paths: List of absolute file paths to track,' adding crucial context beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Track files being edited' and 'Detects conflicts with other sessions.' It uses a specific verb-resource pair and distinguishes from sibling tools like nexo_untrack by focusing on tracking before editing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use: 'MUST call before editing any shared file.' This provides clear context for use but does not explicitly mention when not to use or alternatives beyond the implied opposite (nexo_untrack).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_transcript_readC

Read a full transcript fallback by session id, transcript display name, session_uid, or exact path.

ParametersJSON Schema
NameRequiredDescriptionDefault
clientNo
session_refNo
max_messagesNo
transcript_pathNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It indicates a read operation but fails to explain the 'fallback' behavior, any prerequisites, or side effects. Minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence is concise, but front-loading is acceptable. However, it lacks structure to convey key information like parameter purposes or usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks details on return format, which is critical since there is no output schema. The description does not fully address what a 'full transcript fallback' entails or how to use the parameters effectively, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

0% schema description coverage requires the description to explain parameters, but it only lists possible identifiers (session id, display name, etc.) without mapping them to the actual parameters (client, session_ref, max_messages, transcript_path). max_messages is not mentioned at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it reads a full transcript fallback by various identifiers, clearly indicating the verb and resource. However, it does not differentiate from sibling tools like nexo_transcript_recent or nexo_transcript_search, and 'fallback' is ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The term 'fallback' suggests a secondary option but without clarification. No exclusions or when-not-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_transcript_recentC

List recent Claude Code / Codex transcripts visible to NEXO.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNo
limitNo
clientNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It implies a read-only listing operation but omits details: whether it returns full transcripts or just metadata, any rate limits, authentication needs, or side effects. The description is too vague for an agent to predict behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence, which is concise but at the expense of necessary detail. It could be structured to front-load the core purpose and then elaborate on parameters. As is, it is too sparse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and three parameters, the description is highly incomplete. It fails to specify return format, filtering behavior, or how it relates to sibling tools. A more complete description is needed for reliable agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%; the description explains none of the three parameters (hours, limit, client). It mentions 'recent' but does not link it to the 'hours' parameter. Without any parameter explanation, the agent cannot correctly set inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies the verb 'list' and the resource 'recent transcripts', and adds context 'visible to NEXO'. It implicitly distinguishes from siblings like 'nexo_transcript_read' (single transcript) and 'nexo_transcript_search' (search across transcripts). However, it could be clearer by explicitly differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives (e.g., nexo_transcript_read, nexo_transcript_search). It does not mention any prerequisites or exclusions, leaving the agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_untrackB

Stop tracking files. If no paths given, releases all.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYesYour session ID.
pathsNoFile paths to release. Omit to release all.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden for behavioral disclosure. It only states 'Stop tracking files' without detailing side effects, reversibility, or required permissions. Minimal transparency beyond basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with action, no wasted words. Efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It does not explain return values, confirmation of release, or side effects. For a tool with this complexity, more completeness is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so baseline is 3. The tool description does not add additional meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Stop tracking files') and the resource ('files'). It distinguishes from sibling tools like 'nexo_track' by indicating the opposite operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for stopping file tracking and mentions conditional behavior ('If no paths given, releases all'), but does not explicitly state when to use versus alternatives or provide exclusions. Usage context is implied but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_compensationC

Return the compensation plan for a partially completed workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
checkpoint_limitNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must carry the burden. It only states it returns a plan, but does not disclose whether the operation is read-only, requires permissions, or has side effects. The behavioral context is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but lacks structure. It front-loads the purpose but omits important details, making it under-specified rather than efficiently concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given two parameters, no output schema, and no annotations, the description is too minimal. It fails to explain the return format, the meaning of 'compensation plan,' or how checkpoint_limit affects results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description does not mention either parameter (run_id, checkpoint_limit) or explain their roles. It adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool returns a compensation plan for a partially completed workflow, using specific verb and resource. It distinguishes from sibling tools like nexo_workflow_get which return workflow details or status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The phrase 'partially completed workflow' implies a condition, but there are no when-not-to-use instructions or references to other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_getC

Read the full durable workflow state.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
include_stepsNo
checkpoint_limitNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only says 'Read', implying non-destructive but omits details like whether pagination applies, whether the tool has rate limits, or what 'full state' includes. The parameters suggest limits on steps/checkpoints, but this is not explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise, but it is too brief to convey necessary detail. It achieves efficiency at the cost of completeness, missing key information about parameters and use context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (3 parameters, no output schema, no annotations), the description is severely incomplete. It does not explain return behavior, parameter effects, or how this tool fits among many workflow-related siblings. A helpful description would cover expected output and usage notes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no meaning to the three parameters (run_id, include_steps, checkpoint_limit). With 0% schema description coverage, the tool relies entirely on the description for parameter guidance, but it fails to explain any parameter's role or relationship to the 'full state'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and resource 'full durable workflow state', making the tool's purpose obvious. However, it does not differentiate from siblings like 'nexo_workflow_list' or 'nexo_workflow_open', which may have similar read behaviors. A more specific purpose (e.g., reading state of a single workflow by run_id) would improve clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. Siblings like 'nexo_workflow_list' or 'nexo_checkpoint_read' could serve overlapping queries, but the description does not exclude them or recommend this tool for specific contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_handoffC

Record a durable workflow handoff.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorYes
run_idYes
new_ownerNo
next_actionNo
handoff_noteNo
shared_stateNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states 'durable', implying persistence, but fails to mention side effects, ownership requirements, or whether the handoff is appended or overwrites. The minimal description leaves significant behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, concise and front-loaded. However, its brevity sacrifices informativeness; it could be expanded slightly to cover key details without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and the tool's moderate complexity (6 parameters), the description is incomplete. It does not explain the return value, success conditions, or how a handoff differs from other workflow operations. An agent would need additional domain knowledge to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no information about the six parameters (run_id, actor, new_owner, next_action, handoff_note, shared_state). The agent receives no semantic guidance beyond the parameter names, which are insufficient for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb 'Record' and the resource 'workflow handoff', making the purpose clear. Among sibling tools, it is unique as the only handoff-related tool, distinguishing it effectively. However, it does not explain what a workflow handoff entails, leaving some ambiguity about the exact operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like nexo_workflow_update, nexo_workflow_resume, or nexo_workflow_compensation. The description lacks context on the appropriate scenarios for a workflow handoff, making it hard for an agent to decide when to invoke this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_listC

List durable workflow runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNo
include_closedNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It only states 'list', implying a read operation, but lacks details on side effects, pagination, sorting, or error behavior. This minimal disclosure is insufficient for a listing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at three words, but this brevity sacrifices valuable information. While no extra words are wasted, the under-specification means it is not efficiently conveying necessary details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and numerous sibling workflow tools, the description lacks completeness. It does not explain return values, how results are ordered, or how it differs from other list/search tools like 'nexo_workflow_get', leaving significant gaps for agent understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must explain what parameters like 'limit', 'status', and 'include_closed' mean, but it offers no information. The agent cannot understand how to filter or control the list without additional context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List durable workflow runs' uses a specific verb and resource, clearly indicating the tool's function. It distinguishes from siblings like 'nexo_workflow_get' by focusing on listing multiple runs, though it could be more specific (e.g., 'list all workflow runs with optional filtering').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'nexo_workflow_get' or 'nexo_workflow_open'. It does not mention context, prerequisites, or when to avoid using it, leaving the agent without decision support.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_openC

Open a durable workflow run for long multi-step work.

ParametersJSON Schema
NameRequiredDescriptionDefault
sidYes
goalYes
ownerNo
stepsNo[]
goal_idNo
priorityNonormal
next_actionNo
shared_stateNo{}
workflow_kindNogeneral
idempotency_keyNo
protocol_task_idNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fails to disclose behavioral traits such as whether it starts execution, durability guarantees, or authorization requirements. 'Open' is vague.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very concise single sentence, but under-specified. Effective for front-loading the core purpose but misses important details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 11 parameters, no output schema, and no annotations, the description is grossly incomplete. It does not address return values, optional parameters, or typical usage scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no meaning for the 11 parameters. Agents cannot infer semantics for required 'sid' and 'goal' or optional ones.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it opens a durable workflow run for long multi-step work, aligning with the tool name. However, it does not differentiate from sibling workflow tools like resume or replay, which share similar purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description lacks context for selecting the appropriate tool among workflow-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_replayC

Replay recent checkpoints for a workflow run.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
run_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. 'Replay recent checkpoints' is vague and does not explain side effects, state changes, or safety, leaving the agent uninformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, front-loaded with key information, with no redundant or extra text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, output schema, and parameter descriptions, the description is heavily under-specified. It does not explain return values, behavior of 'replay', or how parameters affect operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It does not mention the 'limit' parameter or clarify 'run_id' beyond the name, failing to add meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'replay' and resource 'checkpoints for a workflow run', clearly stating the tool's action. However, it does not differentiate this tool from similar siblings like nexo_checkpoint_read or nexo_workflow_resume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor on prerequisites or context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_resumeC

Summarize the next actionable step for a workflow run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided and the description does not disclose any behavioral traits such as side effects, permissions, or whether it is read-only; the agent lacks critical info.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence, but it omits essential details; conciseness is achieved at the expense of completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a parameter with zero description, the description is insufficient for an agent to confidently use the tool; it lacks context about output, prerequisites, and usage boundaries.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter run_id lacks explanation in both schema and description; 0% schema description coverage means the description should compensate, but it does not explain what run_id is or how to obtain it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Summarize the next actionable step for a workflow run' clearly identifies the action (summarize) and the resource (next actionable step of a workflow run), distinguishing it from sibling tools like get or replay.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives; no context about prerequisites or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nexo_workflow_updateC

Update a workflow run with a replayable checkpoint.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorNo
ownerNo
run_idYes
summaryNo
evidenceNo
step_keyNo
run_statusNo
step_titleNo
max_retriesNo
next_actionNo
retry_afterNo
state_patchNo
step_statusNo
compensationNo
retry_policyNo
shared_stateNo
checkpoint_labelNo
requires_approvalNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It states 'update' but does not mention whether the operation is destructive, requires specific run state, or has side effects. The term 'replayable checkpoint' hints at replayability but lacks detail on what changes are made.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, but for a tool with 18 parameters and no annotation support, it is under-specified. Concision should not sacrifice completeness; here, important details are missing, making it insufficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 18 undocumented parameters, the description is severely incomplete. It fails to provide enough information for correct tool invocation or understanding behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it lists none of the 18 parameters. Parameters like 'state_patch', 'step_key', and 'run_status' are not explained, leaving the agent without guidance on how to populate them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('update') and the resource ('workflow run'), adding 'replayable checkpoint' which provides specific context. It distinguishes from siblings like nexo_workflow_get or nexo_workflow_replay by indicating a mutation with checkpointing, though differentiation is implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as nexo_workflow_replay or nexo_workflow_resume. No exclusions, prerequisites, or context are provided, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev7.38.7
    • Changednexo_task_close3 fields changed
      • addedInput schema / properties / partial_verification_acknowledged
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / partial_verification_reason
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • addedInput schema / properties / verification_evidence
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
  2. 74 tool updatesv7.37.4
    • Addednexo_api_call
    • Addednexo_capability_explain
    • Addednexo_closure_close
    • Addednexo_closure_item_get
    • Addednexo_closure_link
    • Addednexo_closure_next
    • Addednexo_closure_snapshot
    • Addednexo_closure_status
    • Addednexo_closure_triage
    • Addednexo_closure_verify
    • Addednexo_cortex_decide
    • Addednexo_create_app_token
    • Addednexo_credential_create
    • Addednexo_credential_delete
    • Addednexo_credential_get
    • Addednexo_credential_list
    • Addednexo_credential_update
    • Addednexo_drive_act
    • Addednexo_drive_dismiss
    • Addednexo_drive_reinforce
    • Addednexo_drive_signals
    • Addednexo_embedding_migration_status
    • Changednexo_entity_dossier1 field changed
      • changedInput schema / properties / max_facts / default
        Previous value: -3000New value: +120
    • Addednexo_followup_complete
    • Addednexo_followup_create
    • Addednexo_followup_delete
    • Addednexo_followup_get
    • Addednexo_followup_lifecycle
    • Addednexo_followup_note
    • Addednexo_followup_restore
    • Addednexo_followup_update
    • Addednexo_hook_runs
    • Addednexo_index_add_dir
    • Addednexo_index_dirs
    • Addednexo_index_remove_dir
    • Addednexo_learning_add
    • Addednexo_learning_apply_retroactively
    • Addednexo_learning_delete
    • Addednexo_learning_list
    • Addednexo_learning_quality
    • Addednexo_learning_resolve_candidate
    • Addednexo_learning_search
    • Addednexo_learning_update
    • Addednexo_local_index_filetypes
    • Addednexo_local_index_migrate_roots_v2
    • Addednexo_managed_mcp_status
    • Addednexo_memory_forget
    • Addednexo_opportunity_feedback
    • Addednexo_opportunity_get
    • Addednexo_opportunity_queue
    • Addednexo_opportunity_refresh
    • Addednexo_opportunity_suppress
    • Addednexo_plugin_list
    • Addednexo_plugin_load
    • Addednexo_plugin_remove
    • Changednexo_pre_answer_route3 fields changed
      • changedInput schema / properties / budget_ms / default
        Previous value: -2500New value: +0
      • addedInput schema / properties / surface
        Added value: +{
        +  "default": "pre_answer",
        +  "type": "string"
        +}
      • changedInput schema / properties / token_budget / default
        Previous value: -2500New value: +0
    • Addednexo_product_answer
    • Addednexo_product_capabilities
    • Addednexo_product_knowledge_validate
    • Addednexo_product_surface_status
    • Addednexo_reindex
    • Addednexo_session_log_close
    • Addednexo_session_log_create
    • Changednexo_startup1 field changed
      • addedInput schema / properties / session_provider
        Added value: +{
        +  "default": "",
        +  "description": "Optional provider label such as `anthropic` or `openai`.",
        +  "type": "string"
        +}
    • Addednexo_support_ticket_close
    • Addednexo_support_ticket_create
    • Addednexo_support_ticket_list
    • Addednexo_support_ticket_message
    • Addednexo_support_ticket_read
    • Addednexo_support_ticket_reopen
    • Changednexo_task_close4 fields changed
      • addedInput schema / properties / artifact_hash
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • addedInput schema / properties / last_human_validation_of_artifact_hash
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • addedInput schema / properties / stakes
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
      • addedInput schema / properties / work_type
        Added value: +{
        +  "default": "",
        +  "type": "string"
        +}
    • Addednexo_task_frequency
    • Addednexo_task_list
    • Addednexo_task_log
  3. 100 tool updatesv7.25.2
    • Addednexo_answer
    • Addednexo_ask
    • Addednexo_automation_reconcile
    • Addednexo_automation_supervisor
    • Addednexo_card_match
    • Addednexo_check_answer
    • Addednexo_checkpoint_read
    • Addednexo_checkpoint_save
    • Addednexo_cognitive_control_observatory
    • Addednexo_confidence_check
    • Addednexo_context_packet
    • Addednexo_context_router
    • Addednexo_continuity_audit
    • Addednexo_continuity_compaction_event
    • Addednexo_continuity_resume_bundle
    • Addednexo_continuity_snapshot_read
    • Addednexo_continuity_snapshot_write
    • Addednexo_cortex_check
    • Addednexo_entity_dossier
    • Addednexo_evidence_record
    • Addednexo_evidence_search
    • Addednexo_files
    • Addednexo_goal_get
    • Addednexo_goal_list
    • Addednexo_goal_open
    • Addednexo_goal_update
    • Addednexo_guard_check
    • Addednexo_guardian_rule_override
    • Addednexo_heartbeat
    • Addednexo_hot_context_list
    • Addednexo_intraday_memory_cycle
    • Addednexo_local_asset_get
    • Addednexo_local_asset_neighbors
    • Addednexo_local_context
    • Addednexo_local_index_control
    • Addednexo_local_index_diagnostics_tail
    • Addednexo_local_index_exclusions
    • Addednexo_local_index_models
    • Addednexo_local_index_purge
    • Addednexo_local_index_roots
    • Addednexo_local_index_service_config
    • Addednexo_local_index_status
    • Addednexo_mcp_write_queue_status
    • Addednexo_memory_answer
    • Addednexo_memory_backfill
    • Addednexo_memory_event_list
    • Addednexo_memory_event_stats
    • Addednexo_memory_health
    • Addednexo_memory_maintenance
    • Addednexo_memory_observation_list
    • Addednexo_memory_observation_process
    • Addednexo_memory_observation_stats
    • Addednexo_memory_search
    • Addednexo_memory_timeline
    • Addednexo_menu
    • Addednexo_pre_action_context
    • Addednexo_pre_answer_route
    • Addednexo_protocol_debt_resolve
    • Addednexo_recent_context
    • Addednexo_recent_context_capture
    • Addednexo_recent_context_resolve
    • Addednexo_reminder_complete
    • Addednexo_reminder_create
    • Addednexo_reminder_delete
    • Addednexo_reminder_get
    • Addednexo_reminder_note
    • Addednexo_reminder_restore
    • Addednexo_reminder_update
    • Addednexo_reminders
    • Addednexo_runtime_service_status
    • Addednexo_saved_not_used_audit
    • Addednexo_send
    • Addednexo_session_compliance_state
    • Addednexo_session_diary_read
    • Addednexo_session_diary_write
    • Addednexo_session_export_bundle
    • Addednexo_session_portable_context
    • Addednexo_skill_match
    • Addednexo_smart_startup
    • Addednexo_startup
    • Addednexo_status
    • Addednexo_stop
    • Addednexo_system_catalog
    • Addednexo_task_acknowledge_guard
    • Addednexo_task_close
    • Addednexo_task_open
    • Addednexo_tool_explain
    • Addednexo_track
    • Addednexo_transcript_read
    • Addednexo_transcript_recent
    • Addednexo_transcript_search
    • Addednexo_untrack
    • Addednexo_workflow_compensation
    • Addednexo_workflow_get
    • Addednexo_workflow_handoff
    • Addednexo_workflow_list
    • Addednexo_workflow_open
    • Addednexo_workflow_replay
    • Addednexo_workflow_resume
    • Addednexo_workflow_update
  4. 126 tool updatesv7.23.8
    • Removednexo_answer
    • Removednexo_api_call
    • Removednexo_ask
    • Removednexo_automation_reconcile
    • Removednexo_automation_supervisor
    • Removednexo_card_match
    • Removednexo_check_answer
    • Removednexo_checkpoint_read
    • Removednexo_checkpoint_save
    • Removednexo_confidence_check
    • Removednexo_context_packet
    • Removednexo_context_router
    • Removednexo_continuity_audit
    • Removednexo_continuity_compaction_event
    • Removednexo_continuity_resume_bundle
    • Removednexo_continuity_snapshot_read
    • Removednexo_continuity_snapshot_write
    • Removednexo_cortex_check
    • Removednexo_create_app_token
    • Removednexo_credential_create
    • Removednexo_credential_delete
    • Removednexo_credential_get
    • Removednexo_credential_list
    • Removednexo_credential_update
    • Removednexo_drive_act
    • Removednexo_drive_dismiss
    • Removednexo_drive_reinforce
    • Removednexo_drive_signals
    • Removednexo_entity_dossier
    • Removednexo_evidence_record
    • Removednexo_evidence_search
    • Removednexo_files
    • Removednexo_followup_complete
    • Removednexo_followup_create
    • Removednexo_followup_delete
    • Removednexo_followup_get
    • Removednexo_followup_note
    • Removednexo_followup_restore
    • Removednexo_followup_update
    • Removednexo_guard_check
    • Removednexo_guardian_rule_override
    • Removednexo_heartbeat
    • Removednexo_hook_runs
    • Removednexo_hot_context_list
    • Removednexo_index_add_dir
    • Removednexo_index_dirs
    • Removednexo_index_remove_dir
    • Removednexo_learning_add
    • Removednexo_learning_apply_retroactively
    • Removednexo_learning_delete
    • Removednexo_learning_list
    • Removednexo_learning_quality
    • Removednexo_learning_search
    • Removednexo_learning_update
    • Removednexo_local_asset_get
    • Removednexo_local_asset_neighbors
    • Removednexo_local_context
    • Removednexo_local_index_control
    • Removednexo_local_index_diagnostics_tail
    • Removednexo_local_index_exclusions
    • Removednexo_local_index_models
    • Removednexo_local_index_purge
    • Removednexo_local_index_roots
    • Removednexo_local_index_service_config
    • Removednexo_local_index_status
    • Removednexo_mcp_write_queue_status
    • Removednexo_memory_answer
    • Removednexo_memory_backfill
    • Removednexo_memory_event_list
    • Removednexo_memory_event_stats
    • Removednexo_memory_health
    • Removednexo_memory_maintenance
    • Removednexo_memory_observation_list
    • Removednexo_memory_observation_process
    • Removednexo_memory_observation_stats
    • Removednexo_memory_search
    • Removednexo_memory_timeline
    • Removednexo_menu
    • Removednexo_plugin_list
    • Removednexo_plugin_load
    • Removednexo_plugin_remove
    • Removednexo_pre_action_context
    • Removednexo_pre_answer_route
    • Removednexo_protocol_debt_resolve
    • Removednexo_recent_context
    • Removednexo_recent_context_capture
    • Removednexo_recent_context_resolve
    • Removednexo_reindex
    • Removednexo_reminder_complete
    • Removednexo_reminder_create
    • Removednexo_reminder_delete
    • Removednexo_reminder_get
    • Removednexo_reminder_note
    • Removednexo_reminder_restore
    • Removednexo_reminder_update
    • Removednexo_reminders
    • Removednexo_runtime_service_status
    • Removednexo_saved_not_used_audit
    • Removednexo_send
    • Removednexo_session_compliance_state
    • Removednexo_session_diary_read
    • Removednexo_session_diary_write
    • Removednexo_session_export_bundle
    • Removednexo_session_log_close
    • Removednexo_session_log_create
    • Removednexo_session_portable_context
    • Removednexo_skill_match
    • Removednexo_smart_startup
    • Removednexo_startup
    • Removednexo_status
    • Removednexo_stop
    • Removednexo_system_catalog
    • Removednexo_task_acknowledge_guard
    • Removednexo_task_close
    • Removednexo_task_frequency
    • Removednexo_task_list
    • Removednexo_task_log
    • Removednexo_task_open
    • Removednexo_tool_explain
    • Removednexo_track
    • Removednexo_transcript_read
    • Removednexo_transcript_recent
    • Removednexo_transcript_search
    • Removednexo_untrack
    • Removednexo_workflow_open
    • Removednexo_workflow_update
  5. 1 tool updatev7.23.2
    • Addednexo_automation_reconcile
  6. 10 tool updatesv7.23.0
    • Addednexo_automation_supervisor
    • Addednexo_entity_dossier
    • Addednexo_evidence_record
    • Addednexo_evidence_search
    • Addednexo_mcp_write_queue_status
    • Changednexo_memory_observation_process2 fields changed
      • addedInput schema / properties / backfill_limit
        Added value: +{
        +  "default": 100,
        +  "type": "integer"
        +}
      • addedInput schema / properties / pending_sla_seconds
        Added value: +{
        +  "default": 3600,
        +  "type": "integer"
        +}
    • Addednexo_pre_answer_route
    • Addednexo_saved_not_used_audit
    • Addednexo_session_compliance_state
    • Addednexo_session_diary_write
  7. 1 tool updatev7.21.0
    • Addednexo_runtime_service_status
  8. 166 tool updatesv7.20.23
    • Addednexo_answer
    • Addednexo_api_call
    • Addednexo_ask
    • Removednexo_automation_disable
    • Removednexo_automation_enable
    • Removednexo_automation_instructions
    • Removednexo_automation_schedule
    • Removednexo_automation_status
    • Removednexo_automations_list
    • Addednexo_card_match
    • Addednexo_check_answer
    • Addednexo_checkpoint_read
    • Addednexo_checkpoint_save
    • Removednexo_consolidate
    • Addednexo_context_packet
    • Addednexo_context_router
    • Addednexo_continuity_audit
    • Addednexo_continuity_compaction_event
    • Addednexo_continuity_resume_bundle
    • Addednexo_continuity_snapshot_read
    • Addednexo_continuity_snapshot_write
    • Removednexo_core_schedule_set
    • Removednexo_core_schedule_status
    • Removednexo_core_schedules_list
    • Addednexo_cortex_check
    • Addednexo_create_app_token
    • Addednexo_credential_create
    • Addednexo_credential_delete
    • Addednexo_credential_get
    • Addednexo_credential_list
    • Addednexo_credential_update
    • Addednexo_drive_act
    • Addednexo_drive_dismiss
    • Addednexo_drive_reinforce
    • Addednexo_drive_signals
    • Addednexo_files
    • Addednexo_followup_complete
    • Addednexo_followup_create
    • Addednexo_followup_delete
    • Addednexo_followup_get
    • Addednexo_followup_note
    • Addednexo_followup_restore
    • Addednexo_followup_update
    • Removednexo_goal_get
    • Removednexo_goal_list
    • Removednexo_goal_open
    • Removednexo_goal_update
    • Addednexo_guard_check
    • Addednexo_guardian_rule_override
    • Addednexo_heartbeat
    • Addednexo_hook_runs
    • Addednexo_hot_context_list
    • Addednexo_index_add_dir
    • Addednexo_index_dirs
    • Addednexo_index_remove_dir
    • Addednexo_learning_add
    • Addednexo_learning_apply_retroactively
    • Addednexo_learning_delete
    • Addednexo_learning_list
    • Addednexo_learning_quality
    • Addednexo_learning_search
    • Addednexo_learning_update
    • Addednexo_local_asset_get
    • Addednexo_local_asset_neighbors
    • Addednexo_local_context
    • Addednexo_local_index_control
    • Addednexo_local_index_diagnostics_tail
    • Addednexo_local_index_exclusions
    • Addednexo_local_index_models
    • Addednexo_local_index_purge
    • Addednexo_local_index_roots
    • Addednexo_local_index_service_config
    • Addednexo_local_index_status
    • Addednexo_memory_answer
    • Addednexo_memory_backfill
    • Addednexo_memory_event_list
    • Addednexo_memory_event_stats
    • Addednexo_memory_health
    • Addednexo_memory_maintenance
    • Addednexo_memory_observation_list
    • Addednexo_memory_observation_process
    • Addednexo_memory_observation_stats
    • Removednexo_memory_recall
    • Addednexo_memory_search
    • Addednexo_memory_timeline
    • Addednexo_menu
    • Removednexo_personal_script_remove
    • Removednexo_personal_script_unschedule
    • Addednexo_plugin_list
    • Addednexo_plugin_load
    • Addednexo_plugin_remove
    • Addednexo_pre_action_context
    • Removednexo_preference_delete
    • Removednexo_preference_get
    • Removednexo_preference_list
    • Removednexo_preference_set
    • Removednexo_protocol_debt_list
    • Addednexo_recent_context
    • Addednexo_recent_context_capture
    • Addednexo_recent_context_resolve
    • Removednexo_recover
    • Addednexo_reindex
    • Removednexo_remember
    • Addednexo_reminder_complete
    • Addednexo_reminder_create
    • Addednexo_reminder_delete
    • Addednexo_reminder_get
    • Addednexo_reminder_note
    • Addednexo_reminder_restore
    • Addednexo_reminder_update
    • Addednexo_reminders
    • Removednexo_run_workflow
    • Removednexo_schedule_add
    • Removednexo_schedule_status
    • Addednexo_send
    • Addednexo_session_diary_read
    • Addednexo_session_export_bundle
    • Addednexo_session_log_close
    • Addednexo_session_log_create
    • Addednexo_session_portable_context
    • Removednexo_skill_apply
    • Removednexo_skill_approve
    • Removednexo_skill_compose
    • Removednexo_skill_compose_candidates
    • Removednexo_skill_create
    • Removednexo_skill_evolution_candidates
    • Removednexo_skill_featured
    • Removednexo_skill_get
    • Removednexo_skill_list
    • Removednexo_skill_merge
    • Removednexo_skill_outcome_review
    • Removednexo_skill_promote
    • Removednexo_skill_result
    • Removednexo_skill_retire
    • Removednexo_skill_seed_from_outcome_pattern
    • Removednexo_skill_stats
    • Removednexo_skill_sync
    • Removednexo_skill_test
    • Addednexo_smart_startup
    • Addednexo_startup
    • Removednexo_state_watcher_create
    • Removednexo_state_watcher_list
    • Removednexo_state_watcher_run
    • Removednexo_state_watcher_update
    • Addednexo_status
    • Addednexo_stop
    • Addednexo_system_catalog
    • Addednexo_task_frequency
    • Addednexo_task_list
    • Addednexo_task_log
    • Addednexo_tool_explain
    • Addednexo_track
    • Addednexo_transcript_read
    • Addednexo_transcript_recent
    • Addednexo_transcript_search
    • Addednexo_untrack
    • Removednexo_update
    • Removednexo_user_state
    • Removednexo_user_state_history
    • Removednexo_user_state_stats
    • Removednexo_workflow_compensation
    • Removednexo_workflow_get
    • Removednexo_workflow_handoff
    • Removednexo_workflow_list
    • Removednexo_workflow_replay
    • Removednexo_workflow_resume
  9. 67 tool updatesv7.20.19
    • Addednexo_automation_disable
    • Addednexo_automation_enable
    • Addednexo_automation_instructions
    • Addednexo_automation_schedule
    • Addednexo_automation_status
    • Addednexo_automations_list
    • Addednexo_confidence_check
    • Addednexo_consolidate
    • Addednexo_core_schedule_set
    • Addednexo_core_schedule_status
    • Addednexo_core_schedules_list
    • Addednexo_goal_get
    • Addednexo_goal_list
    • Addednexo_goal_open
    • Addednexo_goal_update
    • Addednexo_memory_recall
    • Addednexo_personal_script_remove
    • Addednexo_personal_script_unschedule
    • Addednexo_preference_delete
    • Addednexo_preference_get
    • Addednexo_preference_list
    • Addednexo_preference_set
    • Addednexo_protocol_debt_list
    • Addednexo_protocol_debt_resolve
    • Addednexo_recover
    • Addednexo_remember
    • Addednexo_run_workflow
    • Addednexo_schedule_add
    • Addednexo_schedule_status
    • Addednexo_skill_apply
    • Addednexo_skill_approve
    • Addednexo_skill_compose
    • Addednexo_skill_compose_candidates
    • Addednexo_skill_create
    • Addednexo_skill_evolution_candidates
    • Addednexo_skill_featured
    • Addednexo_skill_get
    • Addednexo_skill_list
    • Addednexo_skill_match
    • Addednexo_skill_merge
    • Addednexo_skill_outcome_review
    • Addednexo_skill_promote
    • Addednexo_skill_result
    • Addednexo_skill_retire
    • Addednexo_skill_seed_from_outcome_pattern
    • Addednexo_skill_stats
    • Addednexo_skill_sync
    • Addednexo_skill_test
    • Addednexo_state_watcher_create
    • Addednexo_state_watcher_list
    • Addednexo_state_watcher_run
    • Addednexo_state_watcher_update
    • Addednexo_task_acknowledge_guard
    • Addednexo_task_close
    • Addednexo_task_open
    • Addednexo_update
    • Addednexo_user_state
    • Addednexo_user_state_history
    • Addednexo_user_state_stats
    • Addednexo_workflow_compensation
    • Addednexo_workflow_get
    • Addednexo_workflow_handoff
    • Addednexo_workflow_list
    • Addednexo_workflow_open
    • Addednexo_workflow_replay
    • Addednexo_workflow_resume
    • Addednexo_workflow_update
  10. 297 tool updatesv7.20.14
    • Removednexo_adaptive_decay
    • Removednexo_adaptive_history
    • Removednexo_adaptive_mode
    • Removednexo_adaptive_override
    • Removednexo_adaptive_reset
    • Removednexo_adaptive_weights
    • Removednexo_agent_create
    • Removednexo_agent_delete
    • Removednexo_agent_get
    • Removednexo_agent_list
    • Removednexo_agent_update
    • Removednexo_answer
    • Removednexo_artifact_create
    • Removednexo_artifact_delete
    • Removednexo_artifact_find
    • Removednexo_artifact_learn_alias
    • Removednexo_artifact_list
    • Removednexo_artifact_update
    • Removednexo_ask
    • Removednexo_auto_flush_recent
    • Removednexo_auto_flush_stats
    • Removednexo_automation_disable
    • Removednexo_automation_enable
    • Removednexo_automation_instructions
    • Removednexo_automation_schedule
    • Removednexo_automation_status
    • Removednexo_automations_list
    • Removednexo_backup_list
    • Removednexo_backup_now
    • Removednexo_backup_restore
    • Removednexo_card_catalog
    • Removednexo_card_get
    • Removednexo_card_match
    • Removednexo_change_commit
    • Removednexo_change_log
    • Removednexo_change_search
    • Removednexo_check_answer
    • Removednexo_checkpoint_read
    • Removednexo_checkpoint_save
    • Removednexo_claim_add
    • Removednexo_claim_get
    • Removednexo_claim_link
    • Removednexo_claim_lint
    • Removednexo_claim_search
    • Removednexo_claim_stats
    • Removednexo_claim_verify
    • Removednexo_cognitive_archive
    • Removednexo_cognitive_dissonance
    • Removednexo_cognitive_inspect
    • Removednexo_cognitive_metrics
    • Removednexo_cognitive_pin
    • Removednexo_cognitive_quarantine_list
    • Removednexo_cognitive_quarantine_process
    • Removednexo_cognitive_quarantine_promote
    • Removednexo_cognitive_quarantine_reject
    • Removednexo_cognitive_resolve
    • Removednexo_cognitive_restore
    • Removednexo_cognitive_retrieve
    • Removednexo_cognitive_sentiment
    • Removednexo_cognitive_snooze
    • Removednexo_cognitive_stats
    • Removednexo_cognitive_trigger_check
    • Removednexo_cognitive_trigger_create
    • Removednexo_cognitive_trigger_delete
    • Removednexo_cognitive_trigger_list
    • Removednexo_cognitive_trigger_preview
    • Removednexo_cognitive_trigger_rearm
    • Removednexo_cognitive_trust
    • Removednexo_confidence_check
    • Removednexo_consolidate
    • Removednexo_context_packet
    • Removednexo_continuity_audit
    • Removednexo_continuity_compaction_event
    • Removednexo_continuity_resume_bundle
    • Removednexo_continuity_snapshot_read
    • Removednexo_continuity_snapshot_write
    • Removednexo_core_schedule_set
    • Removednexo_core_schedule_status
    • Removednexo_core_schedules_list
    • Removednexo_cortex_check
    • Removednexo_cortex_decide
    • Removednexo_cortex_override
    • Removednexo_cortex_quality
    • Removednexo_cortex_review
    • Removednexo_cortex_stats
    • Removednexo_credential_create
    • Removednexo_credential_delete
    • Removednexo_credential_get
    • Removednexo_credential_list
    • Removednexo_credential_update
    • Removednexo_decision_log
    • Removednexo_decision_outcome
    • Removednexo_decision_search
    • Removednexo_diary_archive_read
    • Removednexo_diary_archive_search
    • Removednexo_doctor
    • Removednexo_drive_act
    • Removednexo_drive_dismiss
    • Removednexo_drive_reinforce
    • Removednexo_drive_signals
    • Removednexo_entity_create
    • Removednexo_entity_delete
    • Removednexo_entity_list
    • Removednexo_entity_search
    • Removednexo_entity_update
    • Removednexo_evolution_approve
    • Removednexo_evolution_history
    • Removednexo_evolution_propose
    • Removednexo_evolution_reject
    • Removednexo_evolution_status
    • Removednexo_files
    • Removednexo_followup_complete
    • Removednexo_followup_create
    • Removednexo_followup_delete
    • Removednexo_followup_get
    • Removednexo_followup_note
    • Removednexo_followup_restore
    • Removednexo_followup_update
    • Removednexo_goal_engine_status
    • Removednexo_goal_get
    • Removednexo_goal_list
    • Removednexo_goal_open
    • Removednexo_goal_profile_get
    • Removednexo_goal_profile_list
    • Removednexo_goal_profile_set
    • Removednexo_goal_update
    • Removednexo_guard_check
    • Removednexo_guard_cross_check
    • Removednexo_guard_file_check
    • Removednexo_guard_log_repetition
    • Removednexo_guard_stats
    • Removednexo_guardian_rule_override
    • Removednexo_heartbeat
    • Removednexo_hook_runs
    • Removednexo_hot_context_list
    • Removednexo_impact_score
    • Removednexo_index_add_dir
    • Removednexo_index_dirs
    • Removednexo_index_remove_dir
    • Removednexo_kg_export
    • Removednexo_kg_neighbors
    • Removednexo_kg_path
    • Removednexo_kg_query
    • Removednexo_kg_stats
    • Removednexo_learning_add
    • Removednexo_learning_apply_retroactively
    • Removednexo_learning_delete
    • Removednexo_learning_list
    • Removednexo_learning_quality
    • Removednexo_learning_search
    • Removednexo_learning_update
    • Removednexo_lifecycle_complete_canonical
    • Removednexo_lifecycle_event
    • Removednexo_lifecycle_status
    • Removednexo_lifecycle_stop_nexo_session
    • Removednexo_lifecycle_wait_for_diary
    • Removednexo_lifecycle_wait_for_stop
    • Removednexo_lifecycle_write_fallback_diary
    • Removednexo_local_asset_get
    • Removednexo_local_asset_neighbors
    • Removednexo_local_context
    • Removednexo_local_index_control
    • Removednexo_local_index_diagnostics_tail
    • Removednexo_local_index_exclusions
    • Removednexo_local_index_models
    • Removednexo_local_index_purge
    • Removednexo_local_index_roots
    • Removednexo_local_index_service_config
    • Removednexo_local_index_status
    • Removednexo_media_memory_add
    • Removednexo_media_memory_get
    • Removednexo_media_memory_search
    • Removednexo_media_memory_stats
    • Removednexo_memory_answer
    • Removednexo_memory_backend_status
    • Removednexo_memory_backfill
    • Removednexo_memory_event_list
    • Removednexo_memory_event_stats
    • Removednexo_memory_export
    • Removednexo_memory_health
    • Removednexo_memory_maintenance
    • Removednexo_memory_observation_list
    • Removednexo_memory_observation_process
    • Removednexo_memory_observation_stats
    • Removednexo_memory_recall
    • Removednexo_memory_review_queue
    • Removednexo_memory_search
    • Removednexo_memory_timeline
    • Removednexo_menu
    • Removednexo_outcome_cancel
    • Removednexo_outcome_check
    • Removednexo_outcome_list
    • Removednexo_outcome_pattern_candidates
    • Removednexo_outcome_pattern_capture
    • Removednexo_outcome_register
    • Removednexo_personal_plugin_create
    • Removednexo_personal_script_create
    • Removednexo_personal_script_remove
    • Removednexo_personal_script_schedules
    • Removednexo_personal_script_unschedule
    • Removednexo_personal_scripts_classify
    • Removednexo_personal_scripts_ensure_schedules
    • Removednexo_personal_scripts_list
    • Removednexo_personal_scripts_reconcile
    • Removednexo_personal_scripts_sync
    • Removednexo_plugin_list
    • Removednexo_plugin_load
    • Removednexo_plugin_remove
    • Removednexo_pre_action_context
    • Removednexo_preference_delete
    • Removednexo_preference_get
    • Removednexo_preference_list
    • Removednexo_preference_set
    • Removednexo_protocol_debt_list
    • Removednexo_protocol_debt_resolve
    • Removednexo_recall
    • Removednexo_recent_context
    • Removednexo_recent_context_capture
    • Removednexo_recent_context_resolve
    • Removednexo_recover
    • Removednexo_reindex
    • Removednexo_remember
    • Removednexo_reminder_complete
    • Removednexo_reminder_create
    • Removednexo_reminder_delete
    • Removednexo_reminder_get
    • Removednexo_reminder_note
    • Removednexo_reminder_restore
    • Removednexo_reminder_update
    • Removednexo_reminders
    • Removednexo_rules_check
    • Removednexo_rules_list
    • Removednexo_rules_migrate
    • Removednexo_run_workflow
    • Removednexo_schedule_add
    • Removednexo_schedule_status
    • Removednexo_send
    • Removednexo_session_diary_read
    • Removednexo_session_diary_write
    • Removednexo_session_export_bundle
    • Removednexo_session_log_close
    • Removednexo_session_log_create
    • Removednexo_session_portable_context
    • Removednexo_skill_apply
    • Removednexo_skill_approve
    • Removednexo_skill_compose
    • Removednexo_skill_compose_candidates
    • Removednexo_skill_create
    • Removednexo_skill_evolution_candidates
    • Removednexo_skill_featured
    • Removednexo_skill_get
    • Removednexo_skill_list
    • Removednexo_skill_match
    • Removednexo_skill_merge
    • Removednexo_skill_outcome_review
    • Removednexo_skill_promote
    • Removednexo_skill_result
    • Removednexo_skill_retire
    • Removednexo_skill_seed_from_outcome_pattern
    • Removednexo_skill_stats
    • Removednexo_skill_sync
    • Removednexo_skill_test
    • Removednexo_smart_startup
    • Removednexo_somatic_check
    • Removednexo_somatic_stats
    • Removednexo_startup
    • Removednexo_state_watcher_create
    • Removednexo_state_watcher_list
    • Removednexo_state_watcher_run
    • Removednexo_state_watcher_update
    • Removednexo_status
    • Removednexo_stop
    • Removednexo_system_catalog
    • Removednexo_task_acknowledge_guard
    • Removednexo_task_close
    • Removednexo_task_frequency
    • Removednexo_task_list
    • Removednexo_task_log
    • Removednexo_task_open
    • Removednexo_tool_explain
    • Removednexo_track
    • Removednexo_transcript_read
    • Removednexo_transcript_recent
    • Removednexo_transcript_search
    • Removednexo_untrack
    • Removednexo_update
    • Removednexo_user_state
    • Removednexo_user_state_history
    • Removednexo_user_state_stats
    • Removednexo_workflow_compensation
    • Removednexo_workflow_get
    • Removednexo_workflow_handoff
    • Removednexo_workflow_list
    • Removednexo_workflow_open
    • Removednexo_workflow_replay
    • Removednexo_workflow_resume
    • Removednexo_workflow_update
  11. 57 tool updatesv7.20.10
    • Addednexo_confidence_check
    • Addednexo_consolidate
    • Addednexo_core_schedule_set
    • Addednexo_goal_get
    • Addednexo_goal_list
    • Addednexo_goal_open
    • Addednexo_goal_update
    • Addednexo_memory_recall
    • Addednexo_preference_delete
    • Addednexo_preference_get
    • Addednexo_preference_list
    • Addednexo_preference_set
    • Addednexo_protocol_debt_list
    • Addednexo_protocol_debt_resolve
    • Addednexo_recover
    • Addednexo_remember
    • Addednexo_run_workflow
    • Addednexo_schedule_add
    • Addednexo_schedule_status
    • Addednexo_skill_apply
    • Addednexo_skill_approve
    • Addednexo_skill_compose
    • Addednexo_skill_compose_candidates
    • Addednexo_skill_create
    • Addednexo_skill_evolution_candidates
    • Addednexo_skill_featured
    • Addednexo_skill_get
    • Addednexo_skill_list
    • Addednexo_skill_match
    • Addednexo_skill_merge
    • Addednexo_skill_outcome_review
    • Addednexo_skill_promote
    • Addednexo_skill_result
    • Addednexo_skill_retire
    • Addednexo_skill_seed_from_outcome_pattern
    • Addednexo_skill_stats
    • Addednexo_skill_sync
    • Addednexo_skill_test
    • Addednexo_state_watcher_create
    • Addednexo_state_watcher_list
    • Addednexo_state_watcher_run
    • Addednexo_state_watcher_update
    • Addednexo_task_acknowledge_guard
    • Addednexo_task_close
    • Addednexo_task_open
    • Addednexo_update
    • Addednexo_user_state
    • Addednexo_user_state_history
    • Addednexo_user_state_stats
    • Addednexo_workflow_compensation
    • Addednexo_workflow_get
    • Addednexo_workflow_handoff
    • Addednexo_workflow_list
    • Addednexo_workflow_open
    • Addednexo_workflow_replay
    • Addednexo_workflow_resume
    • Addednexo_workflow_update

TDQS

C2.9/5.0

Scored across 170 tools

Disambiguation4/5

Every tool has a distinct resource+verb (get/update/delete/list by entity), but with 170 overlapping entities there is inevitably semantic overlap (e.g. multiple 'explain' or 'surface' tools); still, each appears to target a unique function.

Naming Consistency5/5

Uniform nexo_ prefix with snake_case verbs and nouns is extremely consistent across all 170 tools; pattern is predictable.

Tool Count1/5

170 tools is an extreme surface area. Even intentional 'mission control' servers rarely need this many; it overwhelms context and decision-making.

Completeness4/5

For the apparent domain (multi-session memory, checkpoints, events, sessions, products, external APIs), lifecycle coverage is thorough — create/update/close/query pairs exist across the board. Some areas (e.g., plugin removal edge cases) may still have gaps, hence not 5.

Maintenance

ActivityInactive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    F
    maintenance
    Cognitive memory system for AI agents with 129 MCP tools. Persistent 6-tier hierarchical memory (working→short-term→long-term→semantic), Ebbinghaus forgetting curves, dream consolidation, hybrid retrieval (BM25+RRF), goal tracking, emotional recall, knowledge graphs, and a 26-job consciousness daemon. Works with Claude Code, Cursor, and any MCP client.
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Persistent, auditable memory for AI agents. Hybrid BM25 + vector recall with 18 MCP tools, adaptive block metadata (A-MEM), intent-aware routing, contradiction detection, and governance workflows. Zero external dependencies. Drop-in memory for Claude Code and any MCP-compatible agent.
    261 PyPI
    18
    Apache 2.0
  • A
    license
    Not graded
    quality
    D
    maintenance
    Long-term memory for AI agents over MCP — episodic + semantic memory, a temporal knowledge graph, and a dialectic user model, exposed as 32 tools (recall, remember, context, graph, dreaming, peers). Zero dependencies, runs fully offline; leads the LoCoMo benchmark at ~35x fewer LLM calls.
    2
    Apache 2.0