Skip to main content
Glama
event4u-app

@event4u/agent-config

Official

Agent Config — every claim machine-checked, including the counts in these badges

Smoke Public install smoke (3 OS × 2 Node) npm agent-config MCP server MCP Toplist

Skills Rules Commands Guidelines Personas Advisors

How these are counted — one canonical counter, agent-configupdate_counts --check, re-derives all six from the tree and fails CI on a drift of one. Two bases are not what the linked directory shows, so they are stated here rather than left to be inferred: Commands 202 counts every command file recursively (the linked directory holds 61 at its top level), and Rules 120 counts the source rules while the linked projection holds 119 — one rule is dormant and is not projected. Personas 29 excludes the directory README. Counts are of files and directories: none of them measures quality, activation, or adoption.

Try it in 30 seconds — drop one read-only subagent into any repo and watch it gate "done": @production-validator check this branch is actually done. No wizard, no lock-in, nothing else installed — the 30-second wedge ↓ is the whole first step. Start at the proof, not the catalog: event4u-app.github.io/agent-config/proof/.

The trust surface running green — every "verify it yourself" command from a real, CI-re-executed run

Every public claim in this README is machine-checked — verify it yourself. In a market that runs on unbacked headline numbers, this one binds each claim to resolvable evidence or fails its own build.

Choose your experience — developer · founder · content · agency · finance · ops. Add packs. Get a focused command set, not a 500-artefact dump. Bring your own AI provider.

A deep library of skills, commands and governed rules — plus a capability router that loads the right skill on intent and multi-agent orchestration with consensus review. The whole layer is compiled into 20 host agents — of 23 detected, 3 being export-only (Claude Code, Cursor, Augment, Cline, Windsurf, Copilot, Gemini CLI, Codex, Continue, Zed, JetBrains, Aider and more). Resident processes are permitted only under the supervision contract ADR-249 establishes — a policy this repository adopted on 2026-08-27, not a description of anything running today. Six role-shaped entry paths sit on top, so any host becomes a reliable team member — without locking you to a single model or vendor.

What's different

It is both deep and disciplined — and honest about what it deliberately is not:

  • Depth that routes itself — a deep skill and command library, with a capability router that loads the right one on intent, not a 500-artefact context dump.

  • Governance on every host — rules compiled into each tool's native format at projection time; deterministic runtime hooks added on hook-capable hosts. This config-space, host-agnostic governance is the moat (the governance advantage · enforcement by host).

  • Surgical uninstall — removes only its own keys from a shared host config (matched by JSON-pointer + SHA-256), never a neighbour tool's entries.

  • Pack-scoped install — writes the active pack only, not a 500-artefact dump.

What it deliberately is not — the core is a governance layer with optional, individually opt-in embedded engines (code intelligence, gated reach, the setup GUI, the bench lab — ADR-124): no mandatory or always-on daemon, no separate state database, no self-rewriting memory, no auto-build pipeline. Engines are never mandatory, never default-on without measured lift, and terminate with the command that invoked them. The host agent runs the loop; every learned change is human-reviewed; the same layer stays portable across tools. Capability without a process to babysit.

Where this comes from (honest provenance). The skills, rules and personas are distilled from real production work on TypeScript and PHP codebases. The governance mechanics are stack-agnostic, but the domain heuristics are richest where they were forged — treat coverage on other stacks as promising, not proven and tell us where it falls short.

See exactly what works on which host or jump to things you can do in a minute.

Pick your profile — six entry paths

agent-config setup writes profile.id to .agent-settings.yml; each anchor below is the first-screen the wizard sends you to. One README, six entries, no role-detection guesswork.

Profile (profile.id)

Audience

First commands

First skills

👩‍💻 developer

IC engineer

/implement-ticket · /work · /review-changes · /fix · /commit

developer-like-execution · verify-completion-evidence · minimal-safe-diff · systematic-debugging · test-driven-development

✍️ content_creator

Writers, ghostwriters, marketers

/work · /post-as · /ghostwriter · /optimize-prompt · /video:from-script · /video:storyboard

voice-and-tone-design · messaging-architecture · editorial-calendar · release-comms · character-consistency

🚀 founder

Solo / early-stage founder

/work · /feature · /challenge-me · /council

refine-prompt · rice-prioritization · vision-articulation · fundraising-narrative · runway-cognition

🏛 agency

Multi-client delivery shop

/work · /implement-ticket · /refine-ticket · /feature · /roadmap

doc-coauthoring · decision-record · refine-ticket · estimate-ticket · perf-feedback-craft

💼 finance

CFO / fractional finance / FP&A

/work · /council · /challenge-me

dcf-modeling · forecasting · scenario-modeling · unit-economics-modeling · runway-cognition

🛡 ops

RevOps, support, SRE-adjacent

/work · /threat-model · /review-changes · /fix

incident-commander · dashboard-design · logging-monitoring · threat-modeling · launch-readiness

Not sure which one? Run npx @event4u/agent-config init then agent-config setup — the browser wizard asks a single 8-option role question and maps to the closest profile. Source-of-truth: src/agent-src/profiles/ · schema: docs/contracts/profile-system.md. Beyond software: user-types/ (galabau · metalworking · truck — see Beyond software.

Per-profile experience pages (who it's for · first tasks · packs + flows · what is not loaded · examples): developer · content_creator · founder · agency · finance · ops.

Workflows, not raw commands

You don't memorize every command — you run a work journey. Four flows span the developer story end-to-end; each names the command you TYPE to start and the skills it composes:

Flow

Start with

The journey

🔍 Discovery

/feature:plan · /research

explore → plan → estimate → refine, before building

🔨 Implementation

/work · /implement-ticket

plan → implement → verify → commit

🔎 Review

/review-changes · /judge

self-review → judge → quality-fix → threat-model

🚢 Delivery

/commit · /pr:create

commit in chunks → open PR → answer review

Full detail — entry commands, canonical path, composed skills per flow: docs/flows.md. (agent-admin — memory / analytics / config — is platform operation, not a user-work flow.)


Creative Pack — cinematic AI video. script → character-locked image → motion+audio prompt → provider render → stitched clip, with AIV_DRYRUN=true as the cost-safety default. A first-class capability inside the content / creator experience — no longer the package's headline. See /video:from-script.

Legal Pack — not legal advice. The EU/DE legal pack (contract/NDA/DPA review, triage) is a research-and-drafting aid only — it does not provide legal advice, does not replace a qualified lawyer and must not be relied on for any concrete matter. It produces general information and general templates, never individual-case examination. Read LEGAL_NOTICE.md before use.

Full catalog — every skill, rule, command, guideline: docs/catalog.md. The headline is the experience (profile + packs) and the depth behind it.

Use it in your project

Run from a consumer repo — bootstrap via npx, the agent picks up your stack and you ship work end-to-end. New install? Start with the Quickstart. Already installed? Supported tools shows the wired AIs; docs/featured-commands.md lists the end-to-end workflows (/implement-ticket, /work, /commit, /pr:create). Deeper tour: 2-minute demo.

Install scope. Pick one scope per machine — project-local (default, recommended for application repos) or user-global (recommended for tooling repos / dotfiles). The installer refuses a second, conflicting scope via the scope_guard pre-flight. Details: docs/contracts/install-scopes.md. Cleanup when needed: bash src/scripts/cleanup_other_scope.sh --confirm.

Related MCP server: Agent Module

Prove it

Don't take the claims on trust — verify them. docs/proof.md is generated from source: a claim→evidence table (every public claim binds to a resolvable pointer or CI fails), honest-null benchmarks including the runs where the package changed nothing and a "verify it yourself" block you run on a fresh checkout. The proof page fails CI if it drifts from its sources — reproducibility is the proof. Browse it on the deployed docs site — the proof page is the primary entry: event4u-app.github.io/agent-config/proof/. The honest comparison frame lives at docs/us-vs-the-category.md. Freshest measured row: in one post-fix session, advisory context injection cut language-mirror violations 555 → 19 while the two blocking guards went 8 → 0 and 1 → 0 — advisory reduced massively, only blocking eliminated. One session and a post-hoc reading, so a recorded prior rather than a law.

Maintaining a skills catalog yourself? The anti-reskin gate that blocks find-replace re-skin PRs here runs on your catalog too — docs/anti-reskin-gate.md.

Audit-disciplined by construction — every memory consult, decision key and hook concern lands in agents/runtime/state/ so you can replay it. Core principles names the four invariants; What agent-config is — and what it isn't draws the scope boundary.

Contribute

Working on the package itself? Development covers the task ci pipeline, Requirements the toolchain, Maintainer telemetry the opt-in measurement loop. Source-of-truth tree is src/ (src/skills, src/rules, src/agent-src/); never hand-edit .augment/ or dist/agent-src/.

Security. Disclosure policy: SECURITY.md. Threat model: docs/threat-model.md.


Quickstart

Try one thing in 30 seconds — before the full suite, drop in a single self-contained subagent and see the discipline on your own repo:

mkdir -p .claude/agents
curl -fsSL https://raw.githubusercontent.com/event4u-app/agent-config/main/docs/wedge/production-validator/production-validator.md \
  -o .claude/agents/production-validator.md
# then in Claude Code:  @production-validator check this branch is actually done

production-validator is read-only and installs nothing else — it gates "done" by hunting mocks/stubs on the shipped path and demanding real-system evidence (what it does). Like it? Install the full suite:

One command. Detection-driven — your installed AI tools are found and pre-selected. Nothing is written until you click Finish. No YAML by hand.

Those four are structural: they hold on every run, because they are properties of the code path rather than of your machine. How long it takes is not one of them — that is dominated by network and registry latency, which we do not control. The install → doctor wall-clock is measured by CI on every umbrella run and published with its conditions, as evidence; it is never a promised number.

# 1. Install — on a terminal with a display, the browser wizard launches
#    automatically; the same TypeScript installer runs the real install behind it.
npx -y @event4u/agent-config init

# 2. Pick your profile + tools in the wizard, click Finish.
#    (Writes ~/.event4u/agent-config/, ~/.claude/, ~/.cursor/, …)

# 3. First real task — agent refines, plans, verifies.
/work "your first real task"

Headless / CI: init skips the GUI automatically on CI, on a non-TTY, on a headless host, and whenever any CLI-mode flag is present — it then runs the non-interactive installer directly. The full opt-out set is listed once, against the code, in gui-wizard § When the GUI is skipped. Pass flags (--profile=balanced --tools=claude-code,cursor); add --dry-run to preview writes. The GUI and the CLI share one installer (src/scripts/install.ts), so both produce identical results. Reference: docs/wizard.md.

Pick specific AIs: --tools=claude-code,cursor,augment,windsurf,cline,gemini-cli,copilot,roocode,aider,codex,claude-desktop,continue (any subset). Visual picker: add --gui (loopback-bound, CSRF-gated; contract gui-wizard). --gui is an opt-in that forces the wizard past the TTY and headless checks — it does not override CI, AGENT_CONFIG_NO_UI, or a CLI-mode flag; combining it with one of those exits non-zero rather than quietly running the CLI install. On a headless host add --allow-headless and connect a browser to the printed URL.

Verify hook coverage: npx @event4u/agent-config hooks:status prints the per-platform matrix (--strict for CI, --format json for tooling).

Scope (v2.5+): init writes global only — ~/.event4u/agent-config/, ~/.claude/, ~/.cursor/, …. The project tree gets agents/overrides/ only (the bridge marker was retired — ADR-020 amendment 2026-07-13; the global root resolves from ~/.event4u/agent-config). --project is maintainer-only behind AGENT_CONFIG_DEV_MODE=1 (ADR-020, dev-mode).

Migrating from a v1.x install? npx @event4u/agent-config migrate — full notes in docs/migration/v1-to-v2.md.


What agent-config is — and what it isn't

A content layer — skills, rules, commands, guidelines, personas — distributed via npm and projected into every supported AI tool's native config format. It follows the Agent Skills open standard.

It is not an agent runtime. The agent loop, the LLM dispatcher and tool orchestration stay with the host tool (Claude Code, Augment, Cursor, Cline, Windsurf, Gemini CLI, Copilot). Think of it as a playbook and style guide for those tools — not a replacement.

In scope

Out of scope

Skills, rules, commands, guidelines, personas

Agent loop / LLM dispatcher

Multi-tool projection + condensation pipeline

Execution engine inside the package

Memory helpers (memory-add, memory-promote)

Cross-tool observability dashboard

Linters, CI, frontmatter validation against JSON-Schema (contract)

Runtime GUI / web dashboard

Skill orchestration via citations + deterministic helpers

Opinionated automatic skill-resolver (ML / relevance ranking that decides for you)

User-driven projection-time filtering by profile + packs (ADR-040)

A runtime resolver / daemon (mid-session switching — conditional, post-6.0.0)

What your agent is asked to do

Default behavior

With agent-config

Guess and edit blindly

Analyze code before changing it

Drift from project conventions

Follow detected stack conventions

Skip or invent tests

Write tests in the project's framework

Generic commit messages

Conventional Commits with scope + ticket links

Skip quality checks

Run the project's quality pipeline and fix reported errors

Open PRs without context

Structured PR descriptions from Jira / Linear / GitHub

Claim "done" without proof

Verify with real execution before claiming done


2-minute demo — /implement-ticket

The flagship command. Drives a ticket end-to-end through a fixed linear flow — and blocks on ambiguity instead of guessing.

/implement-ticket PROJ-123

The agent runs this sequence:

refine → memory → analyze → plan → implement → test → verify → report
  • Refines the ticket if acceptance criteria are vague.

  • Queries memory for past decisions, invariants, incidents.

  • Plans the change; you confirm before any file is touched.

  • Implements under minimal-safe-diff + scope-control — no drive-by edits.

  • Tests (targeted first, full suite on success).

  • Reviews the diff through four judges (bugs, security, tests, code quality).

  • Reports changes, verdicts, follow-ups — then stops. /commit and /pr:create are suggestions, never auto-run.

Any ambiguity halts the flow with numbered options — never a silent guess. Persona comes from .agent-settings.yml (roles.active_role): senior-engineer (default), qa or advisory (plan-only).

Command reference · Flow contract

Sibling — /work (free-form prompt)

Same engine, no ticket required:

/work add a CSV export endpoint to the audit-log controller

The first pass scores the prompt on five dimensions and routes on the band:

Band

Score

Action

high

≥ 0.8

Silent proceed — AC + assumptions in the report

medium

0.5–0.79

Halts with assumptions report; confirm or edit

low

< 0.5

Halts with one clarifying question on the weakest dimension

After the band gate, the flow is identical to /implement-ticket. Free-form goal → /work; ticket payload → /implement-ticket.

Command reference · refine-prompt skill

After the run: agent-config explain last reconstructs the trace (route · memory · council · halts · provider) — read-only, PII-scrubbed, offline. Docs

Product UI track

UI-shaped work routes to one of three directive sets — ui (full audit→design→apply→review→polish→report), ui-trivial (≤ 1 file, ≤ 5 lines: apply→test→report), mixed (backend + UI: contract→ui→stitch). Existing-UI audit is a hard gate (ui-audit-gate); polish has a 2-round ceiling with a11y precedence. Stack detection → blade-livewire-flux / react-shadcn / vue / plain.

Mental model (1 page) · Flow contract


Customize

Profiles — how much governance gets loaded

Safety floor (non-destructive defaults · ask-before-guessing · mirror-the-user's-language) ships in every profile. What changes is how much extra coaching gets pulled in.

Profile

What you get

When to pick it

minimal

Non-negotiable safety floor only. Cheapest, fastest.

Quick questions · throw-away scripts · CI · tight token budgets

balanced (default)

Safety floor + everyday coaching (sensible defaults, review nudges, common pitfalls).

Day-to-day work

full

Everything, including long-tail rules normally only maintainers need.

Working on agent-config itself · audits · max-fidelity demos

Under the hood: kernel-only · kernel + tier-1 · kernel + tier-1 + tier-2. Details: rule-router · kernel-membership · Configure →.

Stability: STABILITY.md for the full matrix. Work Engine (/work + /implement-ticket): beta. Runtime Dispatcher: stable. Tool Adapters: experimental (full profile only).

.agent-user.md and Ghostwriter — voice primitives

Primitive

Voice

Disclosure

personas/*.md

Review-lens (internal critique)

n/a

.agent-user.md (project root, gitignored)

The maintainer's own voice — /post-as:me

None (you are the author)

agents/reference/ghostwriter/<slug>.md (gitignored)

Documented public figure — /post-as:ghostwriter

Mandatory, non-removable footer

Create the user file interactively: /agents user init (schema). Ghostwriter cluster: /ghostwriter:fetch <url-or-name> runs an attestation gate; private individuals rejected; paywalled / leaked / DM content banned at the schema level.

Self-hosted MCP on Cloudflare — zero local install

Skills, commands, rules and guidelines can be served as an MCP endpoint from your own Cloudflare Worker — any MCP client (Claude Desktop, Claude Code, Cursor, Zed, Continue, hosted agents) talks to it over HTTP. Two auth modes: public (default, OSS read-only deploys) and bearer-auth (operator opt-in, MCP-Token Wrangler secret).

task mcp:cloud:login         # one-time, opens browser
task mcp:cloud:setup         # check → r2-create → r2-verify → whoami
task mcp:cloud:secret-put    # opt in to bearer-auth (recommended for private deploys)

→ Operator walkthrough: mcp-cloud-setup · Per-client config: mcp-client-config · Endpoints: mcp-cloud-endpoints.

Scope — Lite, not Full. The Worker serves read-only governance (skills · commands · rules · guidelines · contexts) as MCP prompts and resources, plus small read-only tools (memory_lookup, chat_history_read, list_*). It does not execute the repository's local scripts (linters, audits, task ci, work-engine hooks) — those require local install per Quickstart.

The built-in local stdio server is listed for discovery in the Glama MCP Registry (agent developers / contributors; requires a local checkout, not a turnkey install — see ADR-067).

Deployment posture

Shape

Status

Path

Single-user workspace

✅ today

npx @event4u/agent-config init — single machine, single user; no remote sync

Small team (3–10 people)

✅ today

Shared agents/overrides/ Git repo + shared NAS for knowledge — no code change, no new server. Recipe: docs/deploy/small-team-recipe.md

Organization mode (SSO · central policy · team context · internal connectors)

⏸ not started

Each shape gated on a recruited customer + funded audit + maintainer ADR. Posture rationale: docs/deploy/team-deployment-posture.md

The Hard Floor on organization-mode features (SSO, central policy, OAuth connectors, team-context) is preserved by design — they stay cancelled until a real first customer + funded security audit lifts them. The small-team recipe is the supported path in the meantime.

The 9.3/10 feedback round (2026-05-25) re-asked for OAuth knowledge connectors, IAM / org governance and organization-shared memory. Each is a stable cancellation row in team-deployment-posture under the same three release gates — recruited team customer · funded audit · maintainer ADR.


Harness expectations

Three classes of install/runtime behaviour look like package bugs but are host-harness behaviour the package cannot control — sibling-plugin namespaces (codex:*, cc-gemini-plugin:*), deferred tools surfaced via ToolSearch and cross-scope skill drift (real bug, fixed in the distribution-channels track). Diagnostics + the package's response: docs/contracts/harness-expectations.md. First step when a skill appears twice: task probe:skills.

Supported tools

Project-installed (npx)

Tool

Rules

Skills

Commands

How it works

Claude Code

Reads .claude/

Cursor

☑️

Reads .cursor/rules/ + commands via AGENTS.md

Cline

☑️

Reads .clinerules/ + commands via AGENTS.md

Windsurf

☑️

Reads .windsurfrules + commands via AGENTS.md

Gemini CLI

☑️

Reads GEMINI.md

GitHub Copilot

☑️

Reads .github/copilot-instructions.md

Roo Code

☑️

Auto-discovers .roo/rules/*.md + AGENTS.md

Codex CLI

☑️

Auto-discovers AGENTS.md

Continue.dev

☑️

Auto-discovers .continue/rules/*.md + AGENTS.md

Aider

📌

Manual read: in .aider.conf.yml

Augment (VSCode/IntelliJ)

📌

Global-only; project writes marker

Claude Desktop

📌

Global-only

✅ native &nbsp; ☑️ text reference (in AGENTS.md, not invokable as native slash-command) &nbsp; 📌 marker only &nbsp; — not available

Team reproducibility: every tool you init is recorded in agents/installed-tools.lock (committed, machine-managed). New team members run npx @event4u/agent-config sync after cloning; CI gates drift with agent-config validate. Schema: installed-tools-manifest.

Plugin-installed (optional, global)

Tool

Install

Augment CLI · Copilot CLI

Install → — rules + skills + commands, marketplace-updated

Claude Code: the marketplace plugin is deprecated (single-surface model). The npx/npm file projection now carries content and the deterministic hooks (registered in a managed ~/.claude/settings.json block by agent-config global / upgrade), so the plugin only duplicates skill/command listings while its git-SHA snapshot rots silently. Existing installs: claude plugin uninstall agent-config@event4u-agent-configagent-config doctor flags the duplicate surface.

Keep the global install current with agent-config upgrade (latest) or agent-config refresh --global (same-version re-install); agent-config doctor flags a missing-from-PATH binary or broken hook wiring. See getting-started § Keeping current · Troubleshooting.

The command surface at a glance

Command

What it does

agent-config init

One-shot install — opens the browser wizard (recommended path or step-by-step)

agent-config init --project

Initialize a project: minimal agents/ bridge + managed .gitignore block

agent-config config

Open the configuration GUI — global settings hub (simple + advanced tiers, search, reset-to-default)

agent-config config --project

Open the project configuration surface

agent-config setup

Re-run the guided onboarding wizard (prefilled from your current state)

agent-config upgrade

Update the global install to the latest release + additively sync settings

agent-config doctor

Read-only health/drift report

Cloud / Hosted-agent surfaces

For platforms where the package's scripts cannot run, artefacts are built for paste-in or upload:

  • Linear AI (Codegen, Charlie, …) — dist/linear/{workspace,team,personal}.md

  • Claude.ai Web Skillsdist/cloud/<skill>.zip

Install →


Works with agent-switch

agent-switch is the companion CLI for running several agent accounts on one machine: it isolates each account in its own profile (CLAUDE_CONFIG_DIR per profile), so switching accounts never means logging out and back in. The two compose — agent-switch isolates the accounts, agent-config governs what the agents do inside them. When agent-config runs under an agent-switch profile it says so in the settings hub, warns before writes that would land in a shared (cross-profile) tree, and accepts a host-supplied config root so its own settings stay profile-scoped.

How the two compose →


Who this is for

Stack-agnostic governance core (orchestration · role modes · command clusters · quality gates · audit-discipline) plus parallel stack-specific skill sets:

Stack

Coverage

Laravel · modern PHP (deepest)

Pest · PHPStan · Rector · ECS · Eloquent · Livewire/Flux · Horizon · Pulse · Reverb · Pennant

Symfony

symfony-workflow (DI · Doctrine · Messenger · voters · Twig) + project-analysis

Next.js App Router

nextjs-patterns (RSC · Server Actions · caching · route handlers) + UI react-shadcn

Zend / Laminas

project-analysis + shared PHP coder/quality skills

React · Node / Express

project-analysis + UI react-shadcn

Vue · plain HTML

UI directive set (vue / plain)

Cross-stack

API design · testing · security · database · Docker · Git · CI · review · threat modeling · observability

Beyond software

The same orchestration core drives non-software trades via user-types/: galabau-field-crew · metalworking-shop · truck-driver. Contribute your own — 5-minute scaffold.


Data governance & domain safety

Three domain-safety rules (domain-safety-pii, domain-safety-disclaimer, domain-safety-retention) act as per-domain output floors across ~12 areas — PII redaction (support / finance / recruiting / marketing), advice disclaimers (legal / financial / medical / consulting), retention guidance (finance / support), ops floors (logging / export). Full surface → rule → floor matrix: docs/safety.md. Beta contracts: memory-visibility-v1 · decision-trace-v1.

Code provenance & license governance

Every diff is checked against a license policy derived from the target repo's own detected license (LICENSE/package.json/composer.json, precedence-ordered; sources disagree → escalate, never guess) and a strict linter over our own borrow ledger (provenance/borrows.jsonldocs/THIRD-PARTY-NOTICES.md) that fails a deny-class license, an unknown license, a missing transformation note or a rename-only-phrased one — wired into ci/ci-strict from day one. A third piece, license-compliance-audit, runs an offline/online similarity scan on demand — a human invokes it deliberately, never a pipeline. This is provenance-governed, license-policy-enforced borrow discipline backed by an audited borrow trail — not a copy detector.

Scope & limits

  • Unconscious training-data reproduction is not detectable at this layer. No tool here — or anywhere — can see what a model's training data contained; this system governs what gets consciously borrowed and recorded, never what a model silently recalls.

  • Detection, where it exists, covers a knowledge base of known OSS only — a subset of all code that has ever existed, never a model's training corpus.

  • No CI-facing detection gate exists. A deterministic scanner (jscpd offline + SCANOSS online) was built and measured against a frozen synthetic corpus, but missed its own pre-registered thresholds (measured: recall 12/16, false positives 2/12, SCANOSS rename-only recall 0/8) — see docs/CLAIMS.md. It ships in no form in CI, not even advisory — only as the on-demand skill above.

  • Rename-only laundering is not detected by anything we ship or evaluated. The ledger's transformation-note check rejects a rename-only-phrased note, but it cannot catch an undisclosed rename-only copy that was never logged.

Reduces and documents risk — never eliminates it.

Maintainer telemetry (opt-in, default-off)

Local-only artefact-engagement log. Set telemetry.artifact_engagement.enabled: true in .agent-settings.yml. Records which skills / rules / commands / guidelines the agent consults during /implement-ticket / /work. JSONL under the project root, nothing uploaded. Reports: npx @event4u/agent-config telemetry:report.

Context-aware command suggestion

When a prompt matches a command's purpose ("setze ticket ABC-123 um" → /implement-ticket), the agent surfaces matches as numbered options — nothing auto-executes. Per-conversation off: /command-suggestion-off. Settings: commands.suggestion.{enabled,blocklist,confidence_floor} in .agent-settings.yml.


Core principles

  • Analyze before implementing — no guessing, no blind edits

  • Verify with real execution — no "should work"

  • Challenge to improve — agents are thought partners, not yes-machines

  • Strict by design — quality over flexibility

  • Zero overhead by default — nothing runs until you ask for it


Documentation

Document

Content

Getting Started

First run, 3-test experience, profiles, next steps

Installation

All install paths, Composer/npm, orchestrator details

Architecture

System layers, content pipeline, tool support matrix

Customization

Overrides, AGENTS.md, agent settings, cost profiles

Quality & CI

Linting, CI pipeline, condensation system

Migration

Per-version upgrade steps

Showcase

More examples & expected behavior

Browse content: all commands · skills catalog · full catalog · llms.txt.


Troubleshooting

First stop for any install problem: agent-config doctor — it flags a missing-from-PATH binary, binary↔plugin version drift, stale orphans and manifest issues, each with a one-line fix hint.

For "why didn't rule/hook X fire?" questions: agent-config routing:doctor — a read-only, live diagnosis that reports every session-start gate as ACTIVE/INACTIVE with the concern's own reason (e.g. session-canary: ACTIVE for "Alex" vs INACTIVE — no name on any settings layer), the platform's concern chain, host hook registration, and router + projection freshness. Deeper hook internals (fail-open/closed posture, last dispatcher feedback per concern): agent-config hooks:doctor.

A new command / skill is missing in Claude Code after an upgrade

Under the single-surface model, agent-config upgrade refreshes the ~/.claude/ file projection — that IS the content surface, so a fresh session picks the new commands up directly. If commands are still missing, the usual cause is a leftover marketplace plugin: it is a git-SHA snapshot that never moves with the npm upgrade and it shadows nothing — it just lists everything twice while lagging behind. Remove it:

claude plugin uninstall agent-config@event4u-agent-config

Then start a new Claude Code session. agent-config doctor reports a leftover plugin as claude-plugin: duplicate surface; hooks are unaffected (they live in a managed ~/.claude/settings.json block — verify with the hook-wiring check).

Skills / commands appear twice in Claude Code

Same cause as above: the deprecated marketplace plugin is installed next to the ~/.claude/ file projection, so every skill lists plain and agent-config:-prefixed. Uninstall the plugin (command above) and start a new session.

agent-config upgrade fails with Unknown argument: --no-ui

Known bug in 8.2.0: upgrade passed a --no-ui flag that the install orchestrator did not accept yet, so the run aborted early. Fixed on main; until the next release, work around it with:

AGENT_CONFIG_NO_UI=1 agent-config global   # refresh the global install, no wizard

Upgrade was interrupted (Ctrl-C, wizard closed, step failed)

Only the initial npm install -g hard-aborts an upgrade. Every later step (global re-deploy with hook registration, settings sync, wrapper + git-hook refresh) runs independently — a single failed step is reported in the end-of-run summary instead of silently skipping the rest. Re-run agent-config upgrade to converge and use agent-config doctor to name anything left in a mixed state.

agent-config: command not found / hooks stopped firing

Runtime hooks resolve the global binary on PATH — a project-local install alone is not enough for them. Reinstall the binary:

npm install -g @event4u/agent-config
agent-config doctor   # verifies PATH + plugin wiring

Project files look stale after a package update

Project-local projections are only rewritten on an explicit refresh:

agent-config refresh             # re-apply the installed version to this project
agent-config refresh --global    # same-version re-install of the global root

More per-version steps: Migration · getting-started § Keeping current.


Development

Working on the package itself? Edit src/ (the source of truth — src/skills, src/rules, src/agent-src/), regenerate trees:

task sync             # regenerate dist/agent-src/ and .augment/
task generate-tools   # regenerate .claude/, .cursor/, .clinerules/, .windsurfrules
task ci               # full pipeline — green before PR
task test             # unit + integration tests
task dev:setup        # boot the onboarding wizard against the working tree

Invoking the CLI from a source checkout: ./agent-config <command> (the maintainer shim at the repo root → scripts/agent-configdist/cli/agent-config.js). npx @event4u/agent-config doesn't resolve in the source repo without a prior npm link, since there's no node_modules/.bin/agent-config symlink — use ./agent-config instead. Build the TS binary with npm run build:cli if dist/cli/agent-config.js is missing.

→ Full project structure and commands: docs/development.md · CONTRIBUTING.md. Stack: TypeScript throughout — CLI, UI, and the build / lint scripts. MCP registry payloads render under dist/mcp/ (submission checklist).


Requirements

  • Node ≥ 20.11npx @event4u/agent-config init is the canonical install path. No Python anywhere on the install path (the Python installer retired with the TypeScript migration).

  • Platform: macOS 12.3+, Linux, WSL2. Git Bash needs Developer Mode for symlinks. Contributors rebuilding .augment/ also need Task.

Windows

Native PowerShell / cmd is not supported for the file install — use WSL2 for the full installed tree. The supported native-Windows surface is the MCP stdio server: point any MCP client at

npx -y @event4u/agent-config mcp-server

and the governance content (prompts, resources, tools) is available without the file install. Porting the bash dispatcher to native Windows is demand-gated: a named Windows adopter who cannot use WSL2 or the MCP path reopens it (see agents/roadmaps/ — road-to-credible-install Phase 3).

Funding

The package is free, MIT, and stays that way — no paid tier, no dual licensing. If it saves you time and you want to chip in, the GitHub Sponsor button at the top of the repo is the whole mechanism. If you would rather not, use it anyway; nothing here is gated on it.

License

MIT.

mcp-name: io.github.event4u-app/agent-config

Available Tools

20 tools
capabilities_indexA

Regenerate CAPABILITIES.yaml, the package's coverage index of skills, rules, commands, and guidelines. Use after adding or removing an artifact to keep the index current. Pass check: true to run in read-only CI mode (fails instead of writing on drift).

ParametersJSON Schema
NameRequiredDescriptionDefault
checkNoWhen true, run in read-only drift-check mode instead of writing the index.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the tool writes/regenerates the index, and importantly explains the read-only CI behavior ('fails instead of writing on drift'). This complements the readOnlyHint: false annotation by giving concrete failure semantics, though it does not detail side effects or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and resource, then usage timing, then the parameter mode. Every sentence earns its place with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter write tool with no output schema, the description fully covers what the tool does, when to use it, and how the parameter changes behavior. The inclusion of failure mode in CI makes it operationally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully describes the 'check' parameter with high coverage, and the description adds practical context by explaining the CI-mode behavior. This reinforces the schema without redundancy, raising it above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Regenerate') and a clear resource ('CAPABILITIES.yaml'), identifying it as the package's coverage index. This distinguishes it from sibling tools like list_skills or list_rules, which query rather than update the index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use after adding or removing an artifact to keep the index current,' which is clear guidance for when to invoke it. It also explains the CI-mode use case with 'check: true', but does not name alternative tools to use instead of this one, such as list_* commands.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_history_appendA

Append one structured entry to the consumer project's chat-history log (a JSONL file). Use to record a decision, note, or phase marker that should persist into a later session or be distilled by mine_session. Writes to the filesystem (agents/runtime/.agent-chat-history by default; agents/.agent-chat-history and .agent-chat-history accepted for back-compat) and returns the written entry plus its resolved target path. Path-scoped: a path outside the allowlist, or any traversal escaping the project root, raises an error before writing. Set dry_run: true to preview the entry and target path without touching disk.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoOptional path override. Must resolve to `agents/runtime/.agent-chat-history` (current default), `agents/.agent-chat-history`, or `.agent-chat-history` under consumer_root.
textYesThe entry body to record.
dry_runNoWhen true, return the entry and resolved target path without writing to disk.
sessionNoOptional 16-char session id to group the entry under. Defaults to the current session.
entry_typeNoShort ``t`` tag categorising the entry (e.g. note, decision, phase). Defaults to ``note``.
min_schema_versionNoRefuse to write if the on-disk history schema is older than this version.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only readOnlyHint: false, annotations are minimal. The description fully discloses the write behavior, filesystem target, return value ('returns the written entry plus its resolved target path'), safety validation ('path outside the allowlist... raises an error before writing'), and the dry_run side-effect-free preview. This goes well beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, front-loaded with the primary action, then usage context, then safety/behavior specifics. Every sentence adds critical information: what it does, when to use it, the write behavior and safety, and the dry_run escape hatch. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter write tool with no output schema and minimal annotations, this description is remarkably complete. It covers the target file locations, the return payload, path restrictions, dry_run behavior, and downstream consumers. An agent can confidently select and invoke this tool without further clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds meaningful behavioral semantics for the 'path' parameter (allowlist and traversal checks) and clarifies dry_run's purpose. It also explains the entry body ('text') in the context of a JSONL log, enhancing understanding beyond the schema's bare field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Append one structured entry to the consumer project's chat-history log (a JSONL file).' This is a specific verb+resource that distinguishes it from siblings like chat_history_read, which retrieves the log, and other memory tools that operate on different stores.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use to record a decision, note, or phase marker that should persist into a later session or be distilled by `mine_session`.' This gives clear context for when to use the tool. It also notes back-compat paths and the dry_run option, but it does not explicitly state when not to use it (e.g., for ephemeral or non-history data), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_history_readA
Read-only

Read recent entries back from the consumer project's chat-history JSONL (agents/runtime/.agent-chat-history; agents/.agent-chat-history accepted for back-compat). Use to recover context from an earlier session — decisions, notes, phase markers — at the start of a new task. Read-only. Returns the resolved file path plus a list of matching entries (newest last). Combine session, last, and entry_type to narrow the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
lastNoReturn only the most recent N entries, after other filters apply.
pathNoOptional history-file path override; defaults to the standard chat-history location under the project root.
aroundNoTimeline anchor: return the entries around this ref (from a detail:"index" row) instead of the filtered list — depth_before/depth_after neighbours plus the anchor. Refs are within-file ordinals; rotation invalidates them.
detailNo'index' returns compact rows (ref, t tag, ~100-char preview, tokens_estimate) instead of full entries — scan first, then re-read the refs you need via `around` or a narrowed filter. History entries are large, so index mode pays off here. 'full' (default) returns complete entries.full
sessionNoFilter to a single 16-char session id.
entry_typeNoFilter by the `t` tag (e.g. note, decision, phase).
depth_afterNoNeighbours after the anchor (with `around`). Defaults to 3.
depth_beforeNoNeighbours before the anchor (with `around`). Defaults to 3.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, it adds valuable context: returns the resolved file path plus matching entries (newest last), describes path resolution with back-compat, warns that rotation invalidates refs, and notes that history entries are large so 'index' mode is helpful. These details give the agent a clear behavioral model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at four sentences, each serving a distinct purpose: what it reads, when to use it, return shape, and filter guidance. No redundant text or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 optional parameters and no output schema, the description covers the core aspects: purpose, file location, filtering hints, and return format. It lacks explicit mention of error conditions or edge cases, but the schema and rich parameter descriptions fill the gap reasonably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds a useful tip about combining 'session', 'last', and 'entry_type' filters, but it doesn't further explain each parameter or add semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Read' with a concrete resource ('chat-history JSONL') and gives exact paths, making the tool's function unambiguous. It differentiates from sibling 'chat_history_append' by emphasizing read-only access and mentions return contents (entries, resolved path).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the intended use case: recovering context from earlier sessions at the start of a new task. It also advises combining filters to narrow results, but it doesn't explicitly mention when not to use it or name alternative tools (e.g., memory_get) for different context sources.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

conformance_checkA
Read-only

Run the consumer conformance contract (doctor --ci plus installed-and-firing checks) and return pass/fail per check. Use to verify a consumer project's install is fully wired before relying on it. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include readOnlyHint=true, and the description reiterates 'Read-only,' with no contradiction. It adds behavioral detail by naming the commands executed and stating that results are returned per check, going beyond the annotation's simple safety flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action, command details, and output type, then a clear usage recommendation. Every word earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter tool with no output schema, the description adequately covers the operation, return semantics (pass/fail per check), and when to use it. There are no essential gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so schema coverage is trivially 100%. No parameter information is needed, and the baseline for zero-parameter tools is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource combination ('Run the consumer conformance contract') and clearly identifies the underlying command (`doctor --ci` plus installed-and-firing checks) and output (pass/fail per check). It distinguishes this from sibling tools like `doctor_report` or `run_tests` by emphasizing verification of a consumer project's install wiring.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit usage scenario: 'Use to verify a consumer project's install is fully wired before relying on it.' It does not mention when not to use the tool or alternatives, but the context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

council_estimateA
Read-only

Estimate the token cost of an AI-council debate over a given input (roadmap, diff, prompt, or file set) without spending — no network call, no billing. Use before deciding whether to authorize a real council run. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoCouncil depth tier to estimate for. Defaults to the configured default depth.
input_pathYesPath to the roadmap, diff, or file to estimate council cost for.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation already marks it as read-only, and the description adds meaningful behavioral context beyond that: 'no network call, no billing' and 'without spending'. This clarifies the side-effect-free nature in a way the annotation alone does not convey. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences and packs in purpose, safety, and timing without any filler. It is front-loaded with the action verb and resource, making it immediately understandable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simple read-only nature and the rich annotations and schema, the description adequately covers purpose, usage timing, and safety. It does not describe the return format or precision of the estimate, and since there is no output schema, that is a minor gap, but not enough to lower the score further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes both parameters, so schema coverage is 100%, giving a baseline of 3. The description adds value by specifying that the input can be a 'roadmap, diff, prompt, or file set', which expands on the schema's narrower 'roadmap, diff, or file' phrasing. Depth is not mentioned, but the schema's enum covers it sufficiently.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Estimate the token cost') and the resource ('AI-council debate over a given input'). It also distinguishes itself from a real council run by emphasizing 'without spending — no network call, no billing', leaving no ambiguity about its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Use before deciding whether to authorize a real council run' explicitly defines when to use this tool. While it provides a clear context, it does not explicitly mention when not to use it or name an alternative tool, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

doctor_reportA
Read-only

Run the consumer-project doctor diagnostic and return a structured health report (install drift, hook wiring, settings schema, discovery manifest freshness). Use to triage a misbehaving install. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already declares the tool is read-only, and the description reinforces this with 'Read-only.' Beyond that, it adds useful behavioral context about the nature of the diagnostic (checks for drift, hook wiring, settings schema, freshness) and the structured form of the output, which is not fully derivable from the annotation alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the action and resource, and contains zero filler or redundant content aside from the harmless repetition of 'Read-only.' Every clause adds information about the tool's purpose, output, or usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only diagnostic tool with no output schema, the description provides all necessary context: what it runs, what the report covers, why you would use it, and its safety profile. The annotations and schema further support completeness, leaving no significant gaps for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline of 4 applies per the rubric. The schema is fully covered (100%) and the description does not need to clarify any parameter details because none exist. The description appropriately focuses on behavior and use case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Run') and names a precise resource ('consumer-project doctor diagnostic'), while detailing the report's contents (install drift, hook wiring, settings schema, discovery manifest freshness). This clearly distinguishes it from siblings like run_tests or telemetry_report by identifying a unique diagnostic target and purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'Use to triage a misbehaving install.' This provides clear context for the appropriate use case. However, it does not mention when not to use it or suggest alternatives among the sibling tools, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lint_skillsA
Read-only

Lint skill, rule, command, guideline, and persona markdown files for frontmatter and structural errors. Use before committing or opening a PR that adds or edits any of those artifacts, to catch schema violations early. Read-only — never writes files or spawns git. Returns the scripts/skill_linter.py --format json payload: a summary object (pass / pass_with_warnings / fail / total counts) and a per-file results array with severity-tagged findings. Pass paths to lint a subset; omit for a full tree scan.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNoRepo-relative paths to lint (files or directories). Empty or missing → full tree scan via gather_all_candidate_files.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description explicitly states the tool is read-only, never writes files, never spawns git, and details the exact return payload structure (summary and per-file results). This adds meaningful behavioral context well beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, front-loading the primary purpose, then usage guidance, then safety/behavior, then return format, and finally parameter usage. Every sentence contributes valuable information with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, when to use, safety/read-only behavior, return payload, and parameter semantics. With a simple one-parameter input schema and no output schema, the description provides all necessary context for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides a thorough description for the 'paths' parameter, including the meaning of empty/missing values (full tree scan). The description merely restates this ('Pass 'paths' to lint a subset; omit for a full tree scan') without adding new information. Schema coverage is 100%, so a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (lint) and a clear resource (skill, rule, command, guideline, and persona markdown files) to state what the tool does. It distinguishes itself from siblings by being the only linting tool among the listed tools, and it mentions 'frontmatter and structural errors' as the focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit context for when to use the tool ('Use before committing or opening a PR that adds or edits any of those artifacts'), which is strong guidance. It does not explicitly name alternatives or say when not to use it, but the sibling tools are clearly unrelated in function, so the omission is minor.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_commandsA
Read-only

Enumerate every slash command the server currently exposes as a prompt, each with its name and description. Use to discover available commands before routing a user request to one. Read-only manifest view, takes no arguments. Returns a count plus a commands array.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral detail beyond the readOnlyHint annotation: it labels itself a 'Read-only manifest view,' states it takes no arguments, and reveals the response shape ('Returns a count plus a commands array'). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two succinct sentences: the first states what it does, the second gives usage context, read-only nature, argument count, and return format. Every sentence earns its place with no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no parameters and no output schema, the description fully covers purpose, usage context, behavior, and return structure. Sibling relationships are implied through the tool name and the usage note, so no further context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, and the schema coverage is 100%, so there is nothing to misinterpret. The description reinforces 'takes no arguments,' matching the empty schema. Baseline for 0 params is 4, and no further parameter explanation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Enumerate') and resource ('every slash command the server currently exposes as a prompt'), and clarifies what's included ('name and description'). This clearly distinguishes it from siblings like list_rules and list_skills, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use: 'Use to discover available commands before routing a user request to one.' This gives clear context. However, it does not mention alternatives or when not to use it (e.g., for rules or skills), so it lacks the explicit exclusion/alternative guidance of a top-tier score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_rulesA
Read-only

Enumerate every behavioral rule the server exposes as a resource, each with its URI, name, and description. Use to discover which rules are in effect, then fetch a body with read_resource_body or resources/read. Read-only manifest view, takes no arguments. Returns a count plus a rules array.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint: true. The description adds that it is a 'read-only manifest view', takes no arguments, and returns a 'count' plus 'rules' array. This offers concrete behavioral detail beyond the annotation without contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the first states purpose and result structure; the second gives usage and return format. No redundant wording, front-loaded with the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-argument list tool with no output schema, the description fully covers purpose, usage, and return shape. It even notes the next step (fetch a body), making it contextually complete relative to its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, and the description explicitly states 'takes no arguments'. Since schema coverage is 100% trivially, the description adds confirmation of the empty parameter list. Baseline for zero params is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Enumerate' and identifies a distinct resource: 'behavioral rule the server exposes as a resource'. It explicitly lists what each rule includes (URI, name, description), distinguishing it from sibling tools like list_commands and list_skills.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear usage context: 'Use to discover which rules are in effect' and explicitly names alternative tools for a follow-up action (read_resource_body or resources/read). This tells the agent when to use this tool and what to do next.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_skillsA
Read-only

Enumerate every skill the server currently exposes as a prompt, each with its name, description, and source. Use to discover which skills are available before suggesting or invoking one. Read-only manifest view, takes no arguments. Returns a count plus a skills array.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation is reinforced by 'Read-only manifest view,' and the description adds behavioral details such as taking no arguments and returning a count plus a skills array. Since there is no output schema, explaining the return shape is valuable context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The three sentences are tightly packed: definition, usage context, and behavior/return. No waste or redundancy; each clause serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only listing tool, the description covers what it returns, when to use it, and its read-only nature. Given the lack of output schema, the return value explanation is sufficient, and the tool's simplicity means no further context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, and the input schema is an empty object. The description explicitly notes 'takes no arguments,' and with no parameters to document, the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Enumerate every skill the server currently exposes as a prompt,' clearly identifying the action and resource. It distinguishes itself from sibling listing tools like list_commands and list_rules by focusing specifically on skills. The mention of name, description, and source further clarifies the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use: 'Use to discover which skills are available before suggesting or invoking one.' However, it does not mention when not to use or explicitly compare with alternative list tools, so it falls short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_getA
Read-only

Batch-fetch FULL memory entries by id — the second half of the index-first retrieval workflow. Call memory_lookup with detail:"index" first, pick the ids whose title/tokens_estimate justify the fetch, then fetch them here in ONE batched call. Unknown ids are reported per-id (ids[]="unknown"), never failing the batch. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYesEntry ids to fetch (from a detail:"index" lookup).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses that unknown ids are reported per-id as 'unknown' and never cause the whole batch to fail—critical error-handling behavior that is not visible in the schema. It also reinforces the read-only nature and highlights the batching capability, adding genuine behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the core purpose, the second gives the workflow and usage order, and the third details error behavior. There is no fluff or repetition; the description is tightly packed and front-loaded with the primary action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter, read-only tool with no output schema, the description fully covers purpose, usage, error handling, and the relationship to its sibling. The 'FULL' qualifier and the per-id unknown example give enough detail for an agent to anticipate the response shape without needing a formal output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes the 'ids' parameter as 'from a detail:"index" lookup', so baseline is 3. The description adds the selection criterion ('title/tokens_estimate justify the fetch') and emphasizes the batched nature of the call, providing marginal but meaningful extra guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Batch-fetch FULL memory entries by id') and explicitly frames the tool as the 'second half of the index-first retrieval workflow', distinguishing it from sibling memory_lookup. It clearly states what the tool does and what it returns (full entries as opposed to index summaries).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit when-to-use workflow: call memory_lookup with detail:'index' first, select ids based on title/tokens_estimate, then fetch them here in ONE batched call. It also names the alternative (memory_lookup) and the order of operations, leaving no ambiguity about usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_lookupA
Read-only

Retrieve engineering-memory entries for one or more memory types, optionally narrowed to specific anchor paths. Use before editing a security-sensitive or historically buggy file to surface prior incidents, ownership, and patterns tied to it. WORKFLOW: call with detail:"index" FIRST — each row carries id, title and tokens_estimate (the cost of fetching it) — then fetch full bodies via memory_get ONLY for the ids you will actually use, batching multiple ids into one call. Reads agents/memory/<type>/*.yml plus the agents/memory/intake/*.jsonl signal log. Read-only. Returns the v1 retrieval envelope: a status field plus per-type slices carrying the matched entries.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysNoOptional anchor paths or globs to match entries against (e.g. a file you are about to edit).
limitNoMaximum entries to return per type. Defaults to 5.
typesYesMemory types to scan, e.g. `historical-patterns`, `incident-learnings`, `ownership`. At least one required.
detailNo'index' returns compact priced rows (id, title, tokens_estimate) instead of full bodies — call this first, then memory_get the ids you need. 'full' (default) returns complete entries.full
token_budgetNoOptional token budget. When set, entries are rendered as one-line compact rows (id, type, confidence, `line`) and the row set is hard-cut at token_budget × 4 chars; omitted hits appear as a top-level `truncation` hint naming a concrete next step. Absent → the envelope is unchanged.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the tool's read-only nature, which aligns with the readOnlyHint annotation but goes beyond it by naming the exact data sources ('Reads `agents/memory/<type>/*.yml` plus the `agents/memory/intake/*.jsonl` signal log') and describing the return envelope ('a `status` field plus per-type `slices`'). This provides valuable context about side effects, data scope, and output structure not present in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the primary purpose, followed by usage context, workflow, data sources, and return format. Every sentence contributes value, and the 'WORKFLOW:' label provides clear structure. Despite its length, it is appropriately sized for the tool's complexity, with zero redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes responsibility for explaining return values, and it does so ('Returns the v1 retrieval envelope: a `status` field plus per-type `slices` carrying the matched entries'). It also covers the `detail` modes and token_budget behavior indirectly through the schema. Considering the tool's complexity and the rich schema, the description is complete enough for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining the strategic use of the `detail` parameter ('call with detail:"index" FIRST') and the purpose of `keys` ('anchor paths to match entries against'). While it doesn't add new meaning for every parameter, the workflow guidance enhances understanding of how to use the parameters effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Retrieve engineering-memory entries for one or more memory types, optionally narrowed to specific anchor paths.' It clearly distinguishes itself from sibling tools like memory_get by explaining its role as a lookup that returns indexes and summaries, with memory_get for full bodies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells when to use the tool: 'Use before editing a security-sensitive or historically buggy file...' It also provides a step-by-step workflow ('call with detail:"index" FIRST... then fetch full bodies via memory_get') and names the alternative tool (memory_get) for fetching full entries. This is explicit guidance about when and how to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_signalA

Record an engineering-memory signal — a short, anchored observation such as a recurring bug pattern or an ownership note — to the monthly intake log agents/memory/intake/signals-YYYY-MM.jsonl. Use to capture a learning tied to a specific file so future memory_lookup calls surface it. Appends to the filesystem and is rate-limited per (type, path) within a rolling window. Returns the recorded signal.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesFree-form signal body — the observation to record.
pathYesRepo-relative anchor path the signal is about.
typeYesMemory type the signal belongs to (e.g. historical-patterns, incident-learnings, ownership).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false (write operation), but the description adds that it 'Appends to the filesystem', is 'rate-limited per (type, path) within a rolling window', and 'Returns the recorded signal'. This clearly discloses side effects and constraints beyond the annotation, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three focused sentences: front-loaded with the action and destination, followed by usage context, behavioral details, and return value. Every sentence adds value, with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with 3 required params and no output schema. The description covers purpose, target file path, append behavior, rate limiting, and return value, which is complete for an agent to decide when and how to invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all three parameters (type, path, body). The description adds general context like 'short, anchored observation' and 'tied to a specific file' but does not provide additional per-parameter semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Record' with a resource ('engineering-memory signal') and a concrete destination (the monthly intake log). It distinguishes from sibling memory_lookup by stating this captures learnings so future lookups can surface them, making the tool's role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states when to use it ('Use to capture a learning tied to a specific file') and gives examples (recurring bug pattern, ownership note). It implies the read counterpart is memory_lookup but does not explicitly exclude alternatives, which prevents a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_statusA
Read-only

Report the memory backend status. Memory is entirely file-backed (agents/memory/); there is no external backend. Read-only, takes no arguments. Returns a status (file), the active backend (file), and a short reason.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds valuable context: memory is entirely file-backed at a specific path, there is no external backend, and it discloses the exact return fields (status, backend, reason). This goes beyond annotation and helps the agent understand expected behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four short sentences, each adding distinct value: purpose, backend context, read-only behavior, and return format. It is front-loaded with the purpose and contains no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only status tool, the description fully covers what it does, how it behaves, and what it returns. It also explains the file-backed architecture, making the tool's behavior predictable even without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is trivially complete. The description reinforces this by stating 'takes no arguments', which is sufficient. No additional parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with 'Report the memory backend status', a specific verb+resource pair that clearly distinguishes this from sibling tools like memory_get or memory_lookup which retrieve data. It unambiguously conveys that this is a status query, not a data access operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context is clear: it is a read-only status check with no arguments, and the file-backed nature is explained. However, it does not explicitly mention when not to use this tool or point to alternative siblings, leaving the user to infer distinctions from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_resource_bodyA
Read-only

Fetch the rendered body of a single resource URI (rule, guideline, or context document) in one call, without the two-step resources/list + resources/read handshake. Use when you already know the URI and want to inline its content into a tool-call result. Read-only. Returns the resource uri, name, description, and full text body.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriYesResource URI to fetch, e.g. `rule://commit-policy`, `guideline://php/patterns/events`, or `context://authority/scope-mechanics`.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already declares the tool is read-only, and the description reinforces this, but more importantly it adds the return payload details ('uri', 'name', 'description', and full text body') which are not in an output schema. This provides useful behavioral context beyond the annotation, though it does not cover error cases or other edge behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose in the first sentence and usage/return details in the second. No fluff or redundancy; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple single-parameter read-only tool with no output schema, and the description sufficiently covers the operation, usage condition, and return format. The sibling tools are unrelated, so no confusion. The description is complete for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes the only parameter 'uri' with examples and format. The description adds minimal additional meaning (e.g., 'single resource URI' and types) but does not go beyond what the schema already provides. Given 100% schema coverage, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Fetch the rendered body of a single resource URI'), specifies the resource types (rule, guideline, context document), and distinguishes itself from the two-step resources/list + resources/read handshake. This is a specific verb+resource definition that sets it apart from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it ('when you already know the URI and want to inline its content into a tool-call result') and references the alternative two-step approach it avoids. This gives clear context and an explicit alternative, fulfilling the dimension perfectly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

roadmap_archiveA

Archive every roadmap that has reached count_open == 0 and was touched on the current branch — git mv to agents/roadmaps/archive/, migrate inbound references, and regenerate the dashboard. Use as the PR-gate sweep before opening a pull request. Mutates the git index (moves tracked files) but never commits or pushes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the readOnlyHint: false annotation by detailing that it mutates the git index, moves tracked files, migrates references, and regenerates the dashboard, while explicitly noting it never commits or pushes. This is valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each earning its place: the first defines the action, the second gives usage context, and the third discloses side effects. It is front-loaded and has no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main actions and usage context. It does not mention return values or behavior when no roadmaps match, but given the absence of an output schema and the tool's mutation focus, the coverage is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter info. The baseline is 4 for zero-parameter tools. The description adds operational semantics but does not need to explain parameters, as none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (archive every roadmap with count_open == 0 touched on the current branch) and the resource (roadmaps). It also names the specific operations (git mv, migrate references, regenerate dashboard) and differentiates from siblings like roadmap_progress and memory_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use as the PR-gate sweep before opening a pull request,' providing a clear when-to-use context. It does not mention when not to use or explicit alternatives, but the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

roadmap_progressA

Regenerate agents/roadmaps-progress.md from the current checkbox state of every active roadmap. Use after landing roadmap work to keep the dashboard in sync without a shell round-trip. Writes the dashboard file inside the project tree. Set dry_run: true to compute counts without writing.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoWhen true, return the computed dashboard without writing the file.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the tool writes a file inside the project tree, which is useful behavioral context beyond the readOnlyHint:false annotation. It also explains the dry_run flag to avoid writing, adding transparency about how to perform a non-destructive run. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each serving a distinct purpose: main action, usage context, and dry-run behavior. It is front-loaded with the core purpose and contains no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one optional parameter, no output schema, and a clear write behavior, the description covers the essential aspects: what it does, when to use it, and how to avoid writing via dry_run. It could be slightly more explicit about return values, but overall it is complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents the `dry_run` parameter with 100% coverage. The description adds minimal extra meaning by stating it 'compute counts without writing', but this largely paraphrases the schema description. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool regenerates `agents/roadmaps-progress.md` from checkbox states of active roadmaps. The verb 'regenerate' and specific file path make the purpose unambiguous, and it differentiates from siblings like `roadmap_archive` by focusing on syncing the dashboard file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use it 'after landing roadmap work' and contrasts it with a 'shell round-trip', giving clear context on when and why to use the tool. It does not list exclusions or alternative tool names, but the guidance is sufficient for typical use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_testsA

Run the consumer project's vitest test suite under a compiled safety envelope: fixed argv (no shell interpolation), 120s timeout, 64KB output cap per stream. Shell-exec pilot per the 2026-07-07 council cut — vitest projects only; other runners (Pest / PHPUnit, pytest, Jest) return an error until a future council round approves them. Pass filter (vitest --testNamePattern) or path (in-tree file or directory) to narrow the run. Returns runner, passed, exit_code, timed_out, truncated stdout/stderr, and duration_ms.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRestrict to tests under this directory.
filterNoRestrict to tests matching this name pattern.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the sparse annotation (readOnlyHint=false), the description reveals the compiled safety envelope: fixed argv, no shell interpolation, 120s timeout, 64KB output cap per stream. It also notes the Shell-exec pilot status and details the return fields, providing substantial behavioral context that annotations alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences front-load the core action and safety constraints, then proceed to restrictions, parameter guidance, and return fields. Every sentence contributes new information with no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two optional parameters and no output schema, the description covers purpose, safety envelope, supported/unsupported runners, parameter usage, and return value shape. The only minor omission is an explicit statement that passing neither parameter runs the full suite, but this is strongly implied by 'narrow the run.'

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the input schema already describes both parameters, the description adds mapping: `filter` maps to vitest --testNamePattern and `path` is an in-tree file/directory, clarifying exact usage. This goes beyond the schema's generic descriptions and helps an agent choose parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Run the consumer project's vitest test suite,' clearly identifying the tool's action and scope. It further distinguishes itself by stating it supports only vitest projects and that other runners return an error, making it unmistakable from a generic test runner.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit exclusions ('other runners ... return an error until a future council round approves them') and explains how to narrow runs with `filter` or `path`. However, it does not name an alternative tool to use for non-vitest projects, so it stops short of full when-to-use-vs-alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_skill_for_taskA
Read-only

Match a free-form task description to the most relevant skills, ranked by a deterministic keyword scorer over SKILL.md frontmatter. Use when a skill you need is not in the catalogue the host delivered — a measured host dropped 402 entries from its model-visible list — so asking by name is impossible while asking by task is not. Read-only: no shell, no writes, and no skill bodies are returned, only names, scores and declared personas.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesFree-form description of the task to match skills against.
limitNoMaximum number of skills to return. Defaults to 5.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses behavioral traits beyond annotations: it states it is read-only, does not return skill bodies, only names, scores, and personas fraction. The annotation readOnlyHint: true is reinforced and expanded with details about no shell, no writes, and what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the main purpose, and each clause adds value. The mention of the host dropping 402 entries provides context but could be trimmed; still it's justified as it explains why this tool exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description informs the agent what is returned (names, scores, personas) and the read-only behavior. It is complete for a recommendation tool with two parameters that are fully described in schema. The additional context about the catalogue gap adds necessary usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already fully describes both parameters (task as string, limit with default and min). The description adds that it ranks by a deterministic keyword scorer, which hints at how task is used, but does not significantly enhance parameter understanding beyond the schema. Baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool matches a free-form task description to relevant skills via a deterministic keyword scorer over SKILL.md frontmatter, and it differentiates from siblings by focusing on task-based search rather than name-based listing (like list_skills). It is specific about verb, resource, and mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: when a skill is not in the catalogue the host delivered, and explains asking by name is impossible while asking by task works. It also implies when not to use (when you can name the skill) and gives a concrete scenario, distinguishing it from alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

telemetry_reportA
Read-only

Return the artefact-engagement telemetry report — essential / useful / retirement-candidate skills and rules ranked by recorded consult+apply signals over a rolling window. Use to see which artifacts are actually load-bearing. Read-only. No-op (empty report) when telemetry recording is disabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
window_daysNoReporting window in days. Defaults to 30.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states 'Read-only', which is consistent with the readOnlyHint annotation, but goes beyond it by adding a significant non-obvious behavior: 'No-op (empty report) when telemetry recording is disabled.' This is valuable context not available in structured data. It also notes the 'rolling window' aspect, adding operational detail about data recency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loaded with the main action, and packs in the purpose, usage context, and an edge-case behavior. There is no redundant phrasing or filler. Every clause contributes to understanding, making it appropriately concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with one optional parameter and no output schema, the description covers the essential aspects: what the report contains, how it is ranked, when to use it, and what happens when telemetry is disabled. It falls slightly short of a 5 because it does not describe the expected return shape (list, object) or field names, but given the simplicity, it is sufficiently complete for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema fully describes the single parameter 'window_days' with a description and default. The description mentions 'over a rolling window' but does not add new details about the parameter itself. Since schema coverage is 100%, the baseline of 3 is appropriate; the description provides minimal additional semantic value beyond what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Return' and identifies the resource as the 'arteffect-engagement telemetry report'. It details the content (skills/rules ranked by consult+apply signals) and clearly distinguishes it from sibling tools like list_skills or capabilities_index, which are general artifact listings. The inclusion of 'essential / useful / retirement-candidate' categories further clarifies the specific purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Use to see which artifacts are actually load-bearing' provides a clear use case for when to invoke this tool. While it does not explicitly name alternatives or exclusions, the focus on telemetry signals implies it is the right choice over generic listing tools. The context is sufficient for an agent to decide appropriately, but lacks an explicit 'when not to use' or alternative tool reference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.4/5.0
Disambiguation5/5

Each tool targets a distinct resource and action: memory lookup/get/signal/status, roadmap archive/progress, chat append/read, list commands/rules/skills, and separate health/check tools. Even overlapping diagnostics like doctor_report and conformance_check have clear, non-conflicting purposes.

Naming Consistency4/5

All names use lowercase snake_case and follow a mostly verb_noun pattern (run_tests, list_skills, memory_get). Some noun-phrase names like memory_status and roadmap_progress are descriptive but consistent in style, with no camelCase or chaotic mixing.

Tool Count4/5

At 19 tools, the set is somewhat large but justified by the broad scope of agent configuration (memory, roadmaps, chat history, diagnostics, listing). It sits at the upper boundary of reasonable, with each tool earning its place.

Completeness4/5

The surface covers core workflows: memory retrieval and intake, roadmap maintenance, chat history persistence, manifest listing, diagnostics, and linting. Minor gaps exist (no memory delete/update, no roadmap creation), but these are workable and likely outside the intended scope.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Cross-agent memory bridge for AI coding assistants. Persistent knowledge graph shared across 10 IDEs (Cursor, Windsurf, Claude Code, Codex, Copilot, Kiro, Antigravity, OpenCode, Trae, Gemini CLI) via MCP. 22 tools including team collaboration, auto-cleanup, mini-skills, session management, and workspace sync. 100% local, zero API keys required.
    9
    1,884
    719
    Apache 2.0
  • A
    license
    A
    quality
    C
    maintenance
    Agent Toolbelt is an MCP server exposing 11 focused API tools for LLM agents — schema generation, text extraction, token counting, CSV conversion, Markdown conversion, URL metadata, regex builder, cron expressions, address normalization, color palettes, and brand kits. Each tool is a focused microservice with structured input/output, WCAG-scored color data, USPS address parsing, and multi-model to
    25
    37
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    A remote MCP server that ships procedural knowledge skills to AI agents, enabling methodologies like brainstorm-first and plan-before-action. Hosted on Cloudflare Workers, it provides resources, prompts, and tools for skill management.

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/event4u-app/agent-config'

If you have feedback or need assistance with the MCP directory API, please join our Discord server