@event4u/agent-config
OfficialThis server provides a comprehensive set of tools for managing an AI agent's configuration, memory, chat history, skills, rules, commands, and project quality.
Chat History & Memory
Append entries to and read from chat history (
chat_history_append,chat_history_read)Mine sessions for patterns, decisions, and learnings (
mine_session)Hybrid memory retrieval over structured files (
memory_lookup)Record memory signals with rate limiting (
memory_signal)Check optional memory package status (
memory_status)
Skills, Rules & Commands Discovery
List all skills, rules, and commands (
list_skills,list_rules,list_commands)Match skills to a free-form task description (
suggest_skill_for_task)Score a user message against the compiled router (
skill_trigger_eval)Suggest relevant commands for a message/context (
suggest_command)Fetch the rendered body of any resource URI (
read_resource_body)
Configuration & Synchronization
Sync
.agent-settings.ymlagainst the current template/profile (sync_agent_settings)Keep the
.gitignoreagent-config block up to date (sync_gitignore)Regenerate the compiled skill/rule/command router from frontmatter (
compile_router)
Quality & Testing
Lint markdown frontmatter for skills, rules, commands, guidelines, and personas (
lint_skills)Run configured quality gates (PHPStan, Rector, ECS, etc.) (
run_quality_checks)Execute the project's test suite, optionally scoped (
run_tests)
Laravel-Specific
Sync the
messages()method of a Laravel FormRequest class, adding missing entries and removing stale ones (update_form_request_messages)
Integrates with GitHub to manage pull requests with structured descriptions and facilitate the implementation workflow.
Integrates with GitHub Copilot to inject project-specific skills and rules, enhancing its behavior for consistent code generation and review.
Integrates with Jira to retrieve ticket information and structure pull request descriptions based on Jira issues.
Integrates with Linear to fetch issue details and structure pull request descriptions accordingly.
Agent Config — every claim machine-checked, including the counts in these badges
How these are counted — one canonical counter, agent-config → update_counts --check, re-derives all six from the tree and fails CI on a drift of one. Two bases are not what the linked directory shows, so they are stated here rather than left to be inferred: Commands 202 counts every command file recursively (the linked directory holds 61 at its top level), and Rules 120 counts the source rules while the linked projection holds 119 — one rule is dormant and is not projected. Personas 29 excludes the directory README. Counts are of files and directories: none of them measures quality, activation, or adoption.
Try it in 30 seconds — drop one read-only subagent into any repo and watch it gate "done": @production-validator check this branch is actually done. No wizard, no lock-in, nothing else installed — the 30-second wedge ↓ is the whole first step. Start at the proof, not the catalog: event4u-app.github.io/agent-config/proof/.

Every public claim in this README is machine-checked — verify it yourself. In a market that runs on unbacked headline numbers, this one binds each claim to resolvable evidence or fails its own build.
Choose your experience — developer · founder · content · agency · finance · ops. Add packs. Get a focused command set, not a 500-artefact dump. Bring your own AI provider.
A deep library of skills, commands and governed rules — plus a capability router that loads the right skill on intent and multi-agent orchestration with consensus review. The whole layer is compiled into 20 host agents — of 23 detected, 3 being export-only (Claude Code, Cursor, Augment, Cline, Windsurf, Copilot, Gemini CLI, Codex, Continue, Zed, JetBrains, Aider and more). Resident processes are permitted only under the supervision contract ADR-249 establishes — a policy this repository adopted on 2026-08-27, not a description of anything running today. Six role-shaped entry paths sit on top, so any host becomes a reliable team member — without locking you to a single model or vendor.
What's different
It is both deep and disciplined — and honest about what it deliberately is not:
Depth that routes itself — a deep skill and command library, with a capability router that loads the right one on intent, not a 500-artefact context dump.
Governance on every host — rules compiled into each tool's native format at projection time; deterministic runtime hooks added on hook-capable hosts. This config-space, host-agnostic governance is the moat (the governance advantage · enforcement by host).
Surgical uninstall — removes only its own keys from a shared host config (matched by JSON-pointer + SHA-256), never a neighbour tool's entries.
Pack-scoped install — writes the active pack only, not a 500-artefact dump.
What it deliberately is not — the core is a governance layer with optional, individually opt-in embedded engines (code intelligence, gated reach, the setup GUI, the bench lab — ADR-124): no mandatory or always-on daemon, no separate state database, no self-rewriting memory, no auto-build pipeline. Engines are never mandatory, never default-on without measured lift, and terminate with the command that invoked them. The host agent runs the loop; every learned change is human-reviewed; the same layer stays portable across tools. Capability without a process to babysit.
Where this comes from (honest provenance). The skills, rules and personas are distilled from real production work on TypeScript and PHP codebases. The governance mechanics are stack-agnostic, but the domain heuristics are richest where they were forged — treat coverage on other stacks as promising, not proven and tell us where it falls short.
See exactly what works on which host or jump to things you can do in a minute.
Pick your profile — six entry paths
agent-config setup writes profile.id to .agent-settings.yml; each
anchor below is the first-screen the wizard sends you to. One README,
six entries, no role-detection guesswork.
Profile ( | Audience | First commands | First skills |
👩💻 | IC engineer |
|
|
Writers, ghostwriters, marketers |
|
| |
🚀 | Solo / early-stage founder |
|
|
🏛 | Multi-client delivery shop |
|
|
💼 | CFO / fractional finance / FP&A |
|
|
🛡 | RevOps, support, SRE-adjacent |
|
|
Not sure which one? Run npx @event4u/agent-config init then
agent-config setup — the browser wizard asks a single 8-option role
question and maps to the closest profile. Source-of-truth:
src/agent-src/profiles/ ·
schema: docs/contracts/profile-system.md.
Beyond software: user-types/
(galabau · metalworking · truck — see Beyond software.
Per-profile experience pages (who it's for · first tasks · packs + flows · what is not loaded · examples): developer · content_creator · founder · agency · finance · ops.
Workflows, not raw commands
You don't memorize every command — you run a work journey. Four flows span the developer story end-to-end; each names the command you TYPE to start and the skills it composes:
Flow | Start with | The journey |
🔍 Discovery |
| explore → plan → estimate → refine, before building |
🔨 Implementation |
| plan → implement → verify → commit |
🔎 Review |
| self-review → judge → quality-fix → threat-model |
🚢 Delivery |
| commit in chunks → open PR → answer review |
Full detail — entry commands, canonical path, composed skills per flow:
docs/flows.md. (agent-admin — memory / analytics / config —
is platform operation, not a user-work flow.)
Creative Pack — cinematic AI video. script → character-locked image → motion+audio prompt → provider render → stitched clip, with
AIV_DRYRUN=trueas the cost-safety default. A first-class capability inside the content / creator experience — no longer the package's headline. See/video:from-script.
Legal Pack — not legal advice. The EU/DE legal pack (contract/NDA/DPA review, triage) is a research-and-drafting aid only — it does not provide legal advice, does not replace a qualified lawyer and must not be relied on for any concrete matter. It produces general information and general templates, never individual-case examination. Read
LEGAL_NOTICE.mdbefore use.
Full catalog — every skill, rule, command, guideline:
docs/catalog.md. The headline is the experience (profile + packs) and the depth behind it.
Use it in your project
Run from a consumer repo — bootstrap via npx, the agent picks up
your stack and you ship work end-to-end. New install? Start with the
Quickstart. Already installed? Supported tools
shows the wired AIs; docs/featured-commands.md
lists the end-to-end workflows (/implement-ticket, /work,
/commit, /pr:create). Deeper tour: 2-minute demo.
Install scope. Pick one scope per machine — project-local (default, recommended for application repos) or user-global (recommended for tooling repos / dotfiles). The installer refuses a second, conflicting scope via the scope_guard pre-flight. Details: docs/contracts/install-scopes.md. Cleanup when needed: bash src/scripts/cleanup_other_scope.sh --confirm.
Related MCP server: Agent Module
Prove it
Don't take the claims on trust — verify them. docs/proof.md
is generated from source: a claim→evidence table (every public claim binds to a
resolvable pointer or CI fails), honest-null benchmarks including the runs where
the package changed nothing and a "verify it yourself" block you run on a fresh
checkout. The proof page fails CI if it drifts from its sources — reproducibility
is the proof. Browse it on the deployed docs site — the proof page is the primary entry:
event4u-app.github.io/agent-config/proof/.
The honest comparison frame lives at
docs/us-vs-the-category.md.
Freshest measured row: in one post-fix session, advisory context injection cut language-mirror violations 555 → 19 while the two blocking guards went 8 → 0 and 1 → 0 — advisory reduced massively, only blocking eliminated. One session and a post-hoc reading, so a recorded prior rather than a law.
Maintaining a skills catalog yourself? The anti-reskin gate that blocks
find-replace re-skin PRs here runs on your catalog too —
docs/anti-reskin-gate.md.
Audit-disciplined by construction — every memory consult, decision
key and hook concern lands in agents/runtime/state/ so you can replay it.
Core principles names the four invariants;
What agent-config is — and what it isn't
draws the scope boundary.
Contribute
Working on the package itself? Development covers the
task ci pipeline, Requirements the toolchain,
Maintainer telemetry the
opt-in measurement loop. Source-of-truth tree is
src/ (src/skills, src/rules, src/agent-src/); never hand-edit .augment/ or dist/agent-src/.
Security. Disclosure policy: SECURITY.md. Threat model: docs/threat-model.md.
Quickstart
Try one thing in 30 seconds — before the full suite, drop in a single self-contained subagent and see the discipline on your own repo:
mkdir -p .claude/agents
curl -fsSL https://raw.githubusercontent.com/event4u-app/agent-config/main/docs/wedge/production-validator/production-validator.md \
-o .claude/agents/production-validator.md
# then in Claude Code: @production-validator check this branch is actually doneproduction-validator is read-only and installs nothing else — it gates "done"
by hunting mocks/stubs on the shipped path and demanding real-system evidence
(what it does). Like it? Install the
full suite:
One command. Detection-driven — your installed AI tools are found and pre-selected. Nothing is written until you click Finish. No YAML by hand.
Those four are structural: they hold on every run, because they are properties
of the code path rather than of your machine. How long it takes is not one of
them — that is dominated by network and registry latency, which we do not
control. The install → doctor wall-clock is measured by CI on every umbrella
run and published with its conditions, as evidence; it is never a promised
number.
# 1. Install — on a terminal with a display, the browser wizard launches
# automatically; the same TypeScript installer runs the real install behind it.
npx -y @event4u/agent-config init
# 2. Pick your profile + tools in the wizard, click Finish.
# (Writes ~/.event4u/agent-config/, ~/.claude/, ~/.cursor/, …)
# 3. First real task — agent refines, plans, verifies.
/work "your first real task"Headless / CI: init skips the GUI automatically on CI, on a non-TTY, on a headless host, and whenever any CLI-mode flag is present — it then runs the non-interactive installer directly. The full opt-out set is listed once, against the code, in gui-wizard § When the GUI is skipped. Pass flags (--profile=balanced --tools=claude-code,cursor); add --dry-run to preview writes. The GUI and the CLI share one installer (src/scripts/install.ts), so both produce identical results. Reference: docs/wizard.md.
Pick specific AIs: --tools=claude-code,cursor,augment,windsurf,cline,gemini-cli,copilot,roocode,aider,codex,claude-desktop,continue (any subset). Visual picker: add --gui (loopback-bound, CSRF-gated; contract gui-wizard). --gui is an opt-in that forces the wizard past the TTY and headless checks — it does not override CI, AGENT_CONFIG_NO_UI, or a CLI-mode flag; combining it with one of those exits non-zero rather than quietly running the CLI install. On a headless host add --allow-headless and connect a browser to the printed URL.
Verify hook coverage: npx @event4u/agent-config hooks:status prints the per-platform matrix (--strict for CI, --format json for tooling).
Scope (v2.5+):
initwrites global only —~/.event4u/agent-config/,~/.claude/,~/.cursor/, …. The project tree getsagents/overrides/only (the bridge marker was retired — ADR-020 amendment 2026-07-13; the global root resolves from~/.event4u/agent-config).--projectis maintainer-only behindAGENT_CONFIG_DEV_MODE=1(ADR-020, dev-mode).
Migrating from a v1.x install? npx @event4u/agent-config migrate — full notes in docs/migration/v1-to-v2.md.
What agent-config is — and what it isn't
A content layer — skills, rules, commands, guidelines, personas — distributed via npm and projected into every supported AI tool's native config format. It follows the Agent Skills open standard.
It is not an agent runtime. The agent loop, the LLM dispatcher and tool orchestration stay with the host tool (Claude Code, Augment, Cursor, Cline, Windsurf, Gemini CLI, Copilot). Think of it as a playbook and style guide for those tools — not a replacement.
In scope | Out of scope |
Skills, rules, commands, guidelines, personas | Agent loop / LLM dispatcher |
Multi-tool projection + condensation pipeline | Execution engine inside the package |
Memory helpers ( | Cross-tool observability dashboard |
Linters, CI, frontmatter validation against JSON-Schema (contract) | Runtime GUI / web dashboard |
Skill orchestration via citations + deterministic helpers | Opinionated automatic skill-resolver (ML / relevance ranking that decides for you) |
User-driven projection-time filtering by profile + packs (ADR-040) | A runtime resolver / daemon (mid-session switching — conditional, post-6.0.0) |
What your agent is asked to do
Default behavior | With agent-config |
Guess and edit blindly | Analyze code before changing it |
Drift from project conventions | Follow detected stack conventions |
Skip or invent tests | Write tests in the project's framework |
Generic commit messages | Conventional Commits with scope + ticket links |
Skip quality checks | Run the project's quality pipeline and fix reported errors |
Open PRs without context | Structured PR descriptions from Jira / Linear / GitHub |
Claim "done" without proof | Verify with real execution before claiming done |
2-minute demo — /implement-ticket
The flagship command. Drives a ticket end-to-end through a fixed linear flow — and blocks on ambiguity instead of guessing.
/implement-ticket PROJ-123The agent runs this sequence:
refine → memory → analyze → plan → implement → test → verify → reportRefines the ticket if acceptance criteria are vague.
Queries memory for past decisions, invariants, incidents.
Plans the change; you confirm before any file is touched.
Implements under
minimal-safe-diff+scope-control— no drive-by edits.Tests (targeted first, full suite on success).
Reviews the diff through four judges (bugs, security, tests, code quality).
Reports changes, verdicts, follow-ups — then stops.
/commitand/pr:createare suggestions, never auto-run.
Any ambiguity halts the flow with numbered options — never a silent guess. Persona comes from .agent-settings.yml (roles.active_role): senior-engineer (default), qa or advisory (plan-only).
→ Command reference · Flow contract
Sibling — /work (free-form prompt)
Same engine, no ticket required:
/work add a CSV export endpoint to the audit-log controllerThe first pass scores the prompt on five dimensions and routes on the band:
Band | Score | Action |
high |
| Silent proceed — AC + assumptions in the report |
medium |
| Halts with assumptions report; confirm or edit |
low |
| Halts with one clarifying question on the weakest dimension |
After the band gate, the flow is identical to /implement-ticket. Free-form goal → /work; ticket payload → /implement-ticket.
→ Command reference · refine-prompt skill
After the run: agent-config explain last reconstructs the trace (route · memory · council · halts · provider) — read-only, PII-scrubbed, offline. Docs
Product UI track
UI-shaped work routes to one of three directive sets — ui (full audit→design→apply→review→polish→report), ui-trivial (≤ 1 file, ≤ 5 lines: apply→test→report), mixed (backend + UI: contract→ui→stitch). Existing-UI audit is a hard gate (ui-audit-gate); polish has a 2-round ceiling with a11y precedence. Stack detection → blade-livewire-flux / react-shadcn / vue / plain.
→ Mental model (1 page) · Flow contract
Customize
Profiles — how much governance gets loaded
Safety floor (non-destructive defaults · ask-before-guessing · mirror-the-user's-language) ships in every profile. What changes is how much extra coaching gets pulled in.
Profile | What you get | When to pick it |
| Non-negotiable safety floor only. Cheapest, fastest. | Quick questions · throw-away scripts · CI · tight token budgets |
| Safety floor + everyday coaching (sensible defaults, review nudges, common pitfalls). | Day-to-day work |
| Everything, including long-tail rules normally only maintainers need. | Working on |
Under the hood: kernel-only · kernel + tier-1 · kernel + tier-1 + tier-2. Details: rule-router · kernel-membership · Configure →.
Stability:
STABILITY.mdfor the full matrix. Work Engine (/work+/implement-ticket): beta. Runtime Dispatcher: stable. Tool Adapters: experimental (fullprofile only).
.agent-user.md and Ghostwriter — voice primitives
Primitive | Voice | Disclosure |
Review-lens (internal critique) | n/a | |
| The maintainer's own voice — | None (you are the author) |
| Documented public figure — | Mandatory, non-removable footer |
Create the user file interactively: /agents user init (schema). Ghostwriter cluster: /ghostwriter:fetch <url-or-name> runs an attestation gate; private individuals rejected; paywalled / leaked / DM content banned at the schema level.
Self-hosted MCP on Cloudflare — zero local install
Skills, commands, rules and guidelines can be served as an MCP endpoint from your own Cloudflare Worker — any MCP client (Claude Desktop, Claude Code, Cursor, Zed, Continue, hosted agents) talks to it over HTTP. Two auth modes: public (default, OSS read-only deploys) and bearer-auth (operator opt-in, MCP-Token Wrangler secret).
task mcp:cloud:login # one-time, opens browser
task mcp:cloud:setup # check → r2-create → r2-verify → whoami
task mcp:cloud:secret-put # opt in to bearer-auth (recommended for private deploys)→ Operator walkthrough: mcp-cloud-setup · Per-client config: mcp-client-config · Endpoints: mcp-cloud-endpoints.
Scope — Lite, not Full. The Worker serves read-only governance (skills · commands · rules · guidelines · contexts) as MCP prompts and resources, plus small read-only tools (
memory_lookup,chat_history_read,list_*). It does not execute the repository's local scripts (linters, audits,task ci, work-engine hooks) — those require local install per Quickstart.
The built-in local stdio server is listed for discovery in the Glama MCP Registry (agent developers / contributors; requires a local checkout, not a turnkey install — see ADR-067).
Deployment posture
Shape | Status | Path |
Single-user workspace | ✅ today |
|
Small team (3–10 people) | ✅ today | Shared |
Organization mode (SSO · central policy · team context · internal connectors) | ⏸ not started | Each shape gated on a recruited customer + funded audit + maintainer ADR. Posture rationale: |
The Hard Floor on organization-mode features (SSO, central policy, OAuth connectors, team-context) is preserved by design — they stay cancelled until a real first customer + funded security audit lifts them. The small-team recipe is the supported path in the meantime.
The 9.3/10 feedback round (2026-05-25) re-asked for OAuth knowledge connectors, IAM / org governance and organization-shared memory. Each is a stable cancellation row in
team-deployment-postureunder the same three release gates — recruited team customer · funded audit · maintainer ADR.
Harness expectations
Three classes of install/runtime behaviour look like package bugs but are host-harness behaviour the package cannot control — sibling-plugin namespaces (codex:*, cc-gemini-plugin:*), deferred tools surfaced via ToolSearch and cross-scope skill drift (real bug, fixed in the distribution-channels track). Diagnostics + the package's response: docs/contracts/harness-expectations.md. First step when a skill appears twice: task probe:skills.
Supported tools
Project-installed (npx)
Tool | Rules | Skills | Commands | How it works |
Claude Code | ✅ | ✅ | ✅ | Reads |
Cursor | ✅ | — | ☑️ | Reads |
Cline | ✅ | — | ☑️ | Reads |
Windsurf | ✅ | — | ☑️ | Reads |
Gemini CLI | ✅ | — | ☑️ | Reads |
GitHub Copilot | ✅ | — | ☑️ | Reads |
Roo Code | ✅ | — | ☑️ | Auto-discovers |
Codex CLI | ✅ | — | ☑️ | Auto-discovers |
Continue.dev | ✅ | — | ☑️ | Auto-discovers |
Aider | 📌 | — | — | Manual |
Augment (VSCode/IntelliJ) | 📌 | — | — | Global-only; project writes marker |
Claude Desktop | 📌 | — | — | Global-only |
✅ native ☑️ text reference (in AGENTS.md, not invokable as native slash-command) 📌 marker only — not available
Team reproducibility: every tool you
initis recorded inagents/installed-tools.lock(committed, machine-managed). New team members runnpx @event4u/agent-config syncafter cloning; CI gates drift withagent-config validate. Schema:installed-tools-manifest.
Plugin-installed (optional, global)
Tool | Install |
Augment CLI · Copilot CLI | Install → — rules + skills + commands, marketplace-updated |
Claude Code: the marketplace plugin is deprecated (single-surface model). The npx/npm file projection now carries content and the deterministic hooks (registered in a managed
~/.claude/settings.jsonblock byagent-config global/upgrade), so the plugin only duplicates skill/command listings while its git-SHA snapshot rots silently. Existing installs:claude plugin uninstall agent-config@event4u-agent-config—agent-config doctorflags the duplicate surface.
Keep the global install current with agent-config upgrade (latest) or
agent-config refresh --global (same-version re-install); agent-config doctor
flags a missing-from-PATH binary or broken hook wiring. See
getting-started § Keeping current ·
Troubleshooting.
The command surface at a glance
Command | What it does |
| One-shot install — opens the browser wizard (recommended path or step-by-step) |
| Initialize a project: minimal |
| Open the configuration GUI — global settings hub (simple + advanced tiers, search, reset-to-default) |
| Open the project configuration surface |
| Re-run the guided onboarding wizard (prefilled from your current state) |
| Update the global install to the latest release + additively sync settings |
| Read-only health/drift report |
Cloud / Hosted-agent surfaces
For platforms where the package's scripts cannot run, artefacts are built for paste-in or upload:
Linear AI (Codegen, Charlie, …) —
dist/linear/{workspace,team,personal}.mdClaude.ai Web Skills —
dist/cloud/<skill>.zip
Works with agent-switch
agent-switch is the companion
CLI for running several agent accounts on one machine: it isolates each account
in its own profile (CLAUDE_CONFIG_DIR per profile), so switching accounts
never means logging out and back in. The two compose — agent-switch isolates
the accounts, agent-config governs what the agents do inside them. When
agent-config runs under an agent-switch profile it says so in the settings hub,
warns before writes that would land in a shared (cross-profile) tree, and
accepts a host-supplied config root so its own settings stay profile-scoped.
Who this is for
Stack-agnostic governance core (orchestration · role modes · command clusters · quality gates · audit-discipline) plus parallel stack-specific skill sets:
Stack | Coverage |
Laravel · modern PHP (deepest) | Pest · PHPStan · Rector · ECS · Eloquent · Livewire/Flux · Horizon · Pulse · Reverb · Pennant |
Symfony |
|
Next.js App Router |
|
Zend / Laminas | project-analysis + shared PHP coder/quality skills |
React · Node / Express | project-analysis + UI |
Vue · plain HTML | UI directive set ( |
Cross-stack | API design · testing · security · database · Docker · Git · CI · review · threat modeling · observability |
Beyond software
The same orchestration core drives non-software trades via user-types/: galabau-field-crew · metalworking-shop · truck-driver. Contribute your own — 5-minute scaffold.
Data governance & domain safety
Three domain-safety rules (domain-safety-pii, domain-safety-disclaimer, domain-safety-retention) act as per-domain output floors across ~12 areas — PII redaction (support / finance / recruiting / marketing), advice disclaimers (legal / financial / medical / consulting), retention guidance (finance / support), ops floors (logging / export). Full surface → rule → floor matrix: docs/safety.md. Beta contracts: memory-visibility-v1 · decision-trace-v1.
Code provenance & license governance
Every diff is checked against a license policy derived from the target repo's own detected license (LICENSE/package.json/composer.json, precedence-ordered; sources disagree → escalate, never guess) and a strict linter over our own borrow ledger (provenance/borrows.jsonl → docs/THIRD-PARTY-NOTICES.md) that fails a deny-class license, an unknown license, a missing transformation note or a rename-only-phrased one — wired into ci/ci-strict from day one. A third piece, license-compliance-audit, runs an offline/online similarity scan on demand — a human invokes it deliberately, never a pipeline. This is provenance-governed, license-policy-enforced borrow discipline backed by an audited borrow trail — not a copy detector.
Scope & limits
Unconscious training-data reproduction is not detectable at this layer. No tool here — or anywhere — can see what a model's training data contained; this system governs what gets consciously borrowed and recorded, never what a model silently recalls.
Detection, where it exists, covers a knowledge base of known OSS only — a subset of all code that has ever existed, never a model's training corpus.
No CI-facing detection gate exists. A deterministic scanner (jscpd offline + SCANOSS online) was built and measured against a frozen synthetic corpus, but missed its own pre-registered thresholds (measured: recall 12/16, false positives 2/12, SCANOSS rename-only recall 0/8) — see
docs/CLAIMS.md. It ships in no form in CI, not even advisory — only as the on-demand skill above.Rename-only laundering is not detected by anything we ship or evaluated. The ledger's transformation-note check rejects a rename-only-phrased note, but it cannot catch an undisclosed rename-only copy that was never logged.
Reduces and documents risk — never eliminates it.
Maintainer telemetry (opt-in, default-off)
Local-only artefact-engagement log. Set telemetry.artifact_engagement.enabled: true in .agent-settings.yml. Records which skills / rules / commands / guidelines the agent consults during /implement-ticket / /work. JSONL under the project root, nothing uploaded. Reports: npx @event4u/agent-config telemetry:report.
Context-aware command suggestion
When a prompt matches a command's purpose ("setze ticket ABC-123 um" → /implement-ticket), the agent surfaces matches as numbered options — nothing auto-executes. Per-conversation off: /command-suggestion-off. Settings: commands.suggestion.{enabled,blocklist,confidence_floor} in .agent-settings.yml.
Core principles
Analyze before implementing — no guessing, no blind edits
Verify with real execution — no "should work"
Challenge to improve — agents are thought partners, not yes-machines
Strict by design — quality over flexibility
Zero overhead by default — nothing runs until you ask for it
Documentation
Document | Content |
First run, 3-test experience, profiles, next steps | |
All install paths, Composer/npm, orchestrator details | |
System layers, content pipeline, tool support matrix | |
Overrides, AGENTS.md, agent settings, cost profiles | |
Linting, CI pipeline, condensation system | |
Per-version upgrade steps | |
More examples & expected behavior |
Browse content: all commands · skills catalog · full catalog · llms.txt.
Troubleshooting
First stop for any install problem: agent-config doctor — it flags a
missing-from-PATH binary, binary↔plugin version drift, stale orphans and
manifest issues, each with a one-line fix hint.
For "why didn't rule/hook X fire?" questions: agent-config routing:doctor
— a read-only, live diagnosis that reports every session-start gate as
ACTIVE/INACTIVE with the concern's own reason (e.g.
session-canary: ACTIVE for "Alex" vs INACTIVE — no name on any settings layer), the platform's concern chain, host hook registration, and router +
projection freshness. Deeper hook internals (fail-open/closed posture, last
dispatcher feedback per concern): agent-config hooks:doctor.
A new command / skill is missing in Claude Code after an upgrade
Under the single-surface model, agent-config upgrade refreshes the
~/.claude/ file projection — that IS the content surface, so a fresh
session picks the new commands up directly. If commands are still missing,
the usual cause is a leftover marketplace plugin: it is a git-SHA
snapshot that never moves with the npm upgrade and it shadows nothing —
it just lists everything twice while lagging behind. Remove it:
claude plugin uninstall agent-config@event4u-agent-configThen start a new Claude Code session. agent-config doctor reports a
leftover plugin as claude-plugin: duplicate surface; hooks are unaffected
(they live in a managed ~/.claude/settings.json block — verify with the
hook-wiring check).
Skills / commands appear twice in Claude Code
Same cause as above: the deprecated marketplace plugin is installed next to
the ~/.claude/ file projection, so every skill lists plain and
agent-config:-prefixed. Uninstall the plugin (command above) and start a
new session.
agent-config upgrade fails with Unknown argument: --no-ui
Known bug in 8.2.0: upgrade passed a --no-ui flag that the install
orchestrator did not accept yet, so the run aborted early. Fixed on main;
until the next release, work around it with:
AGENT_CONFIG_NO_UI=1 agent-config global # refresh the global install, no wizardUpgrade was interrupted (Ctrl-C, wizard closed, step failed)
Only the initial npm install -g hard-aborts an upgrade. Every later step
(global re-deploy with hook registration, settings sync, wrapper + git-hook
refresh) runs independently — a single failed step is reported in the
end-of-run summary instead of silently skipping the rest. Re-run
agent-config upgrade to converge and use agent-config doctor to name
anything left in a mixed state.
agent-config: command not found / hooks stopped firing
Runtime hooks resolve the global binary on PATH — a project-local
install alone is not enough for them. Reinstall the binary:
npm install -g @event4u/agent-config
agent-config doctor # verifies PATH + plugin wiringProject files look stale after a package update
Project-local projections are only rewritten on an explicit refresh:
agent-config refresh # re-apply the installed version to this project
agent-config refresh --global # same-version re-install of the global rootMore per-version steps: Migration · getting-started § Keeping current.
Development
Working on the package itself? Edit src/ (the source of truth — src/skills, src/rules, src/agent-src/), regenerate trees:
task sync # regenerate dist/agent-src/ and .augment/
task generate-tools # regenerate .claude/, .cursor/, .clinerules/, .windsurfrules
task ci # full pipeline — green before PR
task test # unit + integration tests
task dev:setup # boot the onboarding wizard against the working treeInvoking the CLI from a source checkout: ./agent-config <command> (the maintainer shim at the repo root → scripts/agent-config → dist/cli/agent-config.js). npx @event4u/agent-config doesn't resolve in the source repo without a prior npm link, since there's no node_modules/.bin/agent-config symlink — use ./agent-config instead. Build the TS binary with npm run build:cli if dist/cli/agent-config.js is missing.
→ Full project structure and commands: docs/development.md · CONTRIBUTING.md. Stack: TypeScript throughout — CLI, UI, and the build / lint scripts. MCP registry payloads render under dist/mcp/ (submission checklist).
Requirements
Node ≥ 20.11 —
npx @event4u/agent-config initis the canonical install path. No Python anywhere on the install path (the Python installer retired with the TypeScript migration).Platform: macOS 12.3+, Linux, WSL2. Git Bash needs Developer Mode for symlinks. Contributors rebuilding
.augment/also need Task.
Windows
Native PowerShell / cmd is not supported for the file install — use WSL2 for the full installed tree. The supported native-Windows surface is the MCP stdio server: point any MCP client at
npx -y @event4u/agent-config mcp-serverand the governance content (prompts, resources, tools) is available without
the file install. Porting the bash dispatcher to native Windows is
demand-gated: a named Windows adopter who cannot use WSL2 or the MCP
path reopens it (see agents/roadmaps/ — road-to-credible-install Phase 3).
Funding
The package is free, MIT, and stays that way — no paid tier, no dual licensing. If it saves you time and you want to chip in, the GitHub Sponsor button at the top of the repo is the whole mechanism. If you would rather not, use it anyway; nothing here is gated on it.
License
MIT.
mcp-name: io.github.event4u-app/agent-config
Available Tools
20 toolscapabilities_indexA
Regenerate CAPABILITIES.yaml, the package's coverage index of skills, rules, commands, and guidelines. Use after adding or removing an artifact to keep the index current. Pass check: true to run in read-only CI mode (fails instead of writing on drift).
| Name | Required | Description | Default |
|---|---|---|---|
| check | No | When true, run in read-only drift-check mode instead of writing the index. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that the tool writes/regenerates the index, and importantly explains the read-only CI behavior ('fails instead of writing on drift'). This complements the readOnlyHint: false annotation by giving concrete failure semantics, though it does not detail side effects or required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and resource, then usage timing, then the parameter mode. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter write tool with no output schema, the description fully covers what the tool does, when to use it, and how the parameter changes behavior. The inclusion of failure mode in CI makes it operationally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully describes the 'check' parameter with high coverage, and the description adds practical context by explaining the CI-mode behavior. This reinforces the schema without redundancy, raising it above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Regenerate') and a clear resource ('CAPABILITIES.yaml'), identifying it as the package's coverage index. This distinguishes it from sibling tools like list_skills or list_rules, which query rather than update the index.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after adding or removing an artifact to keep the index current,' which is clear guidance for when to invoke it. It also explains the CI-mode use case with 'check: true', but does not name alternative tools to use instead of this one, such as list_* commands.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chat_history_appendA
Append one structured entry to the consumer project's chat-history log (a JSONL file). Use to record a decision, note, or phase marker that should persist into a later session or be distilled by mine_session. Writes to the filesystem (agents/runtime/.agent-chat-history by default; agents/.agent-chat-history and .agent-chat-history accepted for back-compat) and returns the written entry plus its resolved target path. Path-scoped: a path outside the allowlist, or any traversal escaping the project root, raises an error before writing. Set dry_run: true to preview the entry and target path without touching disk.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Optional path override. Must resolve to `agents/runtime/.agent-chat-history` (current default), `agents/.agent-chat-history`, or `.agent-chat-history` under consumer_root. | |
| text | Yes | The entry body to record. | |
| dry_run | No | When true, return the entry and resolved target path without writing to disk. | |
| session | No | Optional 16-char session id to group the entry under. Defaults to the current session. | |
| entry_type | No | Short ``t`` tag categorising the entry (e.g. note, decision, phase). Defaults to ``note``. | |
| min_schema_version | No | Refuse to write if the on-disk history schema is older than this version. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint: false, annotations are minimal. The description fully discloses the write behavior, filesystem target, return value ('returns the written entry plus its resolved target path'), safety validation ('path outside the allowlist... raises an error before writing'), and the dry_run side-effect-free preview. This goes well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, front-loaded with the primary action, then usage context, then safety/behavior specifics. Every sentence adds critical information: what it does, when to use it, the write behavior and safety, and the dry_run escape hatch. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter write tool with no output schema and minimal annotations, this description is remarkably complete. It covers the target file locations, the return payload, path restrictions, dry_run behavior, and downstream consumers. An agent can confidently select and invoke this tool without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds meaningful behavioral semantics for the 'path' parameter (allowlist and traversal checks) and clarifies dry_run's purpose. It also explains the entry body ('text') in the context of a JSONL log, enhancing understanding beyond the schema's bare field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Append one structured entry to the consumer project's chat-history log (a JSONL file).' This is a specific verb+resource that distinguishes it from siblings like chat_history_read, which retrieves the log, and other memory tools that operate on different stores.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use to record a decision, note, or phase marker that should persist into a later session or be distilled by `mine_session`.' This gives clear context for when to use the tool. It also notes back-compat paths and the dry_run option, but it does not explicitly state when not to use it (e.g., for ephemeral or non-history data), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chat_history_readARead-only
Read recent entries back from the consumer project's chat-history JSONL (agents/runtime/.agent-chat-history; agents/.agent-chat-history accepted for back-compat). Use to recover context from an earlier session — decisions, notes, phase markers — at the start of a new task. Read-only. Returns the resolved file path plus a list of matching entries (newest last). Combine session, last, and entry_type to narrow the result.
| Name | Required | Description | Default |
|---|---|---|---|
| last | No | Return only the most recent N entries, after other filters apply. | |
| path | No | Optional history-file path override; defaults to the standard chat-history location under the project root. | |
| around | No | Timeline anchor: return the entries around this ref (from a detail:"index" row) instead of the filtered list — depth_before/depth_after neighbours plus the anchor. Refs are within-file ordinals; rotation invalidates them. | |
| detail | No | 'index' returns compact rows (ref, t tag, ~100-char preview, tokens_estimate) instead of full entries — scan first, then re-read the refs you need via `around` or a narrowed filter. History entries are large, so index mode pays off here. 'full' (default) returns complete entries. | full |
| session | No | Filter to a single 16-char session id. | |
| entry_type | No | Filter by the `t` tag (e.g. note, decision, phase). | |
| depth_after | No | Neighbours after the anchor (with `around`). Defaults to 3. | |
| depth_before | No | Neighbours before the anchor (with `around`). Defaults to 3. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it adds valuable context: returns the resolved file path plus matching entries (newest last), describes path resolution with back-compat, warns that rotation invalidates refs, and notes that history entries are large so 'index' mode is helpful. These details give the agent a clear behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at four sentences, each serving a distinct purpose: what it reads, when to use it, return shape, and filter guidance. No redundant text or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 optional parameters and no output schema, the description covers the core aspects: purpose, file location, filtering hints, and return format. It lacks explicit mention of error conditions or edge cases, but the schema and rich parameter descriptions fill the gap reasonably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds a useful tip about combining 'session', 'last', and 'entry_type' filters, but it doesn't further explain each parameter or add semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Read' with a concrete resource ('chat-history JSONL') and gives exact paths, making the tool's function unambiguous. It differentiates from sibling 'chat_history_append' by emphasizing read-only access and mentions return contents (entries, resolved path).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the intended use case: recovering context from earlier sessions at the start of a new task. It also advises combining filters to narrow results, but it doesn't explicitly mention when not to use it or name alternative tools (e.g., memory_get) for different context sources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
conformance_checkARead-only
Run the consumer conformance contract (doctor --ci plus installed-and-firing checks) and return pass/fail per check. Use to verify a consumer project's install is fully wired before relying on it. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, and the description reiterates 'Read-only,' with no contradiction. It adds behavioral detail by naming the commands executed and stating that results are returned per check, going beyond the annotation's simple safety flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, command details, and output type, then a clear usage recommendation. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter tool with no output schema, the description adequately covers the operation, return semantics (pass/fail per check), and when to use it. There are no essential gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so schema coverage is trivially 100%. No parameter information is needed, and the baseline for zero-parameter tools is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource combination ('Run the consumer conformance contract') and clearly identifies the underlying command (`doctor --ci` plus installed-and-firing checks) and output (pass/fail per check). It distinguishes this from sibling tools like `doctor_report` or `run_tests` by emphasizing verification of a consumer project's install wiring.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit usage scenario: 'Use to verify a consumer project's install is fully wired before relying on it.' It does not mention when not to use the tool or alternatives, but the context is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
council_estimateARead-only
Estimate the token cost of an AI-council debate over a given input (roadmap, diff, prompt, or file set) without spending — no network call, no billing. Use before deciding whether to authorize a real council run. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | Council depth tier to estimate for. Defaults to the configured default depth. | |
| input_path | Yes | Path to the roadmap, diff, or file to estimate council cost for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation already marks it as read-only, and the description adds meaningful behavioral context beyond that: 'no network call, no billing' and 'without spending'. This clarifies the side-effect-free nature in a way the annotation alone does not convey. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences and packs in purpose, safety, and timing without any filler. It is front-loaded with the action verb and resource, making it immediately understandable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simple read-only nature and the rich annotations and schema, the description adequately covers purpose, usage timing, and safety. It does not describe the return format or precision of the estimate, and since there is no output schema, that is a minor gap, but not enough to lower the score further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes both parameters, so schema coverage is 100%, giving a baseline of 3. The description adds value by specifying that the input can be a 'roadmap, diff, prompt, or file set', which expands on the schema's narrower 'roadmap, diff, or file' phrasing. Depth is not mentioned, but the schema's enum covers it sufficiently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Estimate the token cost') and the resource ('AI-council debate over a given input'). It also distinguishes itself from a real council run by emphasizing 'without spending — no network call, no billing', leaving no ambiguity about its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Use before deciding whether to authorize a real council run' explicitly defines when to use this tool. While it provides a clear context, it does not explicitly mention when not to use it or name an alternative tool, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
doctor_reportARead-only
Run the consumer-project doctor diagnostic and return a structured health report (install drift, hook wiring, settings schema, discovery manifest freshness). Use to triage a misbehaving install. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares the tool is read-only, and the description reinforces this with 'Read-only.' Beyond that, it adds useful behavioral context about the nature of the diagnostic (checks for drift, hook wiring, settings schema, freshness) and the structured form of the output, which is not fully derivable from the annotation alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the action and resource, and contains zero filler or redundant content aside from the harmless repetition of 'Read-only.' Every clause adds information about the tool's purpose, output, or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only diagnostic tool with no output schema, the description provides all necessary context: what it runs, what the report covers, why you would use it, and its safety profile. The annotations and schema further support completeness, leaving no significant gaps for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline of 4 applies per the rubric. The schema is fully covered (100%) and the description does not need to clarify any parameter details because none exist. The description appropriately focuses on behavior and use case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Run') and names a precise resource ('consumer-project doctor diagnostic'), while detailing the report's contents (install drift, hook wiring, settings schema, discovery manifest freshness). This clearly distinguishes it from siblings like run_tests or telemetry_report by identifying a unique diagnostic target and purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: 'Use to triage a misbehaving install.' This provides clear context for the appropriate use case. However, it does not mention when not to use it or suggest alternatives among the sibling tools, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lint_skillsARead-only
Lint skill, rule, command, guideline, and persona markdown files for frontmatter and structural errors. Use before committing or opening a PR that adds or edits any of those artifacts, to catch schema violations early. Read-only — never writes files or spawns git. Returns the scripts/skill_linter.py --format json payload: a summary object (pass / pass_with_warnings / fail / total counts) and a per-file results array with severity-tagged findings. Pass paths to lint a subset; omit for a full tree scan.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | No | Repo-relative paths to lint (files or directories). Empty or missing → full tree scan via gather_all_candidate_files. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description explicitly states the tool is read-only, never writes files, never spawns git, and details the exact return payload structure (summary and per-file results). This adds meaningful behavioral context well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, front-loading the primary purpose, then usage guidance, then safety/behavior, then return format, and finally parameter usage. Every sentence contributes valuable information with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, when to use, safety/read-only behavior, return payload, and parameter semantics. With a simple one-parameter input schema and no output schema, the description provides all necessary context for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides a thorough description for the 'paths' parameter, including the meaning of empty/missing values (full tree scan). The description merely restates this ('Pass 'paths' to lint a subset; omit for a full tree scan') without adding new information. Schema coverage is 100%, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (lint) and a clear resource (skill, rule, command, guideline, and persona markdown files) to state what the tool does. It distinguishes itself from siblings by being the only linting tool among the listed tools, and it mentions 'frontmatter and structural errors' as the focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit context for when to use the tool ('Use before committing or opening a PR that adds or edits any of those artifacts'), which is strong guidance. It does not explicitly name alternatives or say when not to use it, but the sibling tools are clearly unrelated in function, so the omission is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_commandsARead-only
Enumerate every slash command the server currently exposes as a prompt, each with its name and description. Use to discover available commands before routing a user request to one. Read-only manifest view, takes no arguments. Returns a count plus a commands array.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral detail beyond the readOnlyHint annotation: it labels itself a 'Read-only manifest view,' states it takes no arguments, and reveals the response shape ('Returns a count plus a commands array'). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two succinct sentences: the first states what it does, the second gives usage context, read-only nature, argument count, and return format. Every sentence earns its place with no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no parameters and no output schema, the description fully covers purpose, usage context, behavior, and return structure. Sibling relationships are implied through the tool name and the usage note, so no further context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, and the schema coverage is 100%, so there is nothing to misinterpret. The description reinforces 'takes no arguments,' matching the empty schema. Baseline for 0 params is 4, and no further parameter explanation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Enumerate') and resource ('every slash command the server currently exposes as a prompt'), and clarifies what's included ('name and description'). This clearly distinguishes it from siblings like list_rules and list_skills, which target different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use: 'Use to discover available commands before routing a user request to one.' This gives clear context. However, it does not mention alternatives or when not to use it (e.g., for rules or skills), so it lacks the explicit exclusion/alternative guidance of a top-tier score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_rulesARead-only
Enumerate every behavioral rule the server exposes as a resource, each with its URI, name, and description. Use to discover which rules are in effect, then fetch a body with read_resource_body or resources/read. Read-only manifest view, takes no arguments. Returns a count plus a rules array.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint: true. The description adds that it is a 'read-only manifest view', takes no arguments, and returns a 'count' plus 'rules' array. This offers concrete behavioral detail beyond the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first states purpose and result structure; the second gives usage and return format. No redundant wording, front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument list tool with no output schema, the description fully covers purpose, usage, and return shape. It even notes the next step (fetch a body), making it contextually complete relative to its simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, and the description explicitly states 'takes no arguments'. Since schema coverage is 100% trivially, the description adds confirmation of the empty parameter list. Baseline for zero params is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Enumerate' and identifies a distinct resource: 'behavioral rule the server exposes as a resource'. It explicitly lists what each rule includes (URI, name, description), distinguishing it from sibling tools like list_commands and list_skills.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear usage context: 'Use to discover which rules are in effect' and explicitly names alternative tools for a follow-up action (read_resource_body or resources/read). This tells the agent when to use this tool and what to do next.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_skillsARead-only
Enumerate every skill the server currently exposes as a prompt, each with its name, description, and source. Use to discover which skills are available before suggesting or invoking one. Read-only manifest view, takes no arguments. Returns a count plus a skills array.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation is reinforced by 'Read-only manifest view,' and the description adds behavioral details such as taking no arguments and returning a count plus a skills array. Since there is no output schema, explaining the return shape is valuable context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The three sentences are tightly packed: definition, usage context, and behavior/return. No waste or redundancy; each clause serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only listing tool, the description covers what it returns, when to use it, and its read-only nature. Given the lack of output schema, the return value explanation is sufficient, and the tool's simplicity means no further context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the input schema is an empty object. The description explicitly notes 'takes no arguments,' and with no parameters to document, the baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Enumerate every skill the server currently exposes as a prompt,' clearly identifying the action and resource. It distinguishes itself from sibling listing tools like list_commands and list_rules by focusing specifically on skills. The mention of name, description, and source further clarifies the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use: 'Use to discover which skills are available before suggesting or invoking one.' However, it does not mention when not to use or explicitly compare with alternative list tools, so it falls short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getARead-only
Batch-fetch FULL memory entries by id — the second half of the index-first retrieval workflow. Call memory_lookup with detail:"index" first, pick the ids whose title/tokens_estimate justify the fetch, then fetch them here in ONE batched call. Unknown ids are reported per-id (ids[]="unknown"), never failing the batch. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | Entry ids to fetch (from a detail:"index" lookup). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that unknown ids are reported per-id as 'unknown' and never cause the whole batch to fail—critical error-handling behavior that is not visible in the schema. It also reinforces the read-only nature and highlights the batching capability, adding genuine behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states the core purpose, the second gives the workflow and usage order, and the third details error behavior. There is no fluff or repetition; the description is tightly packed and front-loaded with the primary action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter, read-only tool with no output schema, the description fully covers purpose, usage, error handling, and the relationship to its sibling. The 'FULL' qualifier and the per-id unknown example give enough detail for an agent to anticipate the response shape without needing a formal output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes the 'ids' parameter as 'from a detail:"index" lookup', so baseline is 3. The description adds the selection criterion ('title/tokens_estimate justify the fetch') and emphasizes the batched nature of the call, providing marginal but meaningful extra guidance beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Batch-fetch FULL memory entries by id') and explicitly frames the tool as the 'second half of the index-first retrieval workflow', distinguishing it from sibling memory_lookup. It clearly states what the tool does and what it returns (full entries as opposed to index summaries).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit when-to-use workflow: call memory_lookup with detail:'index' first, select ids based on title/tokens_estimate, then fetch them here in ONE batched call. It also names the alternative (memory_lookup) and the order of operations, leaving no ambiguity about usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_lookupARead-only
Retrieve engineering-memory entries for one or more memory types, optionally narrowed to specific anchor paths. Use before editing a security-sensitive or historically buggy file to surface prior incidents, ownership, and patterns tied to it. WORKFLOW: call with detail:"index" FIRST — each row carries id, title and tokens_estimate (the cost of fetching it) — then fetch full bodies via memory_get ONLY for the ids you will actually use, batching multiple ids into one call. Reads agents/memory/<type>/*.yml plus the agents/memory/intake/*.jsonl signal log. Read-only. Returns the v1 retrieval envelope: a status field plus per-type slices carrying the matched entries.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | No | Optional anchor paths or globs to match entries against (e.g. a file you are about to edit). | |
| limit | No | Maximum entries to return per type. Defaults to 5. | |
| types | Yes | Memory types to scan, e.g. `historical-patterns`, `incident-learnings`, `ownership`. At least one required. | |
| detail | No | 'index' returns compact priced rows (id, title, tokens_estimate) instead of full bodies — call this first, then memory_get the ids you need. 'full' (default) returns complete entries. | full |
| token_budget | No | Optional token budget. When set, entries are rendered as one-line compact rows (id, type, confidence, `line`) and the row set is hard-cut at token_budget × 4 chars; omitted hits appear as a top-level `truncation` hint naming a concrete next step. Absent → the envelope is unchanged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the tool's read-only nature, which aligns with the readOnlyHint annotation but goes beyond it by naming the exact data sources ('Reads `agents/memory/<type>/*.yml` plus the `agents/memory/intake/*.jsonl` signal log') and describing the return envelope ('a `status` field plus per-type `slices`'). This provides valuable context about side effects, data scope, and output structure not present in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose, followed by usage context, workflow, data sources, and return format. Every sentence contributes value, and the 'WORKFLOW:' label provides clear structure. Despite its length, it is appropriately sized for the tool's complexity, with zero redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes responsibility for explaining return values, and it does so ('Returns the v1 retrieval envelope: a `status` field plus per-type `slices` carrying the matched entries'). It also covers the `detail` modes and token_budget behavior indirectly through the schema. Considering the tool's complexity and the rich schema, the description is complete enough for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining the strategic use of the `detail` parameter ('call with detail:"index" FIRST') and the purpose of `keys` ('anchor paths to match entries against'). While it doesn't add new meaning for every parameter, the workflow guidance enhances understanding of how to use the parameters effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Retrieve engineering-memory entries for one or more memory types, optionally narrowed to specific anchor paths.' It clearly distinguishes itself from sibling tools like memory_get by explaining its role as a lookup that returns indexes and summaries, with memory_get for full bodies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to use the tool: 'Use before editing a security-sensitive or historically buggy file...' It also provides a step-by-step workflow ('call with detail:"index" FIRST... then fetch full bodies via memory_get') and names the alternative tool (memory_get) for fetching full entries. This is explicit guidance about when and how to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_signalA
Record an engineering-memory signal — a short, anchored observation such as a recurring bug pattern or an ownership note — to the monthly intake log agents/memory/intake/signals-YYYY-MM.jsonl. Use to capture a learning tied to a specific file so future memory_lookup calls surface it. Appends to the filesystem and is rate-limited per (type, path) within a rolling window. Returns the recorded signal.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Free-form signal body — the observation to record. | |
| path | Yes | Repo-relative anchor path the signal is about. | |
| type | Yes | Memory type the signal belongs to (e.g. historical-patterns, incident-learnings, ownership). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false (write operation), but the description adds that it 'Appends to the filesystem', is 'rate-limited per (type, path) within a rolling window', and 'Returns the recorded signal'. This clearly discloses side effects and constraints beyond the annotation, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three focused sentences: front-loaded with the action and destination, followed by usage context, behavioral details, and return value. Every sentence adds value, with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with 3 required params and no output schema. The description covers purpose, target file path, append behavior, rate limiting, and return value, which is complete for an agent to decide when and how to invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for all three parameters (type, path, body). The description adds general context like 'short, anchored observation' and 'tied to a specific file' but does not provide additional per-parameter semantics beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Record' with a resource ('engineering-memory signal') and a concrete destination (the monthly intake log). It distinguishes from sibling memory_lookup by stating this captures learnings so future lookups can surface them, making the tool's role clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use it ('Use to capture a learning tied to a specific file') and gives examples (recurring bug pattern, ownership note). It implies the read counterpart is memory_lookup but does not explicitly exclude alternatives, which prevents a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_statusARead-only
Report the memory backend status. Memory is entirely file-backed (agents/memory/); there is no external backend. Read-only, takes no arguments. Returns a status (file), the active backend (file), and a short reason.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds valuable context: memory is entirely file-backed at a specific path, there is no external backend, and it discloses the exact return fields (status, backend, reason). This goes beyond annotation and helps the agent understand expected behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four short sentences, each adding distinct value: purpose, backend context, read-only behavior, and return format. It is front-loaded with the purpose and contains no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only status tool, the description fully covers what it does, how it behaves, and what it returns. It also explains the file-backed architecture, making the tool's behavior predictable even without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete. The description reinforces this by stating 'takes no arguments', which is sufficient. No additional parameter semantics are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Report the memory backend status', a specific verb+resource pair that clearly distinguishes this from sibling tools like memory_get or memory_lookup which retrieve data. It unambiguously conveys that this is a status query, not a data access operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: it is a read-only status check with no arguments, and the file-backed nature is explained. However, it does not explicitly mention when not to use this tool or point to alternative siblings, leaving the user to infer distinctions from the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_resource_bodyARead-only
Fetch the rendered body of a single resource URI (rule, guideline, or context document) in one call, without the two-step resources/list + resources/read handshake. Use when you already know the URI and want to inline its content into a tool-call result. Read-only. Returns the resource uri, name, description, and full text body.
| Name | Required | Description | Default |
|---|---|---|---|
| uri | Yes | Resource URI to fetch, e.g. `rule://commit-policy`, `guideline://php/patterns/events`, or `context://authority/scope-mechanics`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares the tool is read-only, and the description reinforces this, but more importantly it adds the return payload details ('uri', 'name', 'description', and full text body') which are not in an output schema. This provides useful behavioral context beyond the annotation, though it does not cover error cases or other edge behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose in the first sentence and usage/return details in the second. No fluff or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple single-parameter read-only tool with no output schema, and the description sufficiently covers the operation, usage condition, and return format. The sibling tools are unrelated, so no confusion. The description is complete for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes the only parameter 'uri' with examples and format. The description adds minimal additional meaning (e.g., 'single resource URI' and types) but does not go beyond what the schema already provides. Given 100% schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Fetch the rendered body of a single resource URI'), specifies the resource types (rule, guideline, context document), and distinguishes itself from the two-step resources/list + resources/read handshake. This is a specific verb+resource definition that sets it apart from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it ('when you already know the URI and want to inline its content into a tool-call result') and references the alternative two-step approach it avoids. This gives clear context and an explicit alternative, fulfilling the dimension perfectly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
roadmap_archiveA
Archive every roadmap that has reached count_open == 0 and was touched on the current branch — git mv to agents/roadmaps/archive/, migrate inbound references, and regenerate the dashboard. Use as the PR-gate sweep before opening a pull request. Mutates the git index (moves tracked files) but never commits or pushes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint: false annotation by detailing that it mutates the git index, moves tracked files, migrates references, and regenerates the dashboard, while explicitly noting it never commits or pushes. This is valuable behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: the first defines the action, the second gives usage context, and the third discloses side effects. It is front-loaded and has no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main actions and usage context. It does not mention return values or behavior when no roadmaps match, but given the absence of an output schema and the tool's mutation focus, the coverage is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter info. The baseline is 4 for zero-parameter tools. The description adds operational semantics but does not need to explain parameters, as none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (archive every roadmap with count_open == 0 touched on the current branch) and the resource (roadmaps). It also names the specific operations (git mv, migrate references, regenerate dashboard) and differentiates from siblings like roadmap_progress and memory_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use as the PR-gate sweep before opening a pull request,' providing a clear when-to-use context. It does not mention when not to use or explicit alternatives, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
roadmap_progressA
Regenerate agents/roadmaps-progress.md from the current checkbox state of every active roadmap. Use after landing roadmap work to keep the dashboard in sync without a shell round-trip. Writes the dashboard file inside the project tree. Set dry_run: true to compute counts without writing.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | When true, return the computed dashboard without writing the file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that the tool writes a file inside the project tree, which is useful behavioral context beyond the readOnlyHint:false annotation. It also explains the dry_run flag to avoid writing, adding transparency about how to perform a non-destructive run. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each serving a distinct purpose: main action, usage context, and dry-run behavior. It is front-loaded with the core purpose and contains no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one optional parameter, no output schema, and a clear write behavior, the description covers the essential aspects: what it does, when to use it, and how to avoid writing via dry_run. It could be slightly more explicit about return values, but overall it is complete for its simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully documents the `dry_run` parameter with 100% coverage. The description adds minimal extra meaning by stating it 'compute counts without writing', but this largely paraphrases the schema description. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool regenerates `agents/roadmaps-progress.md` from checkbox states of active roadmaps. The verb 'regenerate' and specific file path make the purpose unambiguous, and it differentiates from siblings like `roadmap_archive` by focusing on syncing the dashboard file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use it 'after landing roadmap work' and contrasts it with a 'shell round-trip', giving clear context on when and why to use the tool. It does not list exclusions or alternative tool names, but the guidance is sufficient for typical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_testsA
Run the consumer project's vitest test suite under a compiled safety envelope: fixed argv (no shell interpolation), 120s timeout, 64KB output cap per stream. Shell-exec pilot per the 2026-07-07 council cut — vitest projects only; other runners (Pest / PHPUnit, pytest, Jest) return an error until a future council round approves them. Pass filter (vitest --testNamePattern) or path (in-tree file or directory) to narrow the run. Returns runner, passed, exit_code, timed_out, truncated stdout/stderr, and duration_ms.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Restrict to tests under this directory. | |
| filter | No | Restrict to tests matching this name pattern. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the sparse annotation (readOnlyHint=false), the description reveals the compiled safety envelope: fixed argv, no shell interpolation, 120s timeout, 64KB output cap per stream. It also notes the Shell-exec pilot status and details the return fields, providing substantial behavioral context that annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences front-load the core action and safety constraints, then proceed to restrictions, parameter guidance, and return fields. Every sentence contributes new information with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional parameters and no output schema, the description covers purpose, safety envelope, supported/unsupported runners, parameter usage, and return value shape. The only minor omission is an explicit statement that passing neither parameter runs the full suite, but this is strongly implied by 'narrow the run.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the input schema already describes both parameters, the description adds mapping: `filter` maps to vitest --testNamePattern and `path` is an in-tree file/directory, clarifying exact usage. This goes beyond the schema's generic descriptions and helps an agent choose parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Run the consumer project's vitest test suite,' clearly identifying the tool's action and scope. It further distinguishes itself by stating it supports only vitest projects and that other runners return an error, making it unmistakable from a generic test runner.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit exclusions ('other runners ... return an error until a future council round approves them') and explains how to narrow runs with `filter` or `path`. However, it does not name an alternative tool to use for non-vitest projects, so it stops short of full when-to-use-vs-alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_skill_for_taskARead-only
Match a free-form task description to the most relevant skills, ranked by a deterministic keyword scorer over SKILL.md frontmatter. Use when a skill you need is not in the catalogue the host delivered — a measured host dropped 402 entries from its model-visible list — so asking by name is impossible while asking by task is not. Read-only: no shell, no writes, and no skill bodies are returned, only names, scores and declared personas.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Free-form description of the task to match skills against. | |
| limit | No | Maximum number of skills to return. Defaults to 5. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses behavioral traits beyond annotations: it states it is read-only, does not return skill bodies, only names, scores, and personas fraction. The annotation readOnlyHint: true is reinforced and expanded with details about no shell, no writes, and what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the main purpose, and each clause adds value. The mention of the host dropping 402 entries provides context but could be trimmed; still it's justified as it explains why this tool exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description informs the agent what is returned (names, scores, personas) and the read-only behavior. It is complete for a recommendation tool with two parameters that are fully described in schema. The additional context about the catalogue gap adds necessary usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already fully describes both parameters (task as string, limit with default and min). The description adds that it ranks by a deterministic keyword scorer, which hints at how task is used, but does not significantly enhance parameter understanding beyond the schema. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool matches a free-form task description to relevant skills via a deterministic keyword scorer over SKILL.md frontmatter, and it differentiates from siblings by focusing on task-based search rather than name-based listing (like list_skills). It is specific about verb, resource, and mechanism.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: when a skill is not in the catalogue the host delivered, and explains asking by name is impossible while asking by task works. It also implies when not to use (when you can name the skill) and gives a concrete scenario, distinguishing it from alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
telemetry_reportARead-only
Return the artefact-engagement telemetry report — essential / useful / retirement-candidate skills and rules ranked by recorded consult+apply signals over a rolling window. Use to see which artifacts are actually load-bearing. Read-only. No-op (empty report) when telemetry recording is disabled.
| Name | Required | Description | Default |
|---|---|---|---|
| window_days | No | Reporting window in days. Defaults to 30. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states 'Read-only', which is consistent with the readOnlyHint annotation, but goes beyond it by adding a significant non-obvious behavior: 'No-op (empty report) when telemetry recording is disabled.' This is valuable context not available in structured data. It also notes the 'rolling window' aspect, adding operational detail about data recency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exactly two sentences, front-loaded with the main action, and packs in the purpose, usage context, and an edge-case behavior. There is no redundant phrasing or filler. Every clause contributes to understanding, making it appropriately concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with one optional parameter and no output schema, the description covers the essential aspects: what the report contains, how it is ranked, when to use it, and what happens when telemetry is disabled. It falls slightly short of a 5 because it does not describe the expected return shape (list, object) or field names, but given the simplicity, it is sufficiently complete for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully describes the single parameter 'window_days' with a description and default. The description mentions 'over a rolling window' but does not add new details about the parameter itself. Since schema coverage is 100%, the baseline of 3 is appropriate; the description provides minimal additional semantic value beyond what the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Return' and identifies the resource as the 'arteffect-engagement telemetry report'. It details the content (skills/rules ranked by consult+apply signals) and clearly distinguishes it from sibling tools like list_skills or capabilities_index, which are general artifact listings. The inclusion of 'essential / useful / retirement-candidate' categories further clarifies the specific purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The sentence 'Use to see which artifacts are actually load-bearing' provides a clear use case for when to invoke this tool. While it does not explicitly name alternatives or exclusions, the focus on telemetry signals implies it is the right choice over generic listing tools. The context is sufficient for an agent to decide appropriately, but lacks an explicit 'when not to use' or alternative tool reference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool targets a distinct resource and action: memory lookup/get/signal/status, roadmap archive/progress, chat append/read, list commands/rules/skills, and separate health/check tools. Even overlapping diagnostics like doctor_report and conformance_check have clear, non-conflicting purposes.
All names use lowercase snake_case and follow a mostly verb_noun pattern (run_tests, list_skills, memory_get). Some noun-phrase names like memory_status and roadmap_progress are descriptive but consistent in style, with no camelCase or chaotic mixing.
At 19 tools, the set is somewhat large but justified by the broad scope of agent configuration (memory, roadmaps, chat history, diagnostics, listing). It sits at the upper boundary of reasonable, with each tool earning its place.
The surface covers core workflows: memory retrieval and intake, roadmap maintenance, chat history persistence, manifest listing, diagnostics, and linting. Minor gaps exist (no memory delete/update, no roadmap creation), but these are workable and likely outside the intended scope.
Maintenance
Related MCP Connectors
Governed AI agent skills — one library, distributed to devs and exposed to remote agents over MCP.
System-of-record notebook for AI coding agents: pages, datastores, tasks, skills over MCP.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Agent-native notes, tasks, dev-docs, vaults, sync & handoffs. MCP + OpenAPI dual surface.
Related MCP Servers
- AlicenseAqualityAmaintenanceCross-agent memory bridge for AI coding assistants. Persistent knowledge graph shared across 10 IDEs (Cursor, Windsurf, Claude Code, Codex, Copilot, Kiro, Antigravity, OpenCode, Trae, Gemini CLI) via MCP. 22 tools including team collaboration, auto-cleanup, mini-skills, session management, and workspace sync. 100% local, zero API keys required.91,884719Apache 2.0
- AlicenseAqualityCmaintenanceAgent-native knowledge infrastructure. Deterministic, vertical-specific knowledge bases for autonomous agent consumption via MCP. Ethics modules mapped to EU AI Act articles. Free 24-hour trial.72MIT
- AlicenseAqualityCmaintenanceAgent Toolbelt is an MCP server exposing 11 focused API tools for LLM agents — schema generation, text extraction, token counting, CSV conversion, Markdown conversion, URL metadata, regex builder, cron expressions, address normalization, color palettes, and brand kits. Each tool is a focused microservice with structured input/output, WCAG-scored color data, USPS address parsing, and multi-model to25371MIT
- FlicenseNot gradedqualityBmaintenanceA remote MCP server that ships procedural knowledge skills to AI agents, enabling methodologies like brainstorm-first and plan-before-action. Hosted on Cloudflare Workers, it provides resources, prompts, and tools for skill management.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/event4u-app/agent-config'
If you have feedback or need assistance with the MCP directory API, please join our Discord server