ARNO
Summary: ARNO is an MCP server that gives coding agents an IDE-like toolkit for a repository — symbol-aware reading, revision-guarded editing, atomic multi-edit transactions, and validation through the project's own build/test commands — all without the shell.
Read code (
read_range): files verbatim, whole or by line range ("280-400","280-"), or many ranges/files in one call; token-budgeted withcontinuepaging; dependency sources read-only viadep:<name>/<path>.Find symbols (
find): locate declarations by name and get their bodies in one call, filtered bykind, with multiplequeries,maxLines/budgetlimits, and dependency lookup.Search text (
grep): literal or regex search across the workspace withglob,context,exclude,ignoreCase, multiplequeries, and search inside locked dependency sources.Edit files: replace exact unique text (
replace_text), insert new functions/sections without replacing anything (insert, anchored or appended), create new files with parents (create_file), delete files (delete_file).Let the read discover the relevant code (
find/grep/read_range) rather than guessing and grepping through the shell.Apply many edits atomically (
apply): replace_text, replace_range, replace_symbol, delete_symbol and insert land all-or-none with anchors validated up front, formatting of touched files, and a single validation at the end.Guard against stale edits: every edit accepts
expectedRevisionand/orexpectedDigestpreconditions, so concurrent changes make the edit fail loudly instead of silently clobbering.Validate with the repository's own commands (
check): build (default), typecheck, tests, lint or codegen — discovered from the project's Makefile/npm/cargo/Maven/Gradle/sbt manifests, or declared;dryRunnames the command without running it;wait:falsereturns a job to poll.Run scoped tests (
run_tests): a specific test file, a test name, or the changed files' tests, returning pass/fail and the first failure.Record and replay repository verbs (
declare_command/run_command): declare lint/codegen/release-gate commands once into.arno/commands.json, then run them by name in this and later sessions.Work across git worktrees: most tools take an optional
root=<path>to act in a specific worktree of the repository for that one call.
Runs Gradle builds and tests for Java/Scala projects as validation, and allows declaring custom Gradle commands for repeated use.
Runs Makefile targets for build, typecheck, or test validation, integrating with the repository's own Make-based commands.
Runs npm scripts for validation and searches npm package sources as dependencies, enabling navigation and testing in Node.js projects.
Runs pytest tests with name filters and file scopes, providing pass/fail verdicts and first failure output.
Supports Python repositories with symbol-aware reading, editing, and running tests via pytest, enabling agents to navigate and modify Python codebases.
Supports Rust repositories with symbol-aware reading, editing, and running tests via cargo test, enabling agents to navigate and modify Rust codebases.
Provides symbol-aware navigation, exact references, cross-file rename, and type errors on edit for Scala projects via the Metals language server, and can run sbt builds/tests for validation.
Supports TypeScript repositories with symbol-aware reading, editing, and running tests via jest/vitest, enabling agents to navigate and modify TypeScript codebases.
Runs Vitest tests with name filters and file scopes, providing pass/fail verdicts and first failure output.
ARNO — Agent Repository Navigation & Operations
The IDE for agents
ARNO gives coding agents what an IDE gives you, served over MCP. Read by symbol, edit against a known revision, get the compiler's diagnostics back with the edit, validate with the repository's own commands, see what changed, and revert to a checkpoint — each step one tool call, none of it through the shell.
Symbol-aware reading is how ARNO finds its way around; the transaction is what it is for. Where the shell is still the better tool, use it — the question ARNO has to answer is whether an agent gets more done, at acceptable cost, with it than without (docs/benchmark.md).
Half the tokens, more issues fixed. Claude Code on 12 real closed issues from cobra (Go), ky (TypeScript), requests (Python) and ripgrep (Rust), judged by each upstream fix's own hidden tests:
Claude Code with… | Issues fixed | Tokens per task | Time per task |
its built-in tools | 10 of 12 | 1.38M | 215s |
ARNO in their place | 12 of 12 | 0.68M (−51%) | 163s (−24%) |
Not just a shorter tool list: against a shell trimmed to Bash, Read, Edit and Write, ARNO still used 18–25% fewer tokens and fewer turns, in two separate runs. Sonnet 5, one run per task, four languages — results and caveats · how to set it up.
Status: 0.0.14 shipped, the first release without the Jade names; 0.0.11 and earlier were Jade — early, usable, and looking for feedback. Testing it? Start with the tester guide.
curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | shFor macOS and Linux — Windows isn't supported (here's why, and where to upvote). Run it again to update. Other ways to install.
Related MCP server: Serena
Table of contents
Trying ARNO: a guide for testers
Thanks for testing. Half an hour gets you set up; the useful part is a week or two of your normal work with it switched on, and then telling us how it went — including if you turned it off.
1. Install ARNO and your language servers
macOS and Linux:
curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | shWindows isn't supported, sorry! If you'd like it to be, please 👍 issue #2 or tell us there why it matters to you. (WSL 2 runs Linux, so the Linux build may work there, but we don't test it or take bug reports for it.)
The script picks the build for your OS and CPU, checks it against the
release's checksums and installs it to /usr/local/bin, or ~/.local/bin
when that is not writable — no sudo, no Go toolchain. If that directory is
not on PATH, it offers to add it to your shell profile, and it prints the
absolute path to use as command in your MCP client config. On a first install it then opens
arno-mcp install, a menu that installs the language servers you pick, or
shows how to install them by hand; Enter skips it, and you can run it again
any time. Run the same command again to
update — ARNO tells you when a new release is out (see
Update check). (Prefer Go? go install github.com/julianbei/arno/cmd/arno-mcp@latest works too.) ARNO reads
structure in every language with nothing else installed; exact references,
cross-file rename and type errors on edit need the language's server —
arno-mcp install installs these for you, or by hand:
Language | Install | Notes |
Go |
| First answer about 1.6s. |
Java |
| First answer about 8.5s while it indexes. |
Scala |
| Set |
Kotlin | — | No grammar or server yet: text search only. Tell us if you need it. |
A missing server is never an error: ARNO says which answers are approximate.
2. Point your agent at a real repository
For Claude Code, put this in .mcp.json at the root of the repository you
work in (other hosts: Codex CLI, goose, OpenCode):
{
"mcpServers": {
"arno": {
"type": "stdio",
"command": "arno-mcp",
"args": ["--root", "/absolute/path/to/the/repo", "--tools", "core"],
"alwaysLoad": true
}
}
}That keeps Claude Code's own tools too. For the clearest signal, run some
sessions with ARNO in place of them: claude --tools "" with the same
config (why and trade-offs). Restart
or reconnect the client (/mcp) after installing or upgrading ARNO.
3. Check it came up
Ask the agent: "call arno.capabilities". You should see your languages, each
with server … (not started) or no server (… not installed), plus the build
and test commands ARNO found (mvn, gradle, sbt, go test). If a server
you installed shows as not installed, that is a bug report.
4. What to try
Work as you normally would. If you want a checklist for the first sessions:
Find and follow code: "where is X declared, and who calls it?" —
find,references.Rename across files: a method or class used in several files —
rename(exact with gopls, jdtls or metals running).A change in several places at once: "change the signature and update the callers, then check it builds" —
applywithcheck.Run the tests that matter:
run_testswith a file or test name, orapplywithcheck: "impact".Repeatable commands: have the agent
declare_commandsomething you run often (./gradlew :core:test,sbt "testOnly *ParserSpec"); later sessions reuse it from.arno/commands.json.Undo:
checkpointbefore something risky,revertif it goes wrong.
5. Tell us how it went
When | File this |
After a week or two — or when you turn ARNO off | |
The agent used the shell although an ARNO tool existed | |
A tool gave a wrong answer or failed | |
Something you wish ARNO did |
The templates ask for the output of arno.capabilities and, optionally,
arno.telemetry. Neither contains source code; telemetry is
local only and records no arguments or response text, so both are
safe to paste from a private repository.
Known rough edges on the JVM: no formatter runs for Java or Scala files; large Gradle builds can make jdtls's first answer much slower than 8.5s; Kotlin has no support yet.
Why ARNO exists
An agent that falls back to grep, sed and cat is operating outside any
tooling you control. No revision tracking, no guardrails, no telemetry, no way
to know what it did or why it chose to do it that way. Every shell fallback is
a hole in your visibility.
You cannot fix that by telling the model not to use the shell. The model uses the shell because the shell is cheaper — fewer tokens, fewer round trips, more flexible. So the only durable fix is to make the structural tool the cheaper option, and then measure whether you succeeded.
That is the entire bet, and it is testable. On this repository's own benchmark, ARNO answers seven realistic engineering questions in 0.85x the tokens of the equivalent shell commands. It was 5.63x before responses became plain text instead of JSON — see docs/response-style.md for what changed and why.
The counter-measurement matters as much. One question asked against an unrelated repository came out at 1.36x — worse than the shell — because an ambiguous symbol name forced an extra disambiguation call. Seven scenarios at home and one away disagree, both are honest, and the second is the one that predicts outside use. The 0.0.4 pilot on four outside repositories answers it at a larger scale: used in place of Claude Code's built-in tools, ARNO solved 12 of 12 real issues with 51% fewer tokens, and 18–25% fewer than a shell trimmed to four tools (docs/benchmark-results.md). ARNO is not finished.
ARNO's own development log (docs/feedback.md) records every time its author reached for bash instead, and why. The pattern it found was blunt: the fallbacks that survived longest each closed within two tasks of being named in the log — not when the tool shipped.
Install
With the install script
curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | shDownloads the latest release binary for your OS and
CPU, verifies it against checksums.txt, and installs it without sudo to
/usr/local/bin or ~/.local/bin. Run it again to update: a arno-mcp
already on PATH is replaced where it is, and nothing is downloaded when it
is already current. If the directory is not on PATH, it offers to add it to
your shell profile (ARNO_ADD_TO_PATH=1 does it without asking) and prints
the absolute path to use in your MCP client config. ARNO_VERSION pins a release tag;
ARNO_INSTALL_DIR picks the directory. Read it first if you like:
install.sh.
Windows isn't supported — see issue #2, and give it a 👍 if you'd like that to change.
Language servers: arno-mcp install
arno-mcp install # menu: pick what to install
arno-mcp install --list # what is installed, and how the rest would be
arno-mcp install --servers go,java,scala # install these, no questions
arno-mcp install --all # every missing server this machine can installThe menu lists each language server ARNO can use, whether it is installed, and
the exact command it would run — go install for gopls, brew install jdtls,
cs install metals, npm install -g for TypeScript and Pyright, rustup component add rust-analyzer, gem install ruby-lsp — and confirms before
running anything. A server with no installer on the machine gets instructions
for installing it by hand. The install script opens the menu after a first
install; ARNO_SKIP_SETUP=1 skips it.
For an agent, or any script — there is no terminal to answer a menu, so the same steps come without questions:
# install ARNO and chosen servers in one go
curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | ARNO_SERVERS=go,java sh
arno-mcp install --list --json # state of every server, as JSON
arno-mcp install --servers java,scala --dry-run # the commands, not run
arno-mcp install --servers java,scala # run them--list --json gives each server's key, whether it is installed and
where, the command that would install it on this machine, and manual
steps when there is none. arno.capabilities ends a missing server's line
with the command that installs it. Installing is deliberately not an MCP
tool: global package installs go through the agent's shell, where you approve
them.
Other ways to install
The install script above is the easiest way. These work too.
From Go
go install github.com/julianbei/arno/cmd/arno-mcp@latest # newest
go install github.com/julianbei/arno/cmd/arno-mcp@v0.0.14 # pinned to a tagLands in $GOBIN, or $(go env GOPATH)/bin if that is unset — which is
usually ~/go/bin, and is not on PATH by default. Add it if it is not
there, then confirm:
export PATH="$PATH:$(go env GOPATH)/bin"
arno-mcp --versionIf you would rather not touch PATH, use the absolute path in your MCP client
config instead of the bare arno-mcp shown below.
From a release binary
Prebuilt binaries for linux and darwin on amd64 and arm64 are attached to each
GitHub release, with a
checksums.txt alongside them. Each is built natively on its own platform —
ARNO links tree-sitter through cgo, so the linux builds need a reasonably
current glibc. On an older distro, build from source or use the container
image, which is statically linked against musl.
VERSION=$(curl -fsSL https://api.github.com/repos/julianbei/arno/releases/latest | sed -n 's/.*"tag_name": *"\([^"]*\)".*/\1/p' | head -n 1)
OS=$(uname -s | tr '[:upper:]' '[:lower:]')
ARCH=$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')
curl -fsSL "https://github.com/julianbei/arno/releases/download/${VERSION}/arno-mcp_${VERSION}_${OS}_${ARCH}.tar.gz" \
| tar xz
sudo mv "arno-mcp_${VERSION}_${OS}_${ARCH}" /usr/local/bin/arno-mcpAs an MCP bundle, for Claude Desktop
Since 0.0.11, each GitHub release
also carries arno-mcp_<version>.mcpb, one bundle with the binaries for macOS
and Linux on Intel and ARM. Open it with Claude Desktop, pick the repository
ARNO should work on, and it runs with the core tools; no terminal and no
Docker. Language servers still come from arno-mcp install, or run without
them on tree-sitter alone.
From source
git clone https://github.com/julianbei/arno.git
cd arno
make binary # bin/arno-mcp, version stamped from git describe
make install # or straight onto your PATHOptional: gopls
references and rename use gopls for their exact, compiler-resolved form.
Without it they still work — references degrades to a textual approximation
that says so in the response, and rename refuses rather than guessing.
arno-mcp install --servers go # or: go install golang.org/x/tools/gopls@latestThe same goes for every language below: arno-mcp install shows which servers
are installed and installs the rest (Language servers).
Configure your MCP client
ARNO is a stdio MCP server. Point your client at the binary:
{
"mcpServers": {
"arno": {
"type": "stdio",
"command": "arno-mcp",
"alwaysLoad": true,
"env": {
"ARNO_WORKSPACE_ROOT": "/absolute/path/to/the/repo/arno/should/work/on"
}
}
}
}ARNO_WORKSPACE_ROOT is the repository ARNO inspects and edits. It does not
have to be the ARNO checkout — pointing it somewhere else is the entire point.
A --root /path/to/repo flag takes precedence over the environment variable,
and ARNO prints which of the three sources it used (flag, env, working
directory) at startup, so an agent can never quietly operate on the wrong
repository.
ARNO works on a non-git directory and on a repository with no commits yet. In
both cases it says what is degraded — changes, diff, history and
checkpoint need git — and everything else keeps working.
Let ARNO replace the built-in tools
ARNO saves tokens when it replaces the agent's own tools, not when it is added next to them. Every turn resends the whole tool list, and in Claude Code the built-in tools are about 38k tokens of it. In ARNO's pilot benchmark (12 real issues in cobra, ky, requests and ripgrep, one run each):
Tools | Tasks solved | Tokens per run | Time per run |
Claude Code's built-in tools | 10 of 12 | 1.38M | 215s |
Built-in tools trimmed to Bash, Read, Edit, Write | 10 of 12 | 0.90M | 213s |
ARNO only, core profile | 12 of 12 | 0.68M | 163s |
Both, all built-in tools and ARNO | 11 of 12 | 1.50M | 191s |
Trimming the built-in list is most of the saving on its own. ARNO on top of that used 25% fewer tokens and 24% less time than the trimmed shell, and solved the two tasks both shell setups failed. Given both ARNO and every built-in tool, the agent used Bash for four calls in five and paid for both lists. To run ARNO in place of the built-in tools:
claude --tools "" --mcp-config arno.jsonwith arno.json passing the core profile, which lists twelve tools — read,
edit, validate, and run the commands a repository declares — and keeps the
others callable:
{
"mcpServers": {
"arno": {
"type": "stdio",
"command": "arno-mcp",
"args": ["--root", "/absolute/path/to/the/repo", "--tools", "core"],
"alwaysLoad": true
}
}
}The trade-off is real: without Bash the agent cannot run arbitrary commands.
run_tests, check and repository commands declared with declare_command
cover building, testing and repeatable scripts — a declaration lives in
.arno/commands.json, so later sessions reuse it; a task that needs git operations, network access
or ad-hoc scripts needs the shell back. The numbers above are one run per task
— see docs/benchmark-results.md for the results,
a rerun after the pilot's fixes, and the caveats.
Other hosts
Verified with a live session — find a declaration, insert beside it, run
check — using only ARNO's tools:
Codex CLI (0.154), in ~/.codex/config.toml:
[mcp_servers.arno]
command = "arno-mcp"
args = ["--root", "/absolute/path/to/the/repo", "--tools", "core"]Codex asks before every MCP tool call. With approval_policy = "never" it
refuses them outright; run codex exec --approve-for-me or approve ARNO's
tools interactively.
goose (1.50), for one run:
goose run --with-extension "arno:arno-mcp --root /absolute/path/to/the/repo --tools core" -t "…"or permanently with goose configure → Add Extension → Command-line
Extension, command arno-mcp --root /absolute/path/to/the/repo --tools core.
Verified through goose's claude-code provider.
OpenCode (1.18), in opencode.json at the repository root or in
~/.config/opencode/opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"arno": {
"type": "local",
"command": ["arno-mcp", "--root", "/absolute/path/to/the/repo", "--tools", "core"],
"enabled": true
}
}
}OpenCode prefixes tools with the server name, so they appear as
arno_arno_find and so on. Verified with the github-copilot provider
(Claude Sonnet 5).
Cline and Gemini CLI are not verified yet.
Three things that will confuse you once
Without "alwaysLoad": true, Claude Code may never use ARNO. Claude Code
hides MCP tools behind a tool search by default: the agent sees their names but
not their definitions, and has to search before it can call one. With its own
shell and file tools right there, it does not. In ARNO's benchmark, an agent
given both ARNO and the shell made no ARNO call in three of three runs; the
same setup with alwaysLoad called ARNO directly. Other hosts may have their
own equivalent — check that ARNO's tools are actually being called.
One server serves one repository, but any of its worktrees. ARNO is
pinned to the directory it was started in, and no cd in your shell reaches
it — MCP carries no per-call working directory. A session working in another
git worktree of the same repository says so once with
arno.workspace path=<worktree>; its subagents inherit that. A subagent in a
different worktree passes root=<path> on its own calls. For fleets, set
ARNO_REQUIRE_WORKSPACE=1 so a write from a session that never said which
worktree it works in is refused instead of landing in the shared checkout. See
docs/worktrees.md.
The MCP tool catalog is fixed at connection time. A newly added tool does not appear until the client reconnects. If you upgrade ARNO mid-session and a tool seems missing, reconnect before investigating.
Use the binary, not go run ./cmd/arno-mcp. A go run stanza recompiles at
every process start: measured here at 284–584ms to first handshake against 14ms
for the binary, with a warm build cache. A cold one is seconds. It also means
the server silently changes whenever the source does — useful while hacking on
ARNO itself, confusing everywhere else. This repository's own .mcp.json
deliberately still uses go run for that reason.
Environment variables
Variable | Effect |
| Repository to operate on. Overridden by |
| Emit machine-readable JSON instead of plain text. |
| Disable local usage recording entirely. |
| Keep the telemetry log outside the workspace, one subdirectory per workspace. |
| Let metals import an sbt build so Scala edits get diagnostics. Runs sbt; creates |
| Turn off the daily check for a newer release (Update check). |
| In a repository with several worktrees, refuse writes until the session has named its worktree with |
| How many git worktrees one server keeps open at once (default 3, least recently used evicted first). See docs/worktrees.md. |
The install script reads its own:
Variable | Effect |
| Release to install, e.g. |
| Where to put |
| Language servers to install afterwards without a menu: |
| Skip the language-server step. |
| Add the install directory to the shell profile without asking. |
| Base URL of the releases, for a mirror. |
Use it in a container
ARNO is a child process, not a service, so the useful shape is to copy the binary into your own image rather than run ARNO's:
FROM ghcr.io/julianbei/arno-mcp:latest AS arno
FROM your-project-base
COPY --from=arno /arno-mcp /usr/local/bin/arno-mcp
ENV ARNO_WORKSPACE_ROOT=/workspaceOr build it yourself from the included Dockerfile.
The published image is distroless/static, so it carries no git, no gopls and
no language toolchains. ARNO detects each of those at runtime and degrades with
an explicit message rather than failing, so this still works — you get the
textual references fallback, and changes/diff/history/checkpoint are
off. If you want the full surface, install git and gopls in your image;
ARNO will find them.
The tools
31 tools, in four groups. Every response is plain text, shaped to lead with the decisive line — the answer first, the supporting detail after, raw output only when you ask for it.
Over MCP each one is registered as arno.<name> — arno.outline,
arno.replace_symbol, and so on. The tables below use the bare name for
readability. Most MCP clients show you the prefixed name already, often with
the dot rewritten (Claude Code displays mcp__arno__arno_find). Over the wire
ARNO accepts both arno.find and arno_find.
Inspect
Tool | What it does |
| What ARNO can do in this workspace: per language, grammar or text scan, language server state, formatter; git, validation commands, declared commands. Call it first. |
| File structure — declarations grouped by kind, without reading bodies. |
| Verbatim lines, or a whole file. |
| Locate a declaration and get its body in one call. |
| Literal or regex text search with path globs. The replacement for |
| Find usages. Exact from the language server when one is installed; a name-matched approximation otherwise, and it says which answered. |
| Pull a working set for a query. |
| Assemble the surrounding context for one symbol. |
| Directory structure. |
Symbols are addressed as path::Name, or path::Name@line when a name is
ambiguous. An ambiguous read returns the candidates with their signatures
rather than guessing.
Modify
Tool | What it does |
| Replace a whole declaration. Takes the full |
| Replace exact, unique text. Anchored on content, not line numbers. Like every text edit, returns the edited region as it now reads. |
| Replace an entire file's contents. |
| Create a new file. |
| Delete one file; directories are refused. |
| Delete one declaration. |
| Cross-file rename from the language server; refuses rather than guessing when it cannot be exact. |
| Add text without replacing anything — a new function, a new section, an extra case. Appends with no anchor; places before or after a unique anchor with one. |
| Several edits as one atomic unit — anchors validated up front, all applied or none, one revision bump and one validation at the end. |
Every edit returns consequences, not "success": the revision transition, which
symbols moved, immediate diagnostics, and the IDs of any background validation
it started. Edits accept an expectedRevision precondition; supplying it makes
a stale edit fail loudly instead of silently clobbering a concurrent change.
Validate
Tool | What it does |
| Build, typecheck or tests — discovering the repository's own command rather than assuming one: Makefile target, then npm script, cargo, Maven, Gradle, sbt, pytest/mypy or bundler, by manifest. A project it cannot identify is reported as such rather than run with the wrong toolchain. Every result names the command that ran; |
| Tests scoped to a file, a test name, or the changed files. |
| Run one of the repository's declared commands by name. |
| Add or remove a declared command. |
| Poll a background job. |
| Raw output for a job, on demand. |
Exit status is authoritative. A command that prints a success-looking line and exits non-zero fails.
State
Tool | What it does |
| What moved — by file and by symbol, not just by path. |
| The patch, including untracked files. |
| Which commits touched one symbol, via |
| Mark a revertible point: snapshots the files ARNO edited and records git's |
| Restore those files to a checkpoint. Never moves git, and refuses if a commit landed since the checkpoint. |
| The workspace event stream. |
| How ARNO's own tools have been used in this workspace. |
| List this repository's git worktrees, and change which one the session serves by default. |
Every tool above except workspace, events, telemetry, job_status and
job_output also takes an optional root=<path>: one of this repository's
git worktrees, to act in for that one call. See
docs/worktrees.md.
Project configuration
Discovery guesses how to build and test a repository from its manifests, and the
guess is sometimes wrong: a Makefile's python -m pytest picks the system
interpreter, npm run test runs lint and browser suites for a one-file check,
and a Go module with a TypeScript app beside it has two answers to "build". A
committed .arno/project.json states the answer once, the way an editor keeps
its settings in .vscode/:
{
"areas": [
{ "path": ".", "language": "go",
"build": "go build ./...", "test": "go test ./...",
"testName": "go test -run {name} ./..." },
{ "path": "web", "language": "typescript",
"typecheck": "node_modules/.bin/tsc --noEmit",
"test": "node_modules/.bin/vitest run",
"testFile": "node_modules/.bin/vitest run {file}",
"testName": "node_modules/.bin/vitest run {file} -t {name}" }
],
"env": { "python": ".venv/bin/python", "vars": { "CI": "1" } },
"generated": ["web/dist", "*.pb.go"],
"notes": "Browser tests need Playwright; run unit tests by file."
}Areas are parts of the repository with their own language and commands, each run inside its path.
checkruns a kind in every area that declares it;run_testswith a file uses the deepest area containing that file, with{file}relative to the area and{name}the test name.Empty fields fall back to discovery, so a config can state only what discovery gets wrong. A config that does not parse, names an unknown field or a path outside the workspace fails the check instead of being ignored.
envputs the interpreter's directory first onPATHand adds the variables to every command.generatedpaths are skipped bygrep,findand the workspace tree.notes, with the list of areas, is sent to the agent when a session starts.checkwithdryRunsays when a command comesfrom .arno/project.json.
arno-mcp init --root /path/to/repo drafts the file from discovery, one area,
for you to review and commit; it never overwrites an existing one.
Repository commands
Beyond build and test, every repository has its own verbs — lint, codegen, migrate, release-gate — and an agent that does not know them reaches for the shell. So ARNO lets it record them instead:
declare_command(name: "lint", run: "golangci-lint run ./...")
run_command(name: "lint")They live in .arno/commands.json, which is meant to be committed. It becomes
the repository's declared command vocabulary — written once by whoever (or
whatever) worked out the incantation, replayed by name forever after. Calling
run_command with no name lists what the repository declares; calling it with
an unknown name answers with the commands that do exist, so a wrong guess
teaches rather than fails.
Both answers also list detected candidates: npm scripts and Makefile
targets ARNO can see in the repository's own manifests but nobody has
declared. Detection is read-only — nothing is written until declare_command
says so — and a script check already runs for build, typecheck or tests is
left off the list, since it is not a gap. capabilities shows the same
candidates alongside what it already reports. This closed a failure class the
2026-09-16 benchmark measured directly: an agent guessing run_command names
(test-unit, model_formsets_tests) against a repository that had declared
nothing, each guess a wasted call, when the manifest already answered it.
A validation chain
Repository rules — Semgrep, a custom linter, a licence check — belong in
validation, and they need no integration in ARNO. Declare one command that
runs the steps in order, joined with &&:
{
"validate": {
"run": "go test ./... && semgrep scan --config .semgrep.yml --error",
"description": "tests, then repository rules"
}
}or, without editing the file, declare_command(name: "validate", run: "…").
run_command(name: "validate") runs it inside ARNO, so the run is part of the
session's record. Exit status decides: a rule that fails fails the run, the
steps after it do not run, and the summary leads with the failing output.
Use semgrep scan --error or the equivalent flag of your tool — a tool that
prints findings and exits 0 passes.
Language support
Structure comes from tree-sitter grammars compiled into the binary, so it
works with nothing installed. Semantics come from a real language server,
which you provide — arno-mcp install installs it for you — and ARNO starts
it on first use, reuses it for the session, and shuts it down on exit.
Language | Structure | Semantics, with this installed |
Go | ✅ built in |
|
TypeScript / TSX | ✅ built in |
|
JavaScript | ✅ built in |
|
Rust | ✅ built in |
|
Python | ✅ built in |
|
Ruby | ✅ built in |
|
Java | ✅ built in |
|
Scala | ✅ built in |
|
Everything else | text scan, announced | — |
Every row is verified end-to-end by make conformance, which builds an image
containing all eight servers and runs ARNO against a real repository per
language.
Semantic requests wait for the server to finish indexing (its $/progress
tokens), because an indexing server answers wrongly rather than slowly. When
the primary server declines a rename, ARNO asks the language's installed
alternative: ruby-lsp renames classes but not methods, so Ruby method rename
needs solargraph installed alongside it. With ruby-lsp alone, method
rename refuses and repeats the server's reason.
Every edit response names what checked the file (checked: pyright-langserver)
or why nothing did (not checked: app.py: pyright-langserver is not installed).
Scala needs one opt-in. metals reports errors only after importing the sbt
build, and it asks permission first, because importing runs sbt and creates
.bloop/ and .metals/ in the repository. ARNO declines unless
ARNO_METALS_IMPORT=1 is set, and says so in the edit response. A repository
an editor has already imported needs no setting.
"Structure" is outline, symbol read, edit-by-symbol, grep and search.
"Semantics" is exact references, cross-file rename, and type-level
diagnostics on edit.
ARNO looks for servers on PATH and in the places toolchains actually install
them — ~/go/bin, ~/.cargo/bin, ~/.local/bin, ~/.coursier/bin — because
go install puts gopls somewhere that is not on PATH by default, and a
client that only checked PATH would report Go as unsupported on a machine
that has a working gopls.
A missing server is never an error. ARNO degrades to the behaviour above and says which answer you got.
Design principles
Structure before source. Return the minimum sufficient representation first — outline before full source, summary before raw logs.
Deterministic tools before model reasoning. ARNO orchestrates tree-sitter, git, gopls and the project's own build tooling. It does not reimplement them, and does not guess where they could answer.
Every edit has a precondition and returns consequences. An edit can name the revision it expects and is refused if ARNO's revision has moved; it returns the revision transition, what changed and the diagnostics — not "success". Changes made outside ARNO do not yet move the revision (release plan Phase 5).
Conclusions before logs. The verdict leads. Raw output expands on request.
Semantic operations before textual ones. But textual escape hatches stay available, because the semantic path does not always exist.
State is explicit. Revisions, checkpoints and change sets are objects, not implications.
Validation waits by default, backgrounds on request.
check,run_testsandrun_commandreturn the verdict; a long run can return a job to poll instead.An approximation must announce itself. When ARNO falls back to a text scan or a name-matched graph, the caveat travels with the data, in the response — not in documentation the agent will never read.
ARNO is model- and harness-independent. MCP is an adapter, not the architecture.
Measure agent outcomes, not infrastructure sophistication. Tokens and turns per completed task — and token reduction is worthless if the success rate drops with it.
Repository-native execution. Builds, tests and lint run through the repository's own commands — discovered, or declared in
.arno/commands.json— inside ARNO, so validation is part of the record instead of a shell side trip.Cheaper than the escape hatch. If the shell is easier, faster and cheaper for a workflow, ARNO has failed that workflow. The benchmark, not opinion, says which (docs/benchmark.md).
The longer design document is docs/scope.md.
When the shell is still the right tool
ARNO does not try to match the shell's composability. Using the shell is a decision, not a leak, when the work is one of these:
Git operations: commit, branch, rebase, push, blame. ARNO reads git state (
changes,diff,history) and never moves it.One-off probes:
curla local server, inspect a process, check a port, read an environment variable.Debugging a script or a build system itself, where the question is what a shell command does rather than what the code says.
Installing dependencies and toolchains:
npm install,go install,pip install.Network access of any kind.
What stays on ARNO's side of the line, even though a shell could do it:
Builds, typechecks, tests, lint and codegen. Run them with
check,run_testsor a declared command (declare_command, thenrun_command). Validation run from the shell is validation the change transaction cannot see: no verdict in the edit record, no scoped test runner, no failure summary.Reading and searching code, and editing it. That is where ARNO's revisions, diagnostics and provenance apply.
A command you keep running from the shell for validation belongs in
.arno/commands.json, or in .arno/project.json as an area's build or test
command.
What ARNO does not do yet
Windows. ARNO is built and tested for macOS and Linux only, and we'd rather do those two really well than three halfway. Until further notice we don't build, test or look at Windows. If you'd like ARNO on Windows, please 👍 issue #2 — and if you think this is the wrong call, say so there; honest feedback is welcome. WSL 2 runs Linux, so the Linux build may work there, but it isn't tested.
This list is more useful than the feature list — it tells you what is worth reporting and what is already known. What is planned is in ROADMAP.md.
Nine languages get a real grammar; the rest fall back to a text scan. Go, TypeScript, TSX, JavaScript, Python, Ruby, Java, Scala and Rust are parsed properly. Anything else (Kotlin, Swift, C/C++, C#, PHP, …) is served by a heuristic that finds some declarations and misses others — and the amount it misses varies enormously by language, so treat those outlines as a hint rather than an inventory. ARNO always says which you got (
! no kotlin grammar — …).Semantic features need a language server installed for that language. ARNO speaks LSP to whatever is on the machine (see Language support). With a server,
referencesandrenameare compiler-exact and cross-file. Without one,referencesdegrades to a textual approximation that says so, andrenamerefuses rather than guessing — an approximate reference list is still useful to a reader, but an approximate edit is corruption.No completion, hover or code actions. ARNO's LSP client implements what the tools need — references, rename, diagnostics — not the whole protocol.
No blame, no cross-repo work, no remote execution.
Formatting follows the project's own rules, and only for the files you edited. gofmt and rustfmt always run. prettier (TypeScript/JavaScript), black or ruff (Python) and scalafmt run only when the repository declares them — a prettier config or
package.jsonkey,[tool.black],[tool.ruff.format]or a[format]section inruff.toml,.scalafmt.conf— using, for Node and Python, the project's own binary, because a formatter the project did not choose turns a one-line edit into a whole-file diff. rustfmt reads the nearestrustfmt.tomland the edition fromCargo.toml(or the workspace), and formats the one file: it does not followmoddeclarations into files you did not touch. Ruby, Java, JSON and Markdown are left as edited, and ARNO does not interpret.editorconfig.Revision tracking is ARNO's own counter, not git's. It detects concurrent edits within a session. It is not a VCS. A checkpoint snapshots the files ARNO has edited and records git's
HEAD;revertrestores those files and nothing else, never moves git, and refuses once a commit has landed since the checkpoint — undoing committed work is git's job. Checkpoints do not survive a restart of the server.Not hardened for untrusted input. It runs shell commands you declare and edits files you point it at. Treat it as a development tool, and do not point it at a repository you would not run
makein. What running ARNO inside a sandbox or container does and does not cover:Covered by ARNO itself: reads and writes stay inside the workspace root, symlinks included; dependency sources are read-only; a repository cannot make ARNO launch a binary it ships (declared commands run through the shell you already trust, and a
.arno/project.jsoninterpreter is a path you review in the diff).Covered only by the sandbox: what a declared command, a Makefile target, an npm script or a test suite does when
check,run_testsorrun_commandruns it — network access, files outside the workspace, credentials in the environment. ARNO runs the repository's own commands with your environment; a malicious repository'smake testis as dangerous under ARNO as in your shell.Not covered at all: an agent asked to declare a harmful command, and language servers, which execute project configuration of their own (build scripts, plugins) when they index a workspace.
Stability and versioning
Tool names and required arguments are frozen and enforced by a test. 0.0.3
added no tools and made two arguments optional (query on find, path on
read_range), both backward compatible. Schemas and server instructions are
still read once at connection time, so reconnect after upgrading.
docs/tool-contract.md has the full surface and the
policy on what counts as a breaking change.
What is not frozen: response wording, the .arno/* file formats, the
exact spelling of symbol IDs, and everything under internal/. Treat responses
as text for a model to read, not as a format to parse. ARNO_JSON=1 gives
machine-readable output if you need to parse something.
Telemetry
ARNO records how its own tools are used — call counts, response sizes, timing, and the failure classes that most often precede a caller giving up and using the shell.
It is written to .arno/telemetry.jsonl in your workspace and never
transmitted anywhere. It records no arguments, no response bodies and no
error text — only a 10-character hash of each call's target (path, symbol or
query), so the confusion report can tell a second tool asked about the same
thing. ARNO_TELEMETRY=0 turns it off; telemetry(reset: true) clears it.
ARNO tries not to leave files in a repository it was only asked to work in:
In a git repository, before creating the log, ARNO adds it to
.git/info/exclude— the clone-local ignore file, never committed — unless git already ignores it..gitignoreis never touched. The log does not show up as untracked, so a harness that commits every untracked file does not commit it.ARNO_STATE_DIR=/some/dirmoves the log out of the workspace entirely, into a subdirectory per workspace. Use it when ARNO is rooted at a checkout that something else commits or reviews wholesale.A call ARNO rejects outright (an unknown tool name) never creates the log.
.arno/commands.json is different: it is the repository's declared command
vocabulary, meant to be committed, and is only created when you declare a
command.
Update check
Separate from telemetry, ARNO looks up the newest release tag on GitHub — one
unauthenticated request for releases/latest, carrying nothing about your
workspace or how you use ARNO — at most once a day, in the background, with a
three-second timeout. A failed or offline check also waits a day. When a newer
release exists, it says so only where you asked what you are running:
$ arno-mcp --version
arno-mcp v0.0.12
update available: v0.0.13 (running v0.0.12) · curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | sh, then reconnect your MCP clientand as the second line of arno.capabilities. It never appears in the server
instructions or in other tool responses. ARNO_UPDATE_CHECK=0 turns it off;
it is also off in CI (CI set) and for development builds.
It exists because response cost is invisible to whoever is reading the response. Its first live reading found a tool returning 4.6KB in 704ms on a routine call — something sixteen tasks of hand-written notes had never noticed.
Reporting problems
docs/reporting.md says what makes a useful report. There are four issue templates:
feedback — how it went after some real use, or why you turned it off.
bug — it did the wrong thing.
friction — "I used the shell instead." This is the valuable one.
feature wish — it should be able to do X.
If you are unsure which, pick friction. It is the cheapest to write and the easiest to act on, and "it was just habit" is a real answer — we want it. Every shell fallback is a place ARNO was not worth reaching for, and that is the only signal that reliably improves it.
Before filing a bug, check whether your client has reconnected since the version changed. A stale tool catalog explains a surprising share of "this tool does not exist" and "my fix did not take effect".
Development
make build # go build ./...
make test # go test ./...
make fmt # gofmt -w ./cmd ./internal
make binary # bin/arno-mcp, version-stamped
make install # onto your PATHThe repository declares its own commands in .arno/commands.json, including
release-gate — build, vet, tests and a gofmt check, which is the gate a tag
has to pass. Run it the way an agent would: run_command(name: "release-gate").
Layout:
Path | What lives there |
The MCP stdio server — the entry point that matters. | |
A small CLI for driving the internal API directly. | |
The token benchmark: ARNO against equivalent shell commands. | |
Revisions, change sets, checkpoints, git. | |
Symbol index, outlines, search, grep, references. | |
Mutation, atomic apply, formatting. | |
Immediate feedback on edits; gopls. | |
Async job runner, command discovery. | |
Per-language adapters (Go, TypeScript, Rust). | |
The declared-command registry. | |
Local usage recording. | |
MCP adapter, and the transport-independent internal API. | |
Shared request and response types. |
Contributions are welcome. The one hard rule is principle 8: if you add a code path that approximates, the response has to say so.
License
Apache License 2.0 — see LICENSE. Copyright 2026 Julian Amelung.
Available Tools
12 toolsarno.applyADestructive
Apply several edits as one atomic unit: all land or none do. Ops: replace_text, replace_range, replace_symbol, delete_symbol, insert. Anchors are validated before anything is written, touched files are formatted, and one validation runs at the end instead of one per edit. Prefer this over several single edits when changing more than one site.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| check | No | Run one validation after all edits: build, typecheck, tests (the edited files' tests), or impact (those plus tests of callers of touched declarations). | |
| edits | Yes | Edits to apply in order. | |
| format | No | Format touched files afterwards (default true). | |
| expectedRevision | No | Revision expected before editing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only destructiveHint=true available, the description carries real weight: it discloses atomic all-or-nothing semantics, that anchors are validated before any write occurs, that touched files are auto-formatted, and that validation is batched once at the end. These are behavior facts the annotation does not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the atomicity guarantee, then ops, then execution traits, then the routing hint. No filler and nothing repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the essentials an agent needs: atomicity/failure semantics, pre-write anchor validation, post-write formatting, batched validation, and when to prefer it. No significant gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description nevertheless adds meaning by explaining that validation runs once after all edits (vs. per edit) and that formatting applies to touched files, which clarifies how the check/format parameters behave at execution time.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (apply edits) plus the defining property (one atomic unit: 'all land or none do') and enumerates the supported ops. This is immediately distinguishable from single-edit siblings such as arno.replace_text, arno.insert, and arno.delete_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Prefer this over several single edits when changing more than one site,' which names both the alternative approach and the condition that selects this tool. It also describes the validation model (one check at the end rather than per edit), giving the agent a reason to consolidate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.checkA
Run a validation command on demand and wait for the verdict: kind build (default), typecheck or tests. Uses the repository's own Makefile target, npm script or cargo command when present. Waits by default and returns pass/fail directly. Every result names the command that ran; dryRun names it without running anything. In a repository with several projects, pass target to check one; with no command at the root, the answer lists the projects.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | build (default), typecheck, tests, lint or codegen. lint and codegen run the declared commands of that kind; lint with none declared runs typecheck. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| wait | No | Wait for the result (default true). False returns a job ID to poll. | |
| dryRun | No | Name the command that would run, without running it. | |
| target | No | Project directory inside the workspace to check, e.g. services/api. Omit for the workspace root. | |
| timeoutSeconds | No | Bound on the wait (default 90, max 300). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only destructiveHint=false in annotations, the description carries the behavioral load and does it well: it waits by default and returns pass/fail, always names the command that ran, dryRun names without running, and a root without a command returns a project list. It does not cover timeout/failure behavior in depth, but the core execution semantics are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and defaults, then the fallback behaviors. Dense and mostly waste-free, though several clauses duplicate what the schema already documents, which slightly dilutes it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain returns, and it does: pass/fail verdict, the command name in every result, and a job ID when wait=false. Combined with the multi-project fallback, an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description mostly restates schema content (wait default, dryRun meaning, kind default); its one genuinely additive point is the multi-project/target behavior and the fallback listing when no root command exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Run a validation command') and enumerates the validation kinds it supports (build/typecheck/tests). It is clearly a validation runner, but it never names how it differs from close siblings like arno.run_command, arno.run_tests or arno.declare_command, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: waits by default, pass target to check one project in a multi-project repo, and dryRun to preview. There are no explicit exclusions or named alternatives, but the conditions under which each mode applies are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.create_fileA
Create a brand-new file, and any missing parent directories. Refuses to overwrite an existing one — use replace_text or apply to modify existing content.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path to create. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| content | Yes | File content. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=false; the description supplies the genuinely non-obvious behaviors — implicit parent-directory creation and a hard refusal on existing paths. It stops short of describing the failure mode (error vs. silent no-op) or encoding/permission requirements, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the create semantics are front-loaded and the exclusion/alternative routing follows immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple create tool with no output schema, the description covers the essential behaviors an agent needs: parent creation, overwrite refusal, and alternatives. It omits the return value and error surface, which are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so path, root, and content are already documented by the schema. The description adds no format, encoding, or path-resolution detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a brand-new file') plus a notable scope behavior ('and any missing parent directories'), so the agent knows this is a create-with-parents operation, not a plain write. It explicitly names the siblings that handle the mutation case (replace_text, apply), making it distinguishable from them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when-not condition ('Refuses to overwrite an existing one') and names the two alternative tools for modifying existing content. The routing rule for create vs. modify is fully spelled out with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.declare_commandADestructive
Declare a named command in .arno/commands.json — a reproduction, a benchmark — to run with run_command in this and later sessions; the file is reviewed like any change. Declaring an existing name replaces it; do not redeclare one just to run it.
| Name | Required | Description | Default |
|---|---|---|---|
| run | No | Shell command to run from the workspace root. | |
| kind | No | Optional: lint or codegen; check with that kind runs it. | |
| name | Yes | Command name: lowercase letters, digits, ':', '_' or '-'. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| remove | No | Delete the named command instead of declaring it. | |
| description | No | Optional note on what the command is for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only destructiveHint=true supplied, the description carries the rest and does so well: it discloses persistence across sessions, that the target file is 'reviewed like any change' (i.e., a tracked, reviewable mutation), and that an existing name is overwritten rather than duplicated. It does not contradict the destructive hint. Remaining gaps are minor — no word on failure modes if the name is invalid or the workspace is unset.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and artifact, followed by the operational rules. Every clause carries distinct information (persistence, review, overwrite, non-redeclaration) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter declaration tool with full schema coverage, no output schema, and a destructiveHint annotation, the description supplies the missing pieces an agent needs: persistence scope, overwrite semantics, and the routing rule to run_command. Nothing essential is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (including run, kind, root, and the remove toggle) are already documented in the schema. The description adds no syntax, format, or interaction detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Declare) and resource (a named command in .arno/commands.json), and gives concrete examples of what such commands are (a reproduction, a benchmark). It explicitly separates itself from the sibling run_command by framing declaration as the setup step that run_command later consumes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative and the condition: commands are 'to run with run_command', and it states an explicit when-not — 'do not redeclare one just to run it.' The replacement rule for existing names is also given, so the agent knows redeclaration is only warranted when changing the definition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.delete_fileADestructive
Delete a file, not a directory; fails if it does not exist. Reverting to an earlier checkpoint recreates it. To edit content use replace_text or apply.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path to delete. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare only destructiveHint=true; the description adds genuinely new behavior: the tool errors when the path is absent, and deletion is recoverable because reverting to an earlier checkpoint recreates the file. That recoverability note is the single most decision-relevant fact for an agent about to destroy state, and it is not available anywhere else.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, front-loaded with the core action and its main constraint, then failure behavior, then the alternative-tool routing. No filler and nothing repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive tool with no output schema, the description covers the essentials: what is deleted, what is not (directories), failure mode, recoverability, and where to go for the adjacent edit operation. Nothing an agent needs to invoke it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (path, root), so the schema already carries the semantics. The description contributes only the indirect file-vs-directory constraint on 'path' and says nothing about the 'root'/worktree parameter, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Delete a file') and immediately narrows the scope with 'not a directory', which an agent needs in order to avoid misusing it. It also names the sibling tools (replace_text, apply) that serve the adjacent editing intent, so it is distinguishable from them without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions: only files (not directories), and it fails if the target does not exist. It routes editing intent to replace_text or apply. What it does not state is any prerequisite for deletion (e.g. permissions, whether the file must be closed/unmodified) or an alternative for removing directories.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.findARead-only
Locate declarations by name AND return their source in one call — the fused search-and-read that replaces grep -n 'func X' -A 30. Exact name matches win over substring ones. Use this instead of outline and read_range when you have not located the symbol yet. Pass queries to find several names in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Narrow by kind. func/function, type/struct/class/interface, method, const, var — spellings within a family are equivalent. Empty matches any. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| limit | No | Maximum declarations to return (default 5). Prefer budget. | |
| query | No | Symbol name, exact or partial. | |
| budget | No | Size of the answer in tokens. Cut at whole declarations; the rest is behind continue=<handle>. | |
| queries | No | Several symbol names in one call, instead of query. Each is answered as query would be. | |
| continue | No | Handle from a cut answer: its next page. | |
| maxLines | No | Maximum lines of each body (default 40). Prefer budget. | |
| dependency | No | Look in this dependency's source instead of the workspace, read-only: a crate, Go module, npm or Python package name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds genuinely useful behavioral context beyond that: exact name matches are prioritized over substring matches, and source bodies come back in the same call. It does not elaborate pagination/cutting, though the schema carries budget/continue.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the fused search-and-read purpose. Every sentence carries distinct information (purpose, ranking rule, routing guidance, batch option) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with fully documented parameters and no output schema, the description explains what is returned (declarations with source) and the ranking behavior. Only minor gaps around response shape/pagination remain, which the schema partly covers via budget and continue.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters including query, queries, budget, and continue. The description echoes a couple of these (queries for batch lookup, exact-over-substring matching) but adds little syntax or format detail. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (locate) and resource (declarations) plus the distinctive behavior of returning source in one call. It explicitly differentiates itself from grep, outline, and read_range, so an agent can select it without reading any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('when you have not located the symbol yet') and names the alternatives it supersedes (outline and read_range). The routing condition between this and sibling search/read tools is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.grepARead-only
Literal or regex text search across the workspace, returning path:line matches with optional trailing context — the replacement for grep -rn. Use this for anything that is not a declaration name: struct fields, string literals, error messages, config keys, or any search needing a path filter. Use find instead when you want a declaration and its body. Pass queries to search several patterns in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| glob | No | Restrict by path, e.g. *.go or internal/code/*. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| limit | No | Maximum matches returned (default 40). The true total is always reported. Prefer budget. | |
| query | No | Text to find. | |
| regex | No | Treat query as a regular expression. grep-style \| alternation and \( \) groups work as in grep. | |
| budget | No | Size of the answer in tokens. Cut at whole matches; the rest is behind continue=<handle>. | |
| context | No | Trailing lines to show per match, like grep -A (max 40). | |
| exclude | No | Skip paths containing this substring, e.g. testdata. | |
| queries | No | Several patterns in one call, instead of query. Each is answered as query would be, with the same filters. | |
| continue | No | Handle from a cut answer: its next page. Other arguments except budget are ignored. | |
| dependency | No | Search this dependency's source instead of the workspace, read-only, at the version the project locks: a crate, Go module, npm or Python package name. Matches read as dep:<name>/<path>, which read_range accepts. | |
| ignoreCase | No | Case-insensitive match. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already covers safety, but the description adds the return format (path:line matches with context), which matters since there is no output schema. It does not, however, expand on pagination/cut behavior, which the schema carries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, purpose front-loaded, with usage and alternative routing following. No wasted filler and the most decision-relevant information (what it is, what it isn't) comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter read tool with a fully-specified schema, one annotation, and no output schema, the description covers purpose, return shape, and sibling routing well. It leaves pagination/cut semantics to the schema, which is reasonable, so it is nearly but not entirely self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 12 parameters in detail. The description only lightly reinforces the queries parameter ('pass queries to search several patterns in one call') and adds no syntax beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (literal or regex text across the workspace), plus the returned shape (path:line matches with optional trailing context) and the grep -rn analogy. It explicitly distinguishes itself from the find sibling for declaration searches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use with concrete examples (struct fields, string literals, error messages, config keys), a when-not-to-use rule ('use find instead when you want a declaration and its body'), and a tip to batch patterns via queries. Both the condition and the alternative are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.insertA
Add text to a file without replacing anything — a new function, a new section, an extra case. Use this for additive work instead of rewriting a surrounding symbol. With no anchor it appends to the end of the file; with one it places the text before or after that anchor, refusing if the anchor is absent or matches more than once. Several additions or edits at once belong in apply.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File to add to. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| text | Yes | Text to insert. | |
| anchor | No | Optional. Exact, unique text to place the insertion beside. Omit to append to the end of the file. | |
| position | No | Optional. "before" or "after" the anchor. Defaults to after. | |
| expectedDigest | No | Optional. The digest from the read this edit is based on; the edit is refused if the file changed since, by anyone. | |
| expectedRevision | No | Optional. Revision expected before editing; the edit is rejected if the workspace has moved on. Omit for no precondition. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply destructiveHint=false, so the description carries most of the burden and does well: it discloses the append-by-default behavior, the anchor-placement rule, and the refusal conditions (anchor absent or matching more than once). It does not state whether the target file must already exist or how failures surface, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core behavior, then placement rules, then the routing note. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and minimal annotations, it covers when to use it, default placement, failure modes, and the multi-edit alternative. It stops short of stating whether the file must pre-exist or how the worktree/root resolution affects writes, which an agent would want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so 3 is the baseline, but the description adds real meaning beyond the schema by explaining that an omitted anchor appends to end-of-file and that a non-unique anchor causes refusal — semantics the anchor/position schema entries do not spell out.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add text to a file') with the key scope qualifier 'without replacing anything', then illustrates with concrete cases. It explicitly distinguishes itself from the surrounding-symbol rewrite use case covered by arno.replace_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('additive work instead of rewriting a surrounding symbol') and when-to-use-something-else ('Several additions or edits at once belong in apply'). The alternative sibling is named by name, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.read_rangeARead-only
Read a file verbatim, whole or by line range — the replacement for cat and sed -n. Omit both line numbers to read the whole file, which is how to read go.mod, a Makefile, or any JSON/YAML/TOML config that has no symbols to address. An end line past the end of the file reads to the end. A dependency's source reads as dep:/, read-only. Several ranges, in one file or many, go in one call: {"ranges": [{"path": "a.go", "lines": "280-400"}, {"path": "b.go", "lines": "700-760"}]}.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Repository-relative or workspace-relative file path. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| lines | No | Line range: "280-400", "280-" to the end, or "280". Omit to read the whole file. | |
| budget | No | Size of the read in tokens (default 5000). Cut at whole lines; the rest is behind continue=<handle>. | |
| ranges | No | Several reads in one call, instead of path. A range that fails reports its error without failing the others. | |
| continue | No | Handle from a cut read: the rest of it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes the safety profile, but the description adds real behavior beyond it: an end line past EOF reads to the end, dependency sources appear as dep:<name>/<path> read-only, and a failing range does not abort the others. It does not mention output format, which is acceptable given no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core verb/resource leads, followed by the whole-file rule, the EOF edge case, and finally the batching example. The inline JSON example is long but does genuine work by demonstrating the multi-range shape.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param read-only tool with no output schema, the description covers the operation, the two read modes, multi-range batching, error isolation, and the budget/continue mechanism. Nothing an agent needs to call it correctly appears missing, though the `root` semantics are left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still earns credit by explaining range edge-case semantics ('an end line past the end of the file reads to the end') and by showing a concrete multi-range payload that clarifies how `ranges` and `lines` interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a specific verb and resource with scope: 'Read a file verbatim, whole or by line range,' and frames it via the familiar 'replacement for cat and sed -n'. It never names a sibling (e.g. arno.grep or arno.find) to draw the boundary, so differentiation is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for the whole-file mode ('how to read go.mod, a Makefile, or any JSON/YAML/TOML config that has no symbols to address') and demonstrates batched ranges. It stops short of an explicit when-not/use-instead rule against the symbol-oriented siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.replace_textADestructive
Replace an exact, unique string in a file. An anchor string does not move when the lines around it do, which is why follow-up edits address text rather than line numbers. Refuses when the anchor is absent or matches more than once — extend it with surrounding context to disambiguate. For several sites, use apply: atomic, one validation, no diagnostics from half-done intermediate states.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path to edit. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| newText | Yes | Replacement text. | |
| oldText | Yes | Exact text to replace. Must appear exactly once. | |
| expectedDigest | No | Optional. The digest from the read this edit is based on; the edit is refused if the file changed since, by anyone. | |
| expectedRevision | No | Revision expected before editing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With destructiveHint=true already declared, the description adds genuinely non-obvious behavior: the tool refuses when the anchor is absent or matches more than once, and apply performs one atomic validation with no half-done intermediate states. It does not mention the concurrency guards, though those are documented on expectedDigest/expectedRevision in the schema. Solid added context, minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then rationale, then failure mode, then alternative. Every sentence contributes. The trailing "apply: atomic, one validation, no diagnostics from half-done intermediate states" is slightly compressed/ambiguous in phrasing but still compact for the information delivered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, the description covers the key risks: uniqueness of the anchor, refusal conditions, disambiguation strategy, and the multi-site alternative. Concurrency/expected-state handling is left to the schema, which is acceptable given 100% coverage, but the description never signals that staleness causes refusal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented, including the exact-once requirement for oldText and the digest/revision guards. The description reinforces the anchor concept but introduces no syntax, format, or constraint beyond what the schema states. Baseline 3 applies when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Replace an exact, unique string in a file") with the uniqueness constraint baked in. It explains the anchor model (text doesn't move when lines shift) and thereby distinguishes itself from line-number based siblings like insert. An agent can tell it apart from arno.insert and arno.apply without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names an alternative explicitly and the condition that selects it: "For several sites, use apply: atomic". It also gives a concrete recovery path when the tool refuses ("extend it with surrounding context to disambiguate"). Both the routed-away case and the disambiguation case are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.run_commandADestructive
Run a command the repository declares, by name, and wait for the verdict: pass/fail with the decisive output. Use it instead of a shell for anything check does not cover. No name lists the declared commands. Nothing fits? declare_command it once.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Declared command to run. Omit to list the declared commands. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| wait | No | Wait for the result (default true). False returns a job ID to poll. | |
| timeoutSeconds | No | Bound on the wait (default 90, max 300). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the safety profile is carried by structured data. The description adds the pass/fail verdict framing and the wait-for-result behavior, but never elaborates that executing a repo-declared command can mutate state or what auth/working-tree requirements apply, which is what the destructive hint warrants.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and compact, and each sentence carries a distinct instruction. The telegraphic fragments ('Nothing fits? declare_command it once') verge on cryptic, costing some clarity for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation-capable command runner with no output schema, the description conveys the result shape (pass/fail plus decisive output) and the listing fallback. The destructive/mutating nature implied by annotations is not spelled out, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented (name, root, wait, timeoutSeconds). The description only reinforces 'no name lists the declared commands', which the schema already states; it adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (run) and resource (a declared repository command) with the outcome (pass/fail verdict plus decisive output). It actively distinguishes itself from siblings by naming 'check' and 'declare_command', so an agent can route correctly without opening a sibling schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: use this instead of a shell for what 'check' does not cover, omit 'name' to list declared commands, and fall back to 'declare_command' when nothing fits. When-to-use, when-to-list, and the alternative path are all named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arno.run_testsA
Rerun a failing test, a test file, or changed files' tests; waits for pass/fail and the first failure. Full suite: check kind tests.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | Test file to run, for scope=file (Go: its package). With scope=test, limits the name filter to this file. | |
| root | No | Worktree for this call. Default: the one set with workspace, else the start tree. | |
| test | No | Test name, for scope=test: exact in Go, the runner's name filter elsewhere (jest/vitest -t, ava --match, pytest -k, cargo test <name>). | |
| wait | No | Wait for the result (default true). False returns a job ID to poll. | |
| scope | No | all (default), file, test or changed. | |
| timeoutSeconds | No | Bound on the wait (default 90, max 300). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=false, so the description must carry behavioral weight. It usefully discloses that the call blocks and returns pass/fail plus the first failure, but says nothing about auth needs, cost, or how partial failures surface. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses with zero filler: the action and scopes come first, the sibling routing last. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly mentions the returned pass/fail and first-failure signal, and the schema covers wait/job-ID polling and timeout bounds. It is nearly complete; only edge-case behavior on partial or errored runs is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains file, root, test, wait, scope, and timeoutSeconds in detail (including runner-specific name filters). The description only loosely maps its three rerun scopes onto the scope parameter, adding little beyond the structured fields; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (rerun) and three concrete scopes (a failing test, a test file, changed files' tests), so the agent knows exactly what it executes. It also distinguishes itself from the full-suite path by pointing to the 'check kind tests' sibling, though that routing phrase is compressed and slightly cryptic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear selection criteria for the rerun cases and an explicit alternative for the excluded case (full suite goes to arno.check). It stops short of stating prerequisites or when a rerun is inappropriate (e.g., no prior failure exists).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.0.17- Changed
arno.apply1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.check1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.create_file1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.declare_command1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.delete_file1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.find1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.grep1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.insert1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.read_range1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.replace_text1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.run_command1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
- Changed
arno.run_tests1 field changed- changed
Input schema / properties / root / descriptionPrevious value: -"Worktree of this repository to act in. Defaults to the session's workspace."New value: +"Worktree for this call. Default: the one set with workspace, else the start tree."
12 tool updates
v0.0.15- Changed
arno.apply1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.check1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.create_file1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.declare_command1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.delete_file1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.find1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.grep1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.insert1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.read_range1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.replace_text1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.run_command1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
- Changed
arno.run_tests1 field changed- added
Input schema / properties / rootAdded value: +{ + "description": "Worktree of this repository to act in. Defaults to the session's workspace.", + "type": "string" +}
24 tool updates
v0.0.12- Added
arno.apply - Added
arno.check - Added
arno.create_file - Added
arno.declare_command - Added
arno.delete_file - Added
arno.find - Added
arno.grep - Added
arno.insert - Added
arno.read_range - Added
arno.replace_text - Added
arno.run_command - Added
arno.run_tests - Removed
jade.apply - Removed
jade.check - Removed
jade.create_file - Removed
jade.declare_command - Removed
jade.delete_file - Removed
jade.find - Removed
jade.grep - Removed
jade.insert - Removed
jade.read_range - Removed
jade.replace_text - Removed
jade.run_command - Removed
jade.run_tests
1 tool update
v0.0.11- Changed
jade.run_tests1 field changed- changed
Input schema / properties / scope / descriptionPrevious value: -"One of: all, file, test, changed. Defaults to all."New value: +"all (default), file, test or changed."
12 tool updates
v0.0.10- First observed
jade.apply - First observed
jade.check - First observed
jade.create_file - First observed
jade.declare_command - First observed
jade.delete_file - First observed
jade.find - First observed
jade.grep - First observed
jade.insert - First observed
jade.read_range - First observed
jade.replace_text - First observed
jade.run_command - First observed
jade.run_tests
TDQS
Scored across 12 tools
The tools are mostly distinct: find vs grep vs read_range have clear boundaries, and insert vs replace_text vs apply are differentiated by single vs multiple edits. The command execution trio (run_command, check, run_tests) has slight overlap because multiple tools can run tests or validation, though the descriptions provide guidance on which to choose.
All tools use a consistent arno. prefix and snake_case, which is predictable. However, the action part mixes single verbs (find, insert, apply, check, grep) with verb_noun forms (read_range, run_command, etc.), a minor deviation from a pure verb_noun pattern.
With 12 tools, the set is well-scoped for a coding assistant: reading, searching, editing, file lifecycle, and command execution are each covered without excessive granularity. Each tool earns its place and the count sits comfortably in the typical 3-15 range.
The surface covers file CRUD (create, read, update via replace_text/apply/insert, delete), code search (find, grep), and command/test execution (run_command, check, run_tests, declare_command). Notable gaps include directory listing or filename search, and file move/rename/copy, which agents may need but can sometimes work around.
Maintenance
Related MCP Connectors
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Real-time chat for AI agents. Claude Code, Cursor, Cline and Codex join channels over MCP.
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
MCP Server for Slima - AI Writing IDE for Novel Authors with AI Beta Reader.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceRuns a language server and provides tools for communicating with it. Language servers excel at tasks that LLMs often struggle with, such as precisely understanding types, understanding relationships, and providing accurate symbol references.1,605BSD 3-Clause
- AlicenseAqualityAmaintenanceA fully featured coding agent that uses symbolic operations (enabled by language servers) and works well even in large code bases. Essentially a free to use alternative to Cursor and Windsurf Agents, Cline, Roo Code and others.2927,574 PyPI30,080MIT
- AlicenseAqualityAmaintenanceMCP server that keeps language server sessions warm and routes multiple languages through one process. Agents get persistent cross-file awareness, speculative execution (simulate edits before writing to disk), and 20 skills that encode correct multi-step operations like safe rename, blast-radius analysis, and end-to-end refactoring. Single Go binary, no runtime dependencies.50156MIT
- AlicenseNot gradedqualityBmaintenanceLocal-first code intelligence and safety layer for AI coding agents. MCP server exposes dependency graph, impact analysis, and AST-compressed repo context, backed by typed local memory, patch-scope safety gates, and git-independent transaction rollback.1MIT