Skip to main content
Glama

ARNO — Agent Repository Navigation & Operations

ARNO MCP server – quality and maintenance score on Glama

The IDE for agents

ARNO gives coding agents what an IDE gives you, served over MCP. Read by symbol, edit against a known revision, get the compiler's diagnostics back with the edit, validate with the repository's own commands, see what changed, and revert to a checkpoint — each step one tool call, none of it through the shell.

Symbol-aware reading is how ARNO finds its way around; the transaction is what it is for. Where the shell is still the better tool, use it — the question ARNO has to answer is whether an agent gets more done, at acceptable cost, with it than without (docs/benchmark.md).

Half the tokens, more issues fixed. Claude Code on 12 real closed issues from cobra (Go), ky (TypeScript), requests (Python) and ripgrep (Rust), judged by each upstream fix's own hidden tests:

Claude Code with…

Issues fixed

Tokens per task

Time per task

its built-in tools

10 of 12

1.38M

215s

ARNO in their place

12 of 12

0.68M (−51%)

163s (−24%)

Not just a shorter tool list: against a shell trimmed to Bash, Read, Edit and Write, ARNO still used 18–25% fewer tokens and fewer turns, in two separate runs. Sonnet 5, one run per task, four languages — results and caveats · how to set it up.

Status: 0.0.14 shipped, the first release without the Jade names; 0.0.11 and earlier were Jade — early, usable, and looking for feedback. Testing it? Start with the tester guide.

curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | sh

For macOS and Linux — Windows isn't supported (here's why, and where to upvote). Run it again to update. Other ways to install.


Related MCP server: Serena

Table of contents


Trying ARNO: a guide for testers

Thanks for testing. Half an hour gets you set up; the useful part is a week or two of your normal work with it switched on, and then telling us how it went — including if you turned it off.

1. Install ARNO and your language servers

macOS and Linux:

curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | sh

Windows isn't supported, sorry! If you'd like it to be, please 👍 issue #2 or tell us there why it matters to you. (WSL 2 runs Linux, so the Linux build may work there, but we don't test it or take bug reports for it.)

The script picks the build for your OS and CPU, checks it against the release's checksums and installs it to /usr/local/bin, or ~/.local/bin when that is not writable — no sudo, no Go toolchain. If that directory is not on PATH, it offers to add it to your shell profile, and it prints the absolute path to use as command in your MCP client config. On a first install it then opens arno-mcp install, a menu that installs the language servers you pick, or shows how to install them by hand; Enter skips it, and you can run it again any time. Run the same command again to update — ARNO tells you when a new release is out (see Update check). (Prefer Go? go install github.com/julianbei/arno/cmd/arno-mcp@latest works too.) ARNO reads structure in every language with nothing else installed; exact references, cross-file rename and type errors on edit need the language's server — arno-mcp install installs these for you, or by hand:

Language

Install

Notes

Go

go install golang.org/x/tools/gopls@latest

First answer about 1.6s.

Java

brew install jdtls (needs JDK 21+), or your distro's package

First answer about 8.5s while it indexes. check runs mvn or gradle; with only the Gradle wrapper, have the agent declare ./gradlew build once (step 4).

Scala

cs install metals (coursier)

Set ARNO_METALS_IMPORT=1 so metals may import the sbt build (it creates .bloop/ and .metals/). First answer about 22s.

Kotlin

—

No grammar or server yet: text search only. Tell us if you need it.

A missing server is never an error: ARNO says which answers are approximate.

2. Point your agent at a real repository

For Claude Code, put this in .mcp.json at the root of the repository you work in (other hosts: Codex CLI, goose, OpenCode):

{
  "mcpServers": {
    "arno": {
      "type": "stdio",
      "command": "arno-mcp",
      "args": ["--root", "/absolute/path/to/the/repo", "--tools", "core"],
      "alwaysLoad": true
    }
  }
}

That keeps Claude Code's own tools too. For the clearest signal, run some sessions with ARNO in place of them: claude --tools "" with the same config (why and trade-offs). Restart or reconnect the client (/mcp) after installing or upgrading ARNO.

3. Check it came up

Ask the agent: "call arno.capabilities". You should see your languages, each with server … (not started) or no server (… not installed), plus the build and test commands ARNO found (mvn, gradle, sbt, go test). If a server you installed shows as not installed, that is a bug report.

4. What to try

Work as you normally would. If you want a checklist for the first sessions:

  • Find and follow code: "where is X declared, and who calls it?" — find, references.

  • Rename across files: a method or class used in several files — rename (exact with gopls, jdtls or metals running).

  • A change in several places at once: "change the signature and update the callers, then check it builds" — apply with check.

  • Run the tests that matter: run_tests with a file or test name, or apply with check: "impact".

  • Repeatable commands: have the agent declare_command something you run often (./gradlew :core:test, sbt "testOnly *ParserSpec"); later sessions reuse it from .arno/commands.json.

  • Undo: checkpoint before something risky, revert if it goes wrong.

5. Tell us how it went

When

File this

After a week or two — or when you turn ARNO off

Feedback

The agent used the shell although an ARNO tool existed

Friction

A tool gave a wrong answer or failed

Bug

Something you wish ARNO did

Feature wish

The templates ask for the output of arno.capabilities and, optionally, arno.telemetry. Neither contains source code; telemetry is local only and records no arguments or response text, so both are safe to paste from a private repository.

Known rough edges on the JVM: no formatter runs for Java or Scala files; large Gradle builds can make jdtls's first answer much slower than 8.5s; Kotlin has no support yet.


Why ARNO exists

An agent that falls back to grep, sed and cat is operating outside any tooling you control. No revision tracking, no guardrails, no telemetry, no way to know what it did or why it chose to do it that way. Every shell fallback is a hole in your visibility.

You cannot fix that by telling the model not to use the shell. The model uses the shell because the shell is cheaper — fewer tokens, fewer round trips, more flexible. So the only durable fix is to make the structural tool the cheaper option, and then measure whether you succeeded.

That is the entire bet, and it is testable. On this repository's own benchmark, ARNO answers seven realistic engineering questions in 0.85x the tokens of the equivalent shell commands. It was 5.63x before responses became plain text instead of JSON — see docs/response-style.md for what changed and why.

The counter-measurement matters as much. One question asked against an unrelated repository came out at 1.36x — worse than the shell — because an ambiguous symbol name forced an extra disambiguation call. Seven scenarios at home and one away disagree, both are honest, and the second is the one that predicts outside use. The 0.0.4 pilot on four outside repositories answers it at a larger scale: used in place of Claude Code's built-in tools, ARNO solved 12 of 12 real issues with 51% fewer tokens, and 18–25% fewer than a shell trimmed to four tools (docs/benchmark-results.md). ARNO is not finished.

ARNO's own development log (docs/feedback.md) records every time its author reached for bash instead, and why. The pattern it found was blunt: the fallbacks that survived longest each closed within two tasks of being named in the log — not when the tool shipped.


Install

With the install script

curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | sh

Downloads the latest release binary for your OS and CPU, verifies it against checksums.txt, and installs it without sudo to /usr/local/bin or ~/.local/bin. Run it again to update: a arno-mcp already on PATH is replaced where it is, and nothing is downloaded when it is already current. If the directory is not on PATH, it offers to add it to your shell profile (ARNO_ADD_TO_PATH=1 does it without asking) and prints the absolute path to use in your MCP client config. ARNO_VERSION pins a release tag; ARNO_INSTALL_DIR picks the directory. Read it first if you like: install.sh.

Windows isn't supported — see issue #2, and give it a 👍 if you'd like that to change.

Language servers: arno-mcp install

arno-mcp install                           # menu: pick what to install
arno-mcp install --list                    # what is installed, and how the rest would be
arno-mcp install --servers go,java,scala   # install these, no questions
arno-mcp install --all                     # every missing server this machine can install

The menu lists each language server ARNO can use, whether it is installed, and the exact command it would run — go install for gopls, brew install jdtls, cs install metals, npm install -g for TypeScript and Pyright, rustup component add rust-analyzer, gem install ruby-lsp — and confirms before running anything. A server with no installer on the machine gets instructions for installing it by hand. The install script opens the menu after a first install; ARNO_SKIP_SETUP=1 skips it.

For an agent, or any script — there is no terminal to answer a menu, so the same steps come without questions:

# install ARNO and chosen servers in one go
curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | ARNO_SERVERS=go,java sh

arno-mcp install --list --json                   # state of every server, as JSON
arno-mcp install --servers java,scala --dry-run  # the commands, not run
arno-mcp install --servers java,scala            # run them

--list --json gives each server's key, whether it is installed and where, the command that would install it on this machine, and manual steps when there is none. arno.capabilities ends a missing server's line with the command that installs it. Installing is deliberately not an MCP tool: global package installs go through the agent's shell, where you approve them.

Other ways to install

The install script above is the easiest way. These work too.

From Go

go install github.com/julianbei/arno/cmd/arno-mcp@latest    # newest
go install github.com/julianbei/arno/cmd/arno-mcp@v0.0.14   # pinned to a tag

Lands in $GOBIN, or $(go env GOPATH)/bin if that is unset — which is usually ~/go/bin, and is not on PATH by default. Add it if it is not there, then confirm:

export PATH="$PATH:$(go env GOPATH)/bin"
arno-mcp --version

If you would rather not touch PATH, use the absolute path in your MCP client config instead of the bare arno-mcp shown below.

From a release binary

Prebuilt binaries for linux and darwin on amd64 and arm64 are attached to each GitHub release, with a checksums.txt alongside them. Each is built natively on its own platform — ARNO links tree-sitter through cgo, so the linux builds need a reasonably current glibc. On an older distro, build from source or use the container image, which is statically linked against musl.

VERSION=$(curl -fsSL https://api.github.com/repos/julianbei/arno/releases/latest | sed -n 's/.*"tag_name": *"\([^"]*\)".*/\1/p' | head -n 1)
OS=$(uname -s | tr '[:upper:]' '[:lower:]')
ARCH=$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')
curl -fsSL "https://github.com/julianbei/arno/releases/download/${VERSION}/arno-mcp_${VERSION}_${OS}_${ARCH}.tar.gz" \
  | tar xz
sudo mv "arno-mcp_${VERSION}_${OS}_${ARCH}" /usr/local/bin/arno-mcp

As an MCP bundle, for Claude Desktop

Since 0.0.11, each GitHub release also carries arno-mcp_<version>.mcpb, one bundle with the binaries for macOS and Linux on Intel and ARM. Open it with Claude Desktop, pick the repository ARNO should work on, and it runs with the core tools; no terminal and no Docker. Language servers still come from arno-mcp install, or run without them on tree-sitter alone.

From source

git clone https://github.com/julianbei/arno.git
cd arno
make binary          # bin/arno-mcp, version stamped from git describe
make install         # or straight onto your PATH

Optional: gopls

references and rename use gopls for their exact, compiler-resolved form. Without it they still work — references degrades to a textual approximation that says so in the response, and rename refuses rather than guessing.

arno-mcp install --servers go      # or: go install golang.org/x/tools/gopls@latest

The same goes for every language below: arno-mcp install shows which servers are installed and installs the rest (Language servers).


Configure your MCP client

ARNO is a stdio MCP server. Point your client at the binary:

{
  "mcpServers": {
    "arno": {
      "type": "stdio",
      "command": "arno-mcp",
      "alwaysLoad": true,
      "env": {
        "ARNO_WORKSPACE_ROOT": "/absolute/path/to/the/repo/arno/should/work/on"
      }
    }
  }
}

ARNO_WORKSPACE_ROOT is the repository ARNO inspects and edits. It does not have to be the ARNO checkout — pointing it somewhere else is the entire point. A --root /path/to/repo flag takes precedence over the environment variable, and ARNO prints which of the three sources it used (flag, env, working directory) at startup, so an agent can never quietly operate on the wrong repository.

ARNO works on a non-git directory and on a repository with no commits yet. In both cases it says what is degraded — changes, diff, history and checkpoint need git — and everything else keeps working.

Let ARNO replace the built-in tools

ARNO saves tokens when it replaces the agent's own tools, not when it is added next to them. Every turn resends the whole tool list, and in Claude Code the built-in tools are about 38k tokens of it. In ARNO's pilot benchmark (12 real issues in cobra, ky, requests and ripgrep, one run each):

Tools

Tasks solved

Tokens per run

Time per run

Claude Code's built-in tools

10 of 12

1.38M

215s

Built-in tools trimmed to Bash, Read, Edit, Write

10 of 12

0.90M

213s

ARNO only, core profile

12 of 12

0.68M

163s

Both, all built-in tools and ARNO

11 of 12

1.50M

191s

Trimming the built-in list is most of the saving on its own. ARNO on top of that used 25% fewer tokens and 24% less time than the trimmed shell, and solved the two tasks both shell setups failed. Given both ARNO and every built-in tool, the agent used Bash for four calls in five and paid for both lists. To run ARNO in place of the built-in tools:

claude --tools "" --mcp-config arno.json

with arno.json passing the core profile, which lists twelve tools — read, edit, validate, and run the commands a repository declares — and keeps the others callable:

{
  "mcpServers": {
    "arno": {
      "type": "stdio",
      "command": "arno-mcp",
      "args": ["--root", "/absolute/path/to/the/repo", "--tools", "core"],
      "alwaysLoad": true
    }
  }
}

The trade-off is real: without Bash the agent cannot run arbitrary commands. run_tests, check and repository commands declared with declare_command cover building, testing and repeatable scripts — a declaration lives in .arno/commands.json, so later sessions reuse it; a task that needs git operations, network access or ad-hoc scripts needs the shell back. The numbers above are one run per task — see docs/benchmark-results.md for the results, a rerun after the pilot's fixes, and the caveats.

Other hosts

Verified with a live session — find a declaration, insert beside it, run check — using only ARNO's tools:

Codex CLI (0.154), in ~/.codex/config.toml:

[mcp_servers.arno]
command = "arno-mcp"
args = ["--root", "/absolute/path/to/the/repo", "--tools", "core"]

Codex asks before every MCP tool call. With approval_policy = "never" it refuses them outright; run codex exec --approve-for-me or approve ARNO's tools interactively.

goose (1.50), for one run:

goose run --with-extension "arno:arno-mcp --root /absolute/path/to/the/repo --tools core" -t "…"

or permanently with goose configure → Add Extension → Command-line Extension, command arno-mcp --root /absolute/path/to/the/repo --tools core. Verified through goose's claude-code provider.

OpenCode (1.18), in opencode.json at the repository root or in ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "arno": {
      "type": "local",
      "command": ["arno-mcp", "--root", "/absolute/path/to/the/repo", "--tools", "core"],
      "enabled": true
    }
  }
}

OpenCode prefixes tools with the server name, so they appear as arno_arno_find and so on. Verified with the github-copilot provider (Claude Sonnet 5).

Cline and Gemini CLI are not verified yet.

Three things that will confuse you once

Without "alwaysLoad": true, Claude Code may never use ARNO. Claude Code hides MCP tools behind a tool search by default: the agent sees their names but not their definitions, and has to search before it can call one. With its own shell and file tools right there, it does not. In ARNO's benchmark, an agent given both ARNO and the shell made no ARNO call in three of three runs; the same setup with alwaysLoad called ARNO directly. Other hosts may have their own equivalent — check that ARNO's tools are actually being called.

One server serves one repository, but any of its worktrees. ARNO is pinned to the directory it was started in, and no cd in your shell reaches it — MCP carries no per-call working directory. To act in another git worktree of the same repository, pass root=<path> on the tool; a session and its subagents can each name their own, so they do not disturb each other. arno.workspace lists the worktrees. See docs/worktrees.md.

The MCP tool catalog is fixed at connection time. A newly added tool does not appear until the client reconnects. If you upgrade ARNO mid-session and a tool seems missing, reconnect before investigating.

Use the binary, not go run ./cmd/arno-mcp. A go run stanza recompiles at every process start: measured here at 284–584ms to first handshake against 14ms for the binary, with a warm build cache. A cold one is seconds. It also means the server silently changes whenever the source does — useful while hacking on ARNO itself, confusing everywhere else. This repository's own .mcp.json deliberately still uses go run for that reason.

Environment variables

Variable

Effect

ARNO_WORKSPACE_ROOT

Repository to operate on. Overridden by --root.

ARNO_JSON=1

Emit machine-readable JSON instead of plain text.

ARNO_TELEMETRY=0

Disable local usage recording entirely.

ARNO_STATE_DIR

Keep the telemetry log outside the workspace, one subdirectory per workspace.

ARNO_METALS_IMPORT=1

Let metals import an sbt build so Scala edits get diagnostics. Runs sbt; creates .bloop/ and .metals/.

ARNO_UPDATE_CHECK=0

Turn off the daily check for a newer release (Update check).

ARNO_MAX_WORKSPACES

How many git worktrees one server keeps open at once (default 3, least recently used evicted first). See docs/worktrees.md.

The install script reads its own:

Variable

Effect

ARNO_VERSION

Release to install, e.g. v0.0.12. Default: the latest.

ARNO_INSTALL_DIR

Where to put arno-mcp. Default: the directory of the arno-mcp already on PATH, else /usr/local/bin if writable, else ~/.local/bin.

ARNO_SERVERS

Language servers to install afterwards without a menu: go,java,scala,typescript,python,rust,ruby, or all.

ARNO_SKIP_SETUP=1

Skip the language-server step.

ARNO_ADD_TO_PATH=1

Add the install directory to the shell profile without asking.

ARNO_RELEASE_URL

Base URL of the releases, for a mirror.


Use it in a container

ARNO is a child process, not a service, so the useful shape is to copy the binary into your own image rather than run ARNO's:

FROM ghcr.io/julianbei/arno-mcp:latest AS arno

FROM your-project-base
COPY --from=arno /arno-mcp /usr/local/bin/arno-mcp
ENV ARNO_WORKSPACE_ROOT=/workspace

Or build it yourself from the included Dockerfile.

The published image is distroless/static, so it carries no git, no gopls and no language toolchains. ARNO detects each of those at runtime and degrades with an explicit message rather than failing, so this still works — you get the textual references fallback, and changes/diff/history/checkpoint are off. If you want the full surface, install git and gopls in your image; ARNO will find them.


The tools

31 tools, in four groups. Every response is plain text, shaped to lead with the decisive line — the answer first, the supporting detail after, raw output only when you ask for it.

Over MCP each one is registered as arno.<name> — arno.outline, arno.replace_symbol, and so on. The tables below use the bare name for readability. Most MCP clients show you the prefixed name already, often with the dot rewritten (Claude Code displays mcp__arno__arno_find). Over the wire ARNO accepts both arno.find and arno_find.

Inspect

Tool

What it does

capabilities

What ARNO can do in this workspace: per language, grammar or text scan, language server state, formatter; git, validation commands, declared commands. Call it first.

outline

File structure — declarations grouped by kind, without reading bodies.

read_range

Verbatim lines, or a whole file. lines: "280-400" picks a range; ranges reads several files or ranges in one call; an end line past the file reads to the end. dep:<name>/<path> reads a dependency's source, read-only, at the locked version — grep and find take dependency to search it.

find

Locate a declaration and get its body in one call. queries finds several names at once.

grep

Literal or regex text search with path globs. The replacement for grep -rn.

references

Find usages. Exact from the language server when one is installed; a name-matched approximation otherwise, and it says which answered.

retrieve

Pull a working set for a query.

context

Assemble the surrounding context for one symbol.

workspace_tree

Directory structure.

Symbols are addressed as path::Name, or path::Name@line when a name is ambiguous. An ambiguous read returns the candidates with their signatures rather than guessing.

Modify

Tool

What it does

replace_symbol

Replace a whole declaration. Takes the full path::Name@line ID, or just path::Name when that name is unique in the file.

replace_text

Replace exact, unique text. Anchored on content, not line numbers. Like every text edit, returns the edited region as it now reads.

replace_file

Replace an entire file's contents.

create_file

Create a new file.

delete_file

Delete one file; directories are refused.

delete_symbol

Delete one declaration.

rename

Cross-file rename from the language server; refuses rather than guessing when it cannot be exact.

insert

Add text without replacing anything — a new function, a new section, an extra case. Appends with no anchor; places before or after a unique anchor with one.

apply

Several edits as one atomic unit — anchors validated up front, all applied or none, one revision bump and one validation at the end.

Every edit returns consequences, not "success": the revision transition, which symbols moved, immediate diagnostics, and the IDs of any background validation it started. Edits accept an expectedRevision precondition; supplying it makes a stale edit fail loudly instead of silently clobbering a concurrent change.

Validate

Tool

What it does

check

Build, typecheck or tests — discovering the repository's own command rather than assuming one: Makefile target, then npm script, cargo, Maven, Gradle, sbt, pytest/mypy or bundler, by manifest. A project it cannot identify is reported as such rather than run with the wrong toolchain. Every result names the command that ran; dryRun names it without running.

run_tests

Tests scoped to a file, a test name, or the changed files.

run_command

Run one of the repository's declared commands by name.

declare_command

Add or remove a declared command.

job_status

Poll a background job.

job_output

Raw output for a job, on demand.

Exit status is authoritative. A command that prints a success-looking line and exits non-zero fails.

State

Tool

What it does

changes

What moved — by file and by symbol, not just by path.

diff

The patch, including untracked files. since takes any git revision.

history

Which commits touched one symbol, via git log -L.

checkpoint

Mark a revertible point: snapshots the files ARNO edited and records git's HEAD. Not a commit.

revert

Restore those files to a checkpoint. Never moves git, and refuses if a commit landed since the checkpoint.

events

The workspace event stream.

telemetry

How ARNO's own tools have been used in this workspace.

workspace

List this repository's git worktrees, and change which one the session serves by default.

Every tool above except workspace, events, telemetry, job_status and job_output also takes an optional root=<path>: one of this repository's git worktrees, to act in for that one call. See docs/worktrees.md.


Project configuration

Discovery guesses how to build and test a repository from its manifests, and the guess is sometimes wrong: a Makefile's python -m pytest picks the system interpreter, npm run test runs lint and browser suites for a one-file check, and a Go module with a TypeScript app beside it has two answers to "build". A committed .arno/project.json states the answer once, the way an editor keeps its settings in .vscode/:

{
  "areas": [
    { "path": ".", "language": "go",
      "build": "go build ./...", "test": "go test ./...",
      "testName": "go test -run {name} ./..." },
    { "path": "web", "language": "typescript",
      "typecheck": "node_modules/.bin/tsc --noEmit",
      "test": "node_modules/.bin/vitest run",
      "testFile": "node_modules/.bin/vitest run {file}",
      "testName": "node_modules/.bin/vitest run {file} -t {name}" }
  ],
  "env": { "python": ".venv/bin/python", "vars": { "CI": "1" } },
  "generated": ["web/dist", "*.pb.go"],
  "notes": "Browser tests need Playwright; run unit tests by file."
}
  • Areas are parts of the repository with their own language and commands, each run inside its path. check runs a kind in every area that declares it; run_tests with a file uses the deepest area containing that file, with {file} relative to the area and {name} the test name.

  • Empty fields fall back to discovery, so a config can state only what discovery gets wrong. A config that does not parse, names an unknown field or a path outside the workspace fails the check instead of being ignored.

  • env puts the interpreter's directory first on PATH and adds the variables to every command. generated paths are skipped by grep, find and the workspace tree. notes, with the list of areas, is sent to the agent when a session starts.

  • check with dryRun says when a command comes from .arno/project.json.

arno-mcp init --root /path/to/repo drafts the file from discovery, one area, for you to review and commit; it never overwrites an existing one.

Repository commands

Beyond build and test, every repository has its own verbs — lint, codegen, migrate, release-gate — and an agent that does not know them reaches for the shell. So ARNO lets it record them instead:

declare_command(name: "lint", run: "golangci-lint run ./...")
run_command(name: "lint")

They live in .arno/commands.json, which is meant to be committed. It becomes the repository's declared command vocabulary — written once by whoever (or whatever) worked out the incantation, replayed by name forever after. Calling run_command with no name lists what the repository declares; calling it with an unknown name answers with the commands that do exist, so a wrong guess teaches rather than fails.

Both answers also list detected candidates: npm scripts and Makefile targets ARNO can see in the repository's own manifests but nobody has declared. Detection is read-only — nothing is written until declare_command says so — and a script check already runs for build, typecheck or tests is left off the list, since it is not a gap. capabilities shows the same candidates alongside what it already reports. This closed a failure class the 2026-09-16 benchmark measured directly: an agent guessing run_command names (test-unit, model_formsets_tests) against a repository that had declared nothing, each guess a wasted call, when the manifest already answered it.

A validation chain

Repository rules — Semgrep, a custom linter, a licence check — belong in validation, and they need no integration in ARNO. Declare one command that runs the steps in order, joined with &&:

{
  "validate": {
    "run": "go test ./... && semgrep scan --config .semgrep.yml --error",
    "description": "tests, then repository rules"
  }
}

or, without editing the file, declare_command(name: "validate", run: "…").

run_command(name: "validate") runs it inside ARNO, so the run is part of the session's record. Exit status decides: a rule that fails fails the run, the steps after it do not run, and the summary leads with the failing output. Use semgrep scan --error or the equivalent flag of your tool — a tool that prints findings and exits 0 passes.


Language support

Structure comes from tree-sitter grammars compiled into the binary, so it works with nothing installed. Semantics come from a real language server, which you provide — arno-mcp install installs it for you — and ARNO starts it on first use, reuses it for the session, and shuts it down on exit.

Language

Structure

Semantics, with this installed

Go

✅ built in

gopls

TypeScript / TSX

✅ built in

typescript-language-server

JavaScript

✅ built in

typescript-language-server

Rust

✅ built in

rust-analyzer

Python

✅ built in

pyright-langserver, or pylsp / jedi-language-server

Ruby

✅ built in

ruby-lsp, or solargraph

Java

✅ built in

jdtls

Scala

✅ built in

metals

Everything else

text scan, announced

—

Every row is verified end-to-end by make conformance, which builds an image containing all eight servers and runs ARNO against a real repository per language.

Semantic requests wait for the server to finish indexing (its $/progress tokens), because an indexing server answers wrongly rather than slowly. When the primary server declines a rename, ARNO asks the language's installed alternative: ruby-lsp renames classes but not methods, so Ruby method rename needs solargraph installed alongside it. With ruby-lsp alone, method rename refuses and repeats the server's reason.

Every edit response names what checked the file (checked: pyright-langserver) or why nothing did (not checked: app.py: pyright-langserver is not installed).

Scala needs one opt-in. metals reports errors only after importing the sbt build, and it asks permission first, because importing runs sbt and creates .bloop/ and .metals/ in the repository. ARNO declines unless ARNO_METALS_IMPORT=1 is set, and says so in the edit response. A repository an editor has already imported needs no setting.

"Structure" is outline, symbol read, edit-by-symbol, grep and search. "Semantics" is exact references, cross-file rename, and type-level diagnostics on edit.

ARNO looks for servers on PATH and in the places toolchains actually install them — ~/go/bin, ~/.cargo/bin, ~/.local/bin, ~/.coursier/bin — because go install puts gopls somewhere that is not on PATH by default, and a client that only checked PATH would report Go as unsupported on a machine that has a working gopls.

A missing server is never an error. ARNO degrades to the behaviour above and says which answer you got.

Design principles

  1. Structure before source. Return the minimum sufficient representation first — outline before full source, summary before raw logs.

  2. Deterministic tools before model reasoning. ARNO orchestrates tree-sitter, git, gopls and the project's own build tooling. It does not reimplement them, and does not guess where they could answer.

  3. Every edit has a precondition and returns consequences. An edit can name the revision it expects and is refused if ARNO's revision has moved; it returns the revision transition, what changed and the diagnostics — not "success". Changes made outside ARNO do not yet move the revision (release plan Phase 5).

  4. Conclusions before logs. The verdict leads. Raw output expands on request.

  5. Semantic operations before textual ones. But textual escape hatches stay available, because the semantic path does not always exist.

  6. State is explicit. Revisions, checkpoints and change sets are objects, not implications.

  7. Validation waits by default, backgrounds on request. check, run_tests and run_command return the verdict; a long run can return a job to poll instead.

  8. An approximation must announce itself. When ARNO falls back to a text scan or a name-matched graph, the caveat travels with the data, in the response — not in documentation the agent will never read.

  9. ARNO is model- and harness-independent. MCP is an adapter, not the architecture.

  10. Measure agent outcomes, not infrastructure sophistication. Tokens and turns per completed task — and token reduction is worthless if the success rate drops with it.

  11. Repository-native execution. Builds, tests and lint run through the repository's own commands — discovered, or declared in .arno/commands.json — inside ARNO, so validation is part of the record instead of a shell side trip.

  12. Cheaper than the escape hatch. If the shell is easier, faster and cheaper for a workflow, ARNO has failed that workflow. The benchmark, not opinion, says which (docs/benchmark.md).

The longer design document is docs/scope.md.


When the shell is still the right tool

ARNO does not try to match the shell's composability. Using the shell is a decision, not a leak, when the work is one of these:

  • Git operations: commit, branch, rebase, push, blame. ARNO reads git state (changes, diff, history) and never moves it.

  • One-off probes: curl a local server, inspect a process, check a port, read an environment variable.

  • Debugging a script or a build system itself, where the question is what a shell command does rather than what the code says.

  • Installing dependencies and toolchains: npm install, go install, pip install.

  • Network access of any kind.

What stays on ARNO's side of the line, even though a shell could do it:

  • Builds, typechecks, tests, lint and codegen. Run them with check, run_tests or a declared command (declare_command, then run_command). Validation run from the shell is validation the change transaction cannot see: no verdict in the edit record, no scoped test runner, no failure summary.

  • Reading and searching code, and editing it. That is where ARNO's revisions, diagnostics and provenance apply.

A command you keep running from the shell for validation belongs in .arno/commands.json, or in .arno/project.json as an area's build or test command.


What ARNO does not do yet

Windows. ARNO is built and tested for macOS and Linux only, and we'd rather do those two really well than three halfway. Until further notice we don't build, test or look at Windows. If you'd like ARNO on Windows, please 👍 issue #2 — and if you think this is the wrong call, say so there; honest feedback is welcome. WSL 2 runs Linux, so the Linux build may work there, but it isn't tested.

This list is more useful than the feature list — it tells you what is worth reporting and what is already known. What is planned is in ROADMAP.md.

  • Nine languages get a real grammar; the rest fall back to a text scan. Go, TypeScript, TSX, JavaScript, Python, Ruby, Java, Scala and Rust are parsed properly. Anything else (Kotlin, Swift, C/C++, C#, PHP, …) is served by a heuristic that finds some declarations and misses others — and the amount it misses varies enormously by language, so treat those outlines as a hint rather than an inventory. ARNO always says which you got (! no kotlin grammar — …).

  • Semantic features need a language server installed for that language. ARNO speaks LSP to whatever is on the machine (see Language support). With a server, references and rename are compiler-exact and cross-file. Without one, references degrades to a textual approximation that says so, and rename refuses rather than guessing — an approximate reference list is still useful to a reader, but an approximate edit is corruption.

  • No completion, hover or code actions. ARNO's LSP client implements what the tools need — references, rename, diagnostics — not the whole protocol.

  • No blame, no cross-repo work, no remote execution.

  • Formatting runs only where it is safe. gofmt and rustfmt always run. prettier (TypeScript/JavaScript), black or ruff (Python) and scalafmt run only when the repository declares them — its config file, and for Node and Python the project's own binary — because a formatter the project did not choose turns a one-line edit into a whole-file diff. Ruby, Java, JSON and Markdown are left as edited.

  • Revision tracking is ARNO's own counter, not git's. It detects concurrent edits within a session. It is not a VCS. A checkpoint snapshots the files ARNO has edited and records git's HEAD; revert restores those files and nothing else, never moves git, and refuses once a commit has landed since the checkpoint — undoing committed work is git's job. Checkpoints do not survive a restart of the server.

  • Not hardened for untrusted input. It runs shell commands you declare and edits files you point it at. Treat it as a development tool, and do not point it at a repository you would not run make in. What running ARNO inside a sandbox or container does and does not cover:

    • Covered by ARNO itself: reads and writes stay inside the workspace root, symlinks included; dependency sources are read-only; a repository cannot make ARNO launch a binary it ships (declared commands run through the shell you already trust, and a .arno/project.json interpreter is a path you review in the diff).

    • Covered only by the sandbox: what a declared command, a Makefile target, an npm script or a test suite does when check, run_tests or run_command runs it — network access, files outside the workspace, credentials in the environment. ARNO runs the repository's own commands with your environment; a malicious repository's make test is as dangerous under ARNO as in your shell.

    • Not covered at all: an agent asked to declare a harmful command, and language servers, which execute project configuration of their own (build scripts, plugins) when they index a workspace.


Stability and versioning

Tool names and required arguments are frozen and enforced by a test. 0.0.3 added no tools and made two arguments optional (query on find, path on read_range), both backward compatible. Schemas and server instructions are still read once at connection time, so reconnect after upgrading. docs/tool-contract.md has the full surface and the policy on what counts as a breaking change.

What is not frozen: response wording, the .arno/* file formats, the exact spelling of symbol IDs, and everything under internal/. Treat responses as text for a model to read, not as a format to parse. ARNO_JSON=1 gives machine-readable output if you need to parse something.


Telemetry

ARNO records how its own tools are used — call counts, response sizes, timing, and the failure classes that most often precede a caller giving up and using the shell.

It is written to .arno/telemetry.jsonl in your workspace and never transmitted anywhere. It records no arguments, no response bodies and no error text — only a 10-character hash of each call's target (path, symbol or query), so the confusion report can tell a second tool asked about the same thing. ARNO_TELEMETRY=0 turns it off; telemetry(reset: true) clears it.

ARNO tries not to leave files in a repository it was only asked to work in:

  • In a git repository, before creating the log, ARNO adds it to .git/info/exclude — the clone-local ignore file, never committed — unless git already ignores it. .gitignore is never touched. The log does not show up as untracked, so a harness that commits every untracked file does not commit it.

  • ARNO_STATE_DIR=/some/dir moves the log out of the workspace entirely, into a subdirectory per workspace. Use it when ARNO is rooted at a checkout that something else commits or reviews wholesale.

  • A call ARNO rejects outright (an unknown tool name) never creates the log.

.arno/commands.json is different: it is the repository's declared command vocabulary, meant to be committed, and is only created when you declare a command.

Update check

Separate from telemetry, ARNO looks up the newest release tag on GitHub — one unauthenticated request for releases/latest, carrying nothing about your workspace or how you use ARNO — at most once a day, in the background, with a three-second timeout. A failed or offline check also waits a day. When a newer release exists, it says so only where you asked what you are running:

$ arno-mcp --version
arno-mcp v0.0.12
update available: v0.0.13 (running v0.0.12) · curl -fsSL https://raw.githubusercontent.com/julianbei/arno/main/install.sh | sh, then reconnect your MCP client

and as the second line of arno.capabilities. It never appears in the server instructions or in other tool responses. ARNO_UPDATE_CHECK=0 turns it off; it is also off in CI (CI set) and for development builds.

It exists because response cost is invisible to whoever is reading the response. Its first live reading found a tool returning 4.6KB in 704ms on a routine call — something sixteen tasks of hand-written notes had never noticed.


Reporting problems

docs/reporting.md says what makes a useful report. There are four issue templates:

  • feedback — how it went after some real use, or why you turned it off.

  • bug — it did the wrong thing.

  • friction — "I used the shell instead." This is the valuable one.

  • feature wish — it should be able to do X.

If you are unsure which, pick friction. It is the cheapest to write and the easiest to act on, and "it was just habit" is a real answer — we want it. Every shell fallback is a place ARNO was not worth reaching for, and that is the only signal that reliably improves it.

Before filing a bug, check whether your client has reconnected since the version changed. A stale tool catalog explains a surprising share of "this tool does not exist" and "my fix did not take effect".


Development

make build      # go build ./...
make test       # go test ./...
make fmt        # gofmt -w ./cmd ./internal
make binary     # bin/arno-mcp, version-stamped
make install    # onto your PATH

The repository declares its own commands in .arno/commands.json, including release-gate — build, vet, tests and a gofmt check, which is the gate a tag has to pass. Run it the way an agent would: run_command(name: "release-gate").

Layout:

Path

What lives there

cmd/arno-mcp

The MCP stdio server — the entry point that matters.

cmd/arno

A small CLI for driving the internal API directly.

cmd/arno-bench

The token benchmark: ARNO against equivalent shell commands.

internal/workspace

Revisions, change sets, checkpoints, git.

internal/code

Symbol index, outlines, search, grep, references.

internal/edit

Mutation, atomic apply, formatting.

internal/diagnostics

Immediate feedback on edits; gopls.

internal/jobs

Async job runner, command discovery.

internal/languages

Per-language adapters (Go, TypeScript, Rust).

internal/commands

The declared-command registry.

internal/telemetry

Local usage recording.

internal/transport

MCP adapter, and the transport-independent internal API.

internal/protocol

Shared request and response types.

Contributions are welcome. The one hard rule is principle 8: if you add a code path that approximates, the response has to say so.


License

Apache License 2.0 — see LICENSE. Copyright 2026 Julian Amelung.

Available Tools

12 tools
arno.applyA
Destructive

Apply several edits as one atomic unit: all land or none do. Ops: replace_text, replace_range, replace_symbol, delete_symbol, insert. Anchors are validated before anything is written, touched files are formatted, and one validation runs at the end instead of one per edit. Prefer this over several single edits when changing more than one site.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
checkNoRun one validation after all edits: build, typecheck, tests (the edited files' tests), or impact (those plus tests of callers of touched declarations).
editsYesEdits to apply in order.
formatNoFormat touched files afterwards (default true).
expectedRevisionNoRevision expected before editing.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only destructiveHint=true in annotations, the description carries real weight: atomicity ('all land or none do'), pre-write anchor validation, automatic formatting of touched files, and one validation at the end rather than per edit. These are concrete behavioral traits an agent cannot infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the atomic guarantee, then ops, then mechanics, then the routing advice. No filler or restatement of the title/name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-op mutation tool with no output schema, the description covers atomicity, validation timing, formatting, and failure behavior, which is everything an agent needs to decide and call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by clarifying that 'check' runs a single validation after all edits (schema only lists the validation kinds) and that anchors are validated up front. Op names and edit ordering are reinforced but largely duplicate the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (apply several edits as one atomic unit) and enumerates the supported operations (replace_text, replace_range, replace_symbol, delete_symbol, insert). This clearly separates it from single-op siblings like arno.replace_text and arno.insert.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Prefer this over several single edits when changing more than one site,' naming the alternative and the selecting condition. The check options (build, typecheck, tests, impact) further guide correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.checkA

Run a validation command on demand and wait for the verdict: kind build (default), typecheck or tests. Uses the repository's own Makefile target, npm script or cargo command when present. Waits by default and returns pass/fail directly. Every result names the command that ran; dryRun names it without running anything. In a repository with several projects, pass target to check one; with no command at the root, the answer lists the projects.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNobuild (default), typecheck, tests, lint or codegen. lint and codegen run the declared commands of that kind; lint with none declared runs typecheck.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
waitNoWait for the result (default true). False returns a job ID to poll.
dryRunNoName the command that would run, without running it.
targetNoProject directory inside the workspace to check, e.g. services/api. Omit for the workspace root.
timeoutSecondsNoBound on the wait (default 90, max 300).

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=false, so the description carries most of the burden and does so well: it discloses synchronous-by-default behavior, direct pass/fail return, that every result names the executed command, and that dryRun names without running. It omits side effects of running a build (artifact writes) and how to poll the job ID returned when wait=false, but the core behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the wait semantics, then defaults and edge cases; every sentence contributes. It is slightly dense with several clauses per sentence but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, six optional parameters, and only a destructiveHint annotation, the description supplies the essentials: default kind, wait/dryRun behavior, timeout context, and the multi-project fallback. It stops short of describing error surfaces or the polling path for wait=false, which are the remaining gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description meaningfully supplements it: it explains the default kind, the multi-project meaning of target, and that dryRun names the command without executing. It adds interpretation rather than restating field types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Run a validation command on demand and wait for the verdict') and enumerates the kinds (build/typecheck/tests), so the agent knows exactly what is executed. It is clear without opening the schema, though it never distinguishes itself from the sibling arno.run_tests, which overlaps on the 'tests' kind.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the defaults and the multi-project note ('pass target to check one; with no command at the root, the answer lists the projects'), which gives real context. However, there is no explicit when-to-use versus siblings like arno.run_command, arno.run_tests or arno.declare_command, leaving the agent to infer the boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.create_fileA

Create a brand-new file, and any missing parent directories. Refuses to overwrite an existing one — use replace_text or apply to modify existing content.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path to create.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
contentYesFile content.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide destructiveHint=false, so the description must carry the behavior burden and does: it discloses implicit parent-directory creation and the fail-on-existing-file semantics, which are non-obvious. It doesn't cover permissions, error surface, or return value, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the core create-plus-parent-dirs behavior is front-loaded and the exclusion/alternative follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple creation tool with full schema coverage and no output schema, the description covers the key behavioral surprises (mkdir -p, no overwrite) so an agent can call it correctly. Minor gaps remain around permissions and failure responses, keeping it just under complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with three self-documented parameters (path, root, content), so the baseline is 3. The description adds no format, path-resolution, or root-precedence detail beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a brand-new file') and adds scope ('any missing parent directories'). It also names the siblings appropriate for the adjacent case (replace_text, apply), letting an agent distinguish creation from modification without reading schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the when-not condition ('Refuses to overwrite an existing one') and routes to the named alternatives ('use replace_text or apply to modify existing content'). No inference required to pick the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.declare_commandA
Destructive

Declare a named command in .arno/commands.json — a reproduction, a benchmark — to run with run_command in this and later sessions; the file is reviewed like any change. Declaring an existing name replaces it; do not redeclare one just to run it.

ParametersJSON Schema
NameRequiredDescriptionDefault
runNoShell command to run from the workspace root.
kindNoOptional: lint or codegen; check with that kind runs it.
nameYesCommand name: lowercase letters, digits, ':', '_' or '-'.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
removeNoDelete the named command instead of declaring it.
descriptionNoOptional note on what the command is for.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply destructiveHint=true; the description adds real behavioral context beyond that: persistence across sessions, that the manifest file is 'reviewed like any change', and the replace-on-duplicate semantics for an existing name. It stays silent on the deletion path (the remove param), which the annotation's destructive hint implies but the prose never explains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with zero filler; the core function and its target file are front-loaded before the caution about redeclaring, which is exactly the ordering an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return/behavior burden and does reasonably: persistence, review gating, and replacement are covered, and the schema fully documents parameters. The one gap is the removal/destructive path, which the destructiveHint annotation hints at but the prose never spells out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (run, kind, name, root, remove, description) is already documented in the schema; baseline is 3. The description adds only the replace-on-existing-name behavior for `name`, which is modest value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Declare a named command in .arno/commands.json') and explicitly scopes the purpose ('to run with run_command in this and later sessions'). It also distinguishes the tool from the sibling run_command, telling the agent this is for declaration, not execution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (declare so it can be run in this and later sessions), when-not ('do not redeclare one just to run it'), and names the alternative tool (run_command) for execution. Nothing about tool selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.delete_fileA
Destructive

Delete a file, not a directory; fails if it does not exist. Reverting to an earlier checkpoint recreates it. To edit content use replace_text or apply.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path to delete.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry destructiveHint=true, so the safety profile is known; the description goes beyond that by disclosing the error behavior ('fails if it does not exist') and, valuably, that the deletion is recoverable via checkpoint revert. That reversibility detail is exactly the kind of trait annotations cannot express. It does not mention permissions or confirmation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, front-loaded with the core action and its key constraint, and every clause carries distinct information (scope, failure mode, reversibility, alternative). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema, it covers failure behavior, reversibility, and the alternative editing path, which is nearly everything an agent needs. Only the return/confirmation shape and any permission requirements are unstated, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'path' and 'root' documented in the schema itself, so the baseline is 3. The description adds no syntax, format, or defaulting nuance beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (delete) and resource (file), and immediately scopes it negatively with 'not a directory', which rules out the most likely confusion. It also names the sibling tools (replace_text, apply) that handle the adjacent editing case, so an agent can distinguish it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit exclusion and an explicit alternative: to edit content, use replace_text or apply. What's missing is any positive context on when deletion is appropriate or prerequisites, but for a self-evident destructive verb this is close to complete routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.findA
Read-only

Locate declarations by name AND return their source in one call — the fused search-and-read that replaces grep -n 'func X' -A 30. Exact name matches win over substring ones. Use this instead of outline and read_range when you have not located the symbol yet. Pass queries to find several names in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoNarrow by kind. func/function, type/struct/class/interface, method, const, var — spellings within a family are equivalent. Empty matches any.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
limitNoMaximum declarations to return (default 5). Prefer budget.
queryNoSymbol name, exact or partial.
budgetNoSize of the answer in tokens. Cut at whole declarations; the rest is behind continue=<handle>.
queriesNoSeveral symbol names in one call, instead of query. Each is answered as query would be.
continueNoHandle from a cut answer: its next page.
maxLinesNoMaximum lines of each body (default 40). Prefer budget.
dependencyNoLook in this dependency's source instead of the workspace, read-only: a crate, Go module, npm or Python package name.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes this is a safe read. The description adds genuine behavior beyond that: exact name matches outrank substring matches (a result-ordering rule that affects interpretation), and the call fuses search with body retrieval. Pagination via continue is left to the schema, so this is not fully self-contained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core value proposition before the routing advice and the batch hint. The shell-command analogy earns its place by compressing a lot of behavior into a few words. No padding, though the sentence count is near the upper bound for a description this dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-parameter read tool with no output schema, the description covers purpose, sibling routing, ranking behavior, and batch querying. The remaining gaps — result envelope shape and how continue/budget interact — are handled by the schema's own parameter descriptions, so an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a 3 is the baseline and the schema already documents all nine parameters thoroughly. The description still adds semantics the schema does not: that `queries` is a batch alternative to `query` and that match ranking favors exact names, which shapes how an agent should phrase query input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Locate declarations by name") plus a second action ("return their source in one call"), which makes the fused search-and-read behavior clear. The grep -n 'func X' -A 30 analogy immediately distinguishes it from a plain search tool. An agent can tell it apart from grep and read_range without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternatives ("Use this instead of outline and read_range") and the selecting condition ("when you have not located the symbol yet"). It also directs batch use via "Pass queries to find several names in one call." Nothing about when-to-use is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.grepA
Read-only

Literal or regex text search across the workspace, returning path:line matches with optional trailing context — the replacement for grep -rn. Use this for anything that is not a declaration name: struct fields, string literals, error messages, config keys, or any search needing a path filter. Use find instead when you want a declaration and its body. Pass queries to search several patterns in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
globNoRestrict by path, e.g. *.go or internal/code/*.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
limitNoMaximum matches returned (default 40). The true total is always reported. Prefer budget.
queryNoText to find.
regexNoTreat query as a regular expression. grep-style \| alternation and \( \) groups work as in grep.
budgetNoSize of the answer in tokens. Cut at whole matches; the rest is behind continue=<handle>.
contextNoTrailing lines to show per match, like grep -A (max 40).
excludeNoSkip paths containing this substring, e.g. testdata.
queriesNoSeveral patterns in one call, instead of query. Each is answered as query would be, with the same filters.
continueNoHandle from a cut answer: its next page. Other arguments except budget are ignored.
dependencyNoSearch this dependency's source instead of the workspace, read-only, at the version the project locks: a crate, Go module, npm or Python package name. Matches read as dep:<name>/<path>, which read_range accepts.
ignoreCaseNoCase-insensitive match.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description still adds real behavioral context beyond that: result truncation is cut at whole matches with the remainder behind continue=<handle>, the true total is always reported, and dependency search is read-only at the locked version. It stops short of describing match ordering or deduplication, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with what the tool does, then the distinction from `find`, then the multi-pattern tip. No filler; every clause carries routing or behavioral information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, no-output-schema tool with only a readOnly annotation, the description supplies return format, context lines, truncation/pagination, dependency scoping, and sibling routing. The only real gap is that it does not touch the glob/exclude/root/ignoreCase filters, but the schema documents those at 100% coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 12 parameters; baseline would be 3. The description nevertheless adds cross-parameter meaning the schema lacks: that `queries` exists to search several patterns in one call and each is answered 'as query would be, with the same filters', and that `continue` ignores arguments except budget.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('literal or regex text search across the workspace') and states the return shape ('path:line matches with optional trailing context'). It explicitly positions itself as the replacement for `grep -rn` and contrasts with the `find` sibling, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('anything that is not a declaration name: struct fields, string literals, error messages, config keys, or any search needing a path filter') and a clear alternative with its own condition ('use find instead when you want a declaration and its body'). The routing rule against siblings is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.insertA

Add text to a file without replacing anything — a new function, a new section, an extra case. Use this for additive work instead of rewriting a surrounding symbol. With no anchor it appends to the end of the file; with one it places the text before or after that anchor, refusing if the anchor is absent or matches more than once. Several additions or edits at once belong in apply.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile to add to.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
textYesText to insert.
anchorNoOptional. Exact, unique text to place the insertion beside. Omit to append to the end of the file.
positionNoOptional. "before" or "after" the anchor. Defaults to after.
expectedDigestNoOptional. The digest from the read this edit is based on; the edit is refused if the file changed since, by anyone.
expectedRevisionNoOptional. Revision expected before editing; the edit is rejected if the workspace has moved on. Omit for no precondition.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=false, so the description carries the real burden and delivers: it explains append-vs-anchor placement, and discloses two concrete refusal conditions (anchor absent, anchor matches more than once). This is failure-mode disclosure an agent cannot infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with purpose and the sibling routing, then behavior, then the escalation path. No sentence is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an insertion tool with no output schema, the definition covers purpose, alternatives, placement semantics, and refusal conditions — everything needed to invoke it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds inter-parameter semantics the schema states only per-field: that omitting anchor implies append and that position refines placement relative to anchor. It does not, however, explain expectedDigest/expectedRevision beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add text to a file'), then immediately contrasts it with the rewrite behavior of a sibling ('instead of rewriting a surrounding symbol'). An agent can distinguish this from arno.replace_text and arno.apply without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when to use it (additive work) and names the alternative for a different case ('Several additions or edits at once belong in apply'). Both the routing condition and the exclusion are present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.read_rangeA
Read-only

Read a file verbatim, whole or by line range — the replacement for cat and sed -n. Omit both line numbers to read the whole file, which is how to read go.mod, a Makefile, or any JSON/YAML/TOML config that has no symbols to address. An end line past the end of the file reads to the end. A dependency's source reads as dep:/, read-only. Several ranges, in one file or many, go in one call: {"ranges": [{"path": "a.go", "lines": "280-400"}, {"path": "b.go", "lines": "700-760"}]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository-relative or workspace-relative file path.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
linesNoLine range: "280-400", "280-" to the end, or "280". Omit to read the whole file.
budgetNoSize of the read in tokens (default 5000). Cut at whole lines; the rest is behind continue=<handle>.
rangesNoSeveral reads in one call, instead of path. A range that fails reports its error without failing the others.
continueNoHandle from a cut read: the rest of it.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered; the description adds real behavior beyond that — out-of-range end lines clamp to EOF, dependency sources read as dep:<name>/<path> read-only, and a failing range in a batch reports its own error without failing the others. It doesn't explain truncation/continue behavior in prose beyond what the schema says.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and its analogy before the edge cases. Information-dense with no filler, though the multi-range JSON example makes it longer than strictly necessary given the schema already documents the array.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries return-behavior weight and does so reasonably — it covers clamping, per-range error isolation, dependency-path reads, and (via schema) budget/continue. Minor gap: no explicit statement of the response shape for a normal read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning above the schema: it clarifies that omitting both line numbers reads the whole file and supplies a concrete multi-range example ('ranges' array shape) plus the dep:<name>/<path> path convention.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read a file verbatim, whole or by line range') and anchors it against familiar tools ('the replacement for `cat` and `sed -n`'), which cleanly separates it from file-mutating siblings like insert/replace_text and search siblings like find/grep.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditions: omit both line numbers for whole-file reads, which is the route for go.mod, Makefile, or symbol-less config formats — implicitly pointing symbol-addressing work at other tools. It stops short of explicitly naming a sibling (e.g., find/grep) for symbol-based lookup, so it's strong but not fully routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.replace_textA
Destructive

Replace an exact, unique string in a file. An anchor string does not move when the lines around it do, which is why follow-up edits address text rather than line numbers. Refuses when the anchor is absent or matches more than once — extend it with surrounding context to disambiguate. For several sites, use apply: atomic, one validation, no diagnostics from half-done intermediate states.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path to edit.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
newTextYesReplacement text.
oldTextYesExact text to replace. Must appear exactly once.
expectedDigestNoOptional. The digest from the read this edit is based on; the edit is refused if the file changed since, by anyone.
expectedRevisionNoRevision expected before editing.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, so the safety bar is lower, but the description adds the key behavioral guarantee: refusal on absent or ambiguous anchors, and how to fix it. Omits expectedDigest/expectedRevision semantics, which are confined to schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and rationale, then refusal conditions, then a sibling routing note. No waste; minor density in the final sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool without an output schema, the description covers purpose, failure modes, disambiguation, and multi-site routing. Missing digest/revision behavior context, but the schema covers those parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, documenting all six parameters including optional expectedDigest and expectedRevision. The description doesn't add parameter-level detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replace an exact, unique string in a file') and explains the anchor-based motivation. Clearly distinguishable from siblings like arno.insert and arno.create_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a specific routing rule: 'For several sites, use apply: atomic, one validation.' Also explains refusal conditions and how to disambiguate. Lacks explicit when-not-to-use beyond multi-site edits.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.run_commandA
Destructive

Run a command the repository declares, by name, and wait for the verdict: pass/fail with the decisive output. Use it instead of a shell for anything check does not cover. No name lists the declared commands. Nothing fits? declare_command it once.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoDeclared command to run. Omit to list the declared commands.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
waitNoWait for the result (default true). False returns a job ID to poll.
timeoutSecondsNoBound on the wait (default 90, max 300).

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral value by disclosing the return shape (pass/fail plus decisive output) in the absence of an output schema. However, annotations flag destructiveHint=true and the description says nothing about commands mutating the worktree, side effects, or that omitting 'wait' yields a pollable job. Some value added, but the destructive profile is left entirely to the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences with no filler, and the core purpose plus the 'no name lists commands' behavior are front-loaded. The telegraphic fragments ('Nothing fits? declare_command it once.') are efficient but slightly cryptic in tone.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, routing, listing, and escalation, and the schema handles root/wait/timeout defaults. With no output schema, the description supplies the return contract. What is missing is any hint about destructive side effects, which matters given destructiveHint=true.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented (name omission, root default, wait polling, timeout bounds). The description restates only the name-omission behavior, adding no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — run a repository-declared command by name — plus the outcome contract (pass/fail with decisive output). It explicitly contrasts itself with 'check' and with a raw shell, so an agent can place it among the siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives direct routing rules: use instead of a shell for anything 'check' does not cover, omit the name to list declarations, and escalate to 'declare_command' when no declared command fits. When-to-use and the alternative (declare_command) are both named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arno.run_testsA

Rerun a failing test, a test file, or changed files' tests; waits for pass/fail and the first failure. Full suite: check kind tests.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNoTest file to run, for scope=file (Go: its package). With scope=test, limits the name filter to this file.
rootNoWorktree of this repository to act in. Defaults to the session's workspace.
testNoTest name, for scope=test: exact in Go, the runner's name filter elsewhere (jest/vitest -t, ava --match, pytest -k, cargo test <name>).
waitNoWait for the result (default true). False returns a job ID to poll.
scopeNoall (default), file, test or changed.
timeoutSecondsNoBound on the wait (default 90, max 300).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only destructiveHint=false, so the description carries most of the behavioral burden. It discloses that the call blocks for pass/fail and surfaces the first failure, which is useful return information with no output schema to lean on, but it says nothing about side effects of running tests or behavior when wait=false beyond what the wait parameter already documents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight clauses with the primary action front-loaded and the alternative at the end; there is essentially no filler. The compressed phrasing of 'check kind tests' is slightly terse but does not waste space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterized test runner with full schema coverage but no output schema or rich annotations, the description covers the main scopes, notes the waiting behavior, and points to the sibling for full suites. It is largely sufficient, missing only edge behavior for wait=false and side effects, which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are fully documented in the schema itself. The description only restates scoping at a high level and adds no syntax or format detail beyond what the schema already provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Rerun ... tests') and enumerates the scope variants it covers (a failing test, a test file, changed files' tests), which lets an agent distinguish it from generic siblings like run_command. It explicitly routes the full-suite case to another tool ('check kind tests'), though the default scope=all in the schema creates mild tension with that routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditions for use (rerunning a failing test, a file, or changed files) and names an alternative for the excluded case (full suite -> check kind tests). It stops short of explicit when-not guidance beyond that one case, so it is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.0.15
    • Changedarno.apply1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.check1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.create_file1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.declare_command1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.delete_file1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.find1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.grep1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.insert1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.read_range1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.replace_text1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.run_command1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
    • Changedarno.run_tests1 field changed
      • addedInput schema / properties / root
        Added value: +{
        +  "description": "Worktree of this repository to act in. Defaults to the session's workspace.",
        +  "type": "string"
        +}
  2. 24 tool updatesv0.0.12
    • Addedarno.apply
    • Addedarno.check
    • Addedarno.create_file
    • Addedarno.declare_command
    • Addedarno.delete_file
    • Addedarno.find
    • Addedarno.grep
    • Addedarno.insert
    • Addedarno.read_range
    • Addedarno.replace_text
    • Addedarno.run_command
    • Addedarno.run_tests
    • Removedjade.apply
    • Removedjade.check
    • Removedjade.create_file
    • Removedjade.declare_command
    • Removedjade.delete_file
    • Removedjade.find
    • Removedjade.grep
    • Removedjade.insert
    • Removedjade.read_range
    • Removedjade.replace_text
    • Removedjade.run_command
    • Removedjade.run_tests
  3. 1 tool updatev0.0.11
    • Changedjade.run_tests1 field changed
      • changedInput schema / properties / scope / description
        Previous value: -"One of: all, file, test, changed. Defaults to all."New value: +"all (default), file, test or changed."
  4. 12 tool updatesv0.0.10
    • First observedjade.apply
    • First observedjade.check
    • First observedjade.create_file
    • First observedjade.declare_command
    • First observedjade.delete_file
    • First observedjade.find
    • First observedjade.grep
    • First observedjade.insert
    • First observedjade.read_range
    • First observedjade.replace_text
    • First observedjade.run_command
    • First observedjade.run_tests

TDQS

A4.1/5.0

Scored across 12 tools

Disambiguation4/5

Most tools have clearly distinct purposes (read_range vs find vs grep, insert vs replace_text vs apply), and descriptions cross-reference each other to route selection. Minor overlap remains among run_command, check, and run_tests, which all execute-and-report but differ mainly in scope/purpose.

Naming Consistency4/5

All names are snake_case, and most follow a verb_noun or verb pattern (read_range, replace_text, create_file, delete_file, declare_command, run_tests). A handful are bare verbs (find, insert, apply, check, grep), which is a minor deviation but still coherent and readable.

Tool Count5/5

12 tools is well-scoped for a repository coding agent, with each tool covering a distinct workflow (navigate, edit, execute, validate). No obvious redundancy or bloat.

Completeness4/5

Covers the full read/search/edit/execute/validate lifecycle including atomic multi-edits and checkpoints. Minor gaps like file move/rename or directory listing exist but are workable around via shell-style commands or create+delete.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    Runs a language server and provides tools for communicating with it. Language servers excel at tasks that LLMs often struggle with, such as precisely understanding types, understanding relationships, and providing accurate symbol references.
    1,600
    BSD 3-Clause
  • A
    license
    A
    quality
    A
    maintenance
    A fully featured coding agent that uses symbolic operations (enabled by language servers) and works well even in large code bases. Essentially a free to use alternative to Cursor and Windsurf Agents, Cline, Roo Code and others.
    29
    31,102 PyPI
    29,768
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    MCP server that keeps language server sessions warm and routes multiple languages through one process. Agents get persistent cross-file awareness, speculative execution (simulate edits before writing to disk), and 20 skills that encode correct multi-step operations like safe rename, blast-radius analysis, and end-to-end refactoring. Single Go binary, no runtime dependencies.
    50
    156
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local-first code intelligence and safety layer for AI coding agents. MCP server exposes dependency graph, impact analysis, and AST-compressed repo context, backed by typed local memory, patch-scope safety gates, and git-independent transaction rollback.
    1
    MIT