trailstone
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@trailstonerecord a decision: sessions use signed HttpOnly cookies, not JWT, scope src/auth/"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
trailstone
A decision ledger that lives in your repo. When a decision is reversed, trailstone flags every file the decision governed, so it gets re-checked before you build further — in your AI agent's context the moment it opens one, and at the push. Most flagged files will still be fine; the point is that the one that isn't no longer slips through green.
npx trailstone demo # the whole idea, on a throwaway repo, in ten secondsReal work with an AI is weeks of sessions. Every message steers a little, and what was decided three sessions ago quietly stops governing the work. Worse: when a decision is reversed, everything built on the old one stays finished, green, and wrong.
trailstone is one file in your repo and one script. No server, no account, no telemetry.
$ trailstone reverse d_6d0bf686 "Sessions use a signed HttpOnly cookie, not a JWT header"
d_259ab7a1 recorded (supersedes d_6d0bf686).
now stale (1):
src/auth/session.ts
$ git push
⚠️ STALE — these files were last committed BEFORE a decision governing them was reversed.
- src/auth/session.ts — was: Sessions use JWT in an Authorization header, not cookies
→ now: Sessions use a signed HttpOnly cookie, not a JWT header (Dana, 2026-09-06)
→ exit 1: the push is blocked.Your agent sees the same warning before it edits the file, with the old text and the new one, so it re-checks the work instead of building on ground that moved.
What it does and doesn't catch (read this before judging it)
It fires when a decision changes while code still rests on the old one — not when code
changes. A git revert edits the very files in scope, and any edit clears the flag, so
reverting a commit fires nothing. The tool is for the case a revert can't see: you decide "cookies,
not JWT" in conversation, and three files written under the old rule are still sitting there green.
A flag means "governed by," not "broken." Reversing a directory-scoped decision flags every
file under it. In real use most of them still comply with the replacement and clear in seconds —
the flag's job is to force the re-check that catches the one file that doesn't. If you write
--scope src/ you will get a big, noisy flag; if you write --scope src/auth/session.ts you get
a precise one. decide tells you the file count up front so you can narrow it before a future
reversal makes you re-check them all:
$ trailstone decide "Sessions use JWT headers, not cookies" --scope src/auth/
d_a1b2c3d4 recorded. scope src/auth/ covers 4 tracked files — a reversal will flag all 4 to re-check.Scope discipline is the whole game: a tight scope is a sharp tool, a whole-directory scope is an alarm you'll learn to ignore.
It only watches the files you scope. A decision's prose can still drift in a file no decision names — a README that describes an approach you've since reversed, a stale comment. That is real drift the tool will not flag, because it tracks scoped files, not every place a decision is described. Scope the docs you want watched, and treat the rest as your job.
Related MCP server: repo-history
Install
npm i -g trailstone # or use npx trailstone <command> everywhere below
cd your-repo # install must run INSIDE the repo
trailstone install # Claude Code hooks (global) + this repo's pre-push guard + AGENTS.md
trailstone init --goal "what this project is"
git add .trailstone AGENTS.md && git commit -m "trailstone: ledger"
trailstone doctor # confirms it is actually watchingRun install from inside the repo. Outside one it can only write the global Claude Code
hooks — you would get the warnings but not the pre-push guard that enforces them. It now
says so when that happens, and doctor tells you either way.
A repo without .trailstone/decisions.yml is silent — every hook exits immediately, so installing
globally costs nothing on repos that never opted in.
In CI
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # the guard compares commit dates; it needs the history
- uses: vishwamdhavale/trailstone@v0.2.0The convention
One file, committed with your code: .trailstone/decisions.yml. A YAML list, one
entry per decision. The PR that adds a line is the review.
# Trailstone decision ledger. One entry per decision: what was decided, why, and the
# paths it governs. Decisions are immutable — reverse one with a new entry carrying
# `supersedes: <id>`. Editing this file in a PR IS the review.
- id: d_6d0bf686
at: 2026-09-01T11:04:12.882Z
by: Dana
decision: Sessions use JWT in an Authorization header, not cookies
why: the API is called from a CLI too
scope:
- src/auth/
- id: d_259ab7a1 # the reversal — a new entry, never an edit
at: 2026-09-06T09:20:03.114Z
by: Dana
decision: Sessions use a signed HttpOnly cookie, not a JWT header
why: XSS token theft; the CLI gets a separate PAT
scope:
- src/auth/
supersedes: d_6d0bf686
- kind: validation # "I re-read this file; it still holds"
id: v_968dc1c7
at: 2026-09-06T09:41:55.001Z
by: Dana
decisionId: d_6d0bf686
scope:
- src/auth/session.tsFields: id (generated), at (ISO), by (git user.name), decision (what was
chosen, naming the alternative), why (optional), scope (globs / paths / dirs it
governs), supersedes (the id this reverses), status (proposed or rejected —
absent means it binds). Validation rows carry kind: validation, decisionId, and
the scope they re-checked, plus wrong: true for a false-positive fire.
Three rules, and that is the whole engine:
Immutable. You never edit a decision. Changing your mind is a new entry with
supersedes: <id>. A decision that something supersedes stops binding.Stale = a tracked file whose last commit predates a reversal of a decision whose
scopematches it. Scope match is exact path, directory prefix, or glob — never fuzzy.Clears three ways: edit the file in the working tree, commit it after the reversal, or add a
validationrow naming the decision (dated after the reversal).
status: proposed binds nothing: not in force, never flags anything stale.
What the hooks do
install adds four Claude Code hooks to ~/.claude/settings.json and writes
.git/hooks/pre-push in the current repo (if a pre-push already exists it prints
the line to add instead of clobbering it).
Hook | What it injects |
| One line: repo name, decisions in force, proposals pending — plus the stale block if any. |
| Decisions relevant to this prompt: the ones governing files you already touched this session, then a lexical top-up (≥2 shared words) — each labelled with why it surfaced. Plus stale. Capped at 5 + 5 proposals, never padded; silent when nothing matches. |
| Before the write lands: the decisions governing that exact file, and a stale warning if it has one. Once per file per session (remembered in a tmpfile, last 50 files), so a loop of edits does not repeat itself. |
| Passive capture, on by default. Detaches immediately — the session never waits — and writes anything it judged a real decision as |
Budget: the caps above mean a typical injection is a handful of lines; the largest is SessionStart with a long stale list, which is one line per stale file. Nothing dumps the ledger.
pre-push runs stale and exits 1 when a stale file is about to be pushed.
Other agents — Codex, Cursor, Claude Desktop, …
The ledger is a plain file in your repo, so anything can read it. Two portable surfaces ship today.
1. Agent rules — works everywhere. trailstone install writes a short block into
AGENTS.md (and into .cursor/rules/trailstone.mdc if the repo already uses Cursor),
telling any agent to run governing <file> before it edits and to honor what comes back.
Commit it, and every agent your team uses reads it at session start. --no-rules skips it.
2. An MCP server — any MCP client. trailstone mcp speaks MCP over stdio, with six
tools: list_decisions, governing, stale, decide, reverse, validate. No
dependencies, nothing to install beyond this script.
{
"mcpServers": {
"trailstone": {
"command": "npx",
"args": ["-y", "trailstone", "mcp", "--repo", "/absolute/path/to/your/repo"]
}
}
}--repo (or the TRAILSTONE_REPO env var) matters for desktop clients, which launch MCP
servers with an arbitrary working directory. A terminal agent already sitting in the repo
can leave it out. Leave it out and the server still starts — the first tool call then tells
the agent to pass its project path as the repo argument, which agents do recover from.
Claude Desktop: quit it before you edit the file. Claude Desktop holds its MCP config in memory and writes
claude_desktop_config.jsonback on quit, so an edit made while it is running is silently reverted when you close it — the app restarts looking perfectly configured, pointed wherever it was before. Quit it fully, then edit, then start it: that survives. (Its Settings → Local MCP servers → Edit config button just reveals the file, and on Linux may only open the folder.) Cursor and Codex read their config files normally and can be edited any time.On Claude Desktop, consider leaving
--repoout. It has no notion of a "current project", so a pinned repo means every question is answered from that one repo without ever saying so. With no--repothe first call asks the agent for the project path, and the repo is then explicit in the conversation.
3. Cursor hooks — push, not pull. If you have Cursor, install also writes
.cursor/hooks.json (merging into any hooks you already have). Cursor additionally imports
Claude Code's hooks and calls them under its own event names, which Trailstone answers too — so
push works whether or not you have Claude Code. Both routes are verified against a live Cursor
agent. sessionStart injects the
goal, the decisions in force and anything stale; preToolUse denies a write to a stale file
exactly once, handing the agent the reversal, and allows the retry. That one-time deny is the
only way Cursor lets a hook reach the agent before an edit — and it fires only on files that are
already blocked at push, so it interrupts nothing that was not going to be stopped anyway.
The honest limit. Push — the warning arriving unasked, before the edit — works in Claude Code and Cursor. Codex, Windsurf and Claude Desktop are pull-only: the agent has to ask, and agents do not reliably remember to. A hook shim for one of those is the most valuable contribution to this project.
CLI
init create .trailstone/decisions.yml (commit it)
decide "<what>" --why "<why>" --scope src/auth/,src/api/tokens.ts
reverse <id> "<new decision>" [--why ...] [--scope ...] # inherits the old scope if omitted
list [--all] decisions in force (--all includes superseded/proposed/rejected)
governing <file> which decisions bind this file
validate <id> --scope <file> "I re-checked it; it holds" — clears the stale flag
validate <id> --scope <file> --wrong the fire was a false positive; the file never rested on it
stats fires by surface, and precision (right vs wrong)
report [--anon] [--json] paste-ready summary: ledger, fires, precision, current stale
demo [--keep] the whole loop on a throwaway repo, in ten seconds
proposed pending captures
ratify <id> / reject <id> accept or drop one
stale the guard: prints stale files, exit 1 (exit 0 = clean)
doctor is Trailstone watching this repo? exit 1 if installed but blind
capture-health the last 5 judge runs; exit 1 if the last one failed
hook Claude Code hook dispatch (stdin JSON) — not for humans
install wire the hooks + pre-push
uninstall remove that wiring (leaves your ledger alone)Real example:
trailstone decide "Sessions use JWT in an Authorization header, not cookies" \
--why "the API is called from a CLI too" --scope src/auth/
trailstone reverse d_6d0bf686 "Sessions use a signed HttpOnly cookie, not a JWT header"
# → now stale (1): src/auth/session.tsRecording decisions
Explicit first. You run decide when you make a call. Your agent runs the same
command the moment it says "decided" — the trailstone-decide skill in
.claude/skills/trailstone-decide/ tells it when and how. An explicit decision is a real
choice that forecloses an alternative; state the alternative in the text
("X, not Y") so the reversal reads as a diff.
Passive second — and it is what actually fills the ledger. The Stop hook is on by
default: at the end of every turn it detaches a cheap claude -p judge (haiku, ~$0.05–0.10
per turn) that appends anything that looks like a decision with status: proposed. It needs
the claude CLI on your PATH; without it the hook exits silently. Turn it off with
TRAILSTONE_CAPTURE=0. A judge run that dies (expired auth, timeout) is invisible by design — the
worker is detached — so every run logs one line to ~/.trailstone/capture.log; capture-health
prints the last five and exits 1 if the last one failed.
If Trailstone seems quiet, run doctor before believing the ledger is empty. Every hook here
fails open — it can never block your prompt — which means a hook that is blind looks exactly
like a hook with nothing to say. doctor reports the repo, the ledger and whether it is
committed, the four hooks, the pre-push guard and the last judge run, and exits 1 if anything
is missing. The most common cause is the simplest: your session is sitting one directory
above the repo, so there is no .git to walk up to. doctor looks one level down and names
the directory you should be in.
Proposals bind nothing: not in force, never flag anything stale, until a human ratifies. Review them as what they are — a diff in a file you own:
keep it: delete the
status: proposedline in your editor, orratify <id>drop it: delete the entry, or
reject <id>to keep the record
What it deliberately does not do
No server, no token, no account. Git already gives provenance (blame), membership (who can push), review (the PR), and sync (clone).
No plan, no steps, no DAG. It tracks decisions and files, not work.
No semantic judging of what is stale. Staleness is git timestamps and glob membership, full stop.
No fuzzy matching on anything that gates. The lexical prompt match only surfaces a decision to the agent; it can never flag a file or fail a push. A false stale flag is worse than a missed one.
It does not follow file renames. Scope is a path;
git mv-ing a governed file to a new name silently drops it from that decision's scope until you re-scope (record a new decision naming the new path). Following renames would mean heuristic rename detection in the gate, and the gate stays exact, never heuristic — so this is on you: after a rename, re-scope the decision.It does not auto-merge concurrent decisions. Two branches that each append a decision produce an ordinary git conflict in
.trailstone/decisions.yml; resolve it by keeping both entries. The append-only list makes this the easy kind of conflict, but it is still a manual resolve.
The one metric
Fires that were right versus fires that were wrong. Nothing else.
What it looked like on a real project. One reversal on a 79-file product repo flagged
9 governed files: 7 were re-checked and still held, 1 genuinely needed rework, and 1 was
a false positive the developer marked --wrong. 89% precision. That is the shape to
expect — most flagged files still comply, and the flag's job is to force the re-check that
finds the one that does not. (Measured, not projected; it was a deliberate re-validation
sweep after a real decision changed, not a surprise catch.)
When a stale warning fires and the file genuinely needed re-checking, validate <id> --scope <file> records the good fire. When it fired on a file that never rested on
that decision, validate <id> --scope <file> --wrong records the false positive.
stats prints fires by surface and the resulting precision. If precision drops,
scopes are too broad — narrow them; do not soften the rule.
Fires are the rare event; surfacings are the denominator. A healthy repo goes a
long time with zero fires — that does not mean the tool did nothing, it means every
decision it surfaced still held. So stats also counts every time a governing decision
was actually put in front of someone — split into pushed to the agent unasked
(the session/prompt/edit hooks, MCP) and pulled on demand (governing, list). A
clean repo then reads surfaced 40×, fired 0 instead of an empty log — the difference
between "kept the agent on course 40 times" and "was never even consulted". (This is a
count of exposure, not proof the agent obeyed — that only a reversal it honors can show.)
report is the same numbers in a form you can paste into an issue; report --anon
drops every name and reduces each path to its extension, so it is shareable from a
private repo. That is how a maintainer learns whether the fires were any good.
Roadmap
A GitHub App turning
staleinto check-run annotations on the exact lines.import/export so a hosted, cross-repo team view can read the same yml.Codex, Cursor, and other harnesses via their own hook shims over the same file.
A
git blame-shapedhistory <file>: every decision that ever governed it.Scope suggestions from the diff, so
deciderarely needs--scopetyped by hand.
Contributing
VISION.md — the idea, the invariants, and an honest list of what this does not do. Read it before judging the tool or proposing a change.
CONTRIBUTING.md — how to work on it, and what gets declined on principle (dependency cascade, servers, fuzzy gates).
AGENTS.md — instructions for AI agents contributing to this repo.
SECURITY.md — the whole surface, including the one thing that leaves your machine (the optional capture judge).
The most useful thing you can send back is not a star — it is trailstone report --anon:
how often the stale warning fired on your repo, and whether those fires were right.
This server cannot be deployed
Maintenance
Related MCP Connectors
Append-only decisions with provenance, supersession, retrieval, and audited MCP actions.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Manage portable AI agent playbooks, Agent Skills, MCP configurations, personas, and memory.
Team docs served to AI agents over MCP - search, Markdown reads, version pinning, read audit.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables preserving, querying, and managing engineering decisions with reasoning and evidence through CLI, HTTP API, and MCP tools.16-
- AlicenseNot gradedqualityBmaintenanceExposes a repo's historical engineering memory (decisions, landmines, guardrails) via MCP tools for AI coding agents.MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI coding agents and hosts to enforce deterministic repository boundaries via MCP, providing structured reads, supervised edits, snapshots, audits, and recovery with machine-readable evidence.MIT
- AlicenseNot gradedqualityAmaintenanceProvides MCP tools that give AI agents persistent, append-only memory in a git repository, letting them record observations, decisions, and corrections while retrieving context briefs, current facts, conflicts, and traceable event history without a vector database.4 npmMIT