Stellar Jay
Stellar Jay
Safe write access for AI agents. Stellar Jay is where your agents keep business data. Every change is kept, attributed to the agent that made it, and can be undone. Think of it as Git for your business data.
Hosted: AvianSuite runs Stellar Jay for you. $20/month per store, with a 14-day free trial.
Self-hosted: free and open source under the AGPL. The quick start below gets a store running on your machine in a few minutes.
Agent setup
Connect an agent with the MCP server. On AvianSuite, add the remote server:
# Claude Code
claude mcp add --transport http aviansuite https://mcp.aviansuite.com/mcp// Cursor: .cursor/mcp.json
{"mcpServers": {"aviansuite": {"url": "https://mcp.aviansuite.com/mcp"}}}In Claude and ChatGPT, add a custom connector with the same URL.
For a self-hosted store, run the local server instead:
{
"mcpServers": {
"aviansuite": {
"command": "stellarjay-mcp",
"env": {"STELLARJAY_URL": "https://store.example.com", "STELLARJAY_TOKEN": "writer-token"}
}
}
}Install it with go install github.com/kyle-visner/stellarjay/cmd/stellarjay-mcp@latest.
Tools and details: docs/mcp.md.
Related MCP server: ForkMind
Why
Clients want agents that do the work, not just read about it. But when an agent writes straight into a CRM or ticketing system, one bad decision or runaway loop can overwrite or delete records, and there is often no way back.
Stellar Jay is built so that can't happen:
Nothing is overwritten or deleted. A correction or retraction is a new entry, and the original stays in history.
Every change has a name on it. Each agent gets its own token, so you can see exactly which agent changed what, and when.
Any change can be undone. Roll back everything one agent did in a time window, with a preview first.
History is tamper-evident. If anyone rewrites or removes past entries, it shows.
Retries are safe. An agent that retries after a timeout doesn't create duplicates, and a write based on stale information is refused.
Your data stays flexible. Facts are JSON, so new fields and new kinds of records need no migrations.
Stellar Jay records what agents say happened. It doesn't decide whether a fact is true; it makes sure a wrong one stays visible and correctable.
Who it's for
Developers, consultants, and small teams moving from read-only copilots to agents that are allowed to act: operations, accounting, approvals, support, and other work where the data matters. Each store serves one organization. Many agents and apps can share it.
Quick start (self-hosted)
Requires Go 1.22 or later.
1. Install and create secrets.
go install github.com/kyle-visner/stellarjay/cmd/stellarjay-server@latest
stellarjay-server init ./secretsinit creates a data encryption key and prints an admin, writer, and reader
token once. Save them in a password manager.
2. Start the server.
export STELLARJAY_DATA_DIR=./data
export STELLARJAY_DATA_KEY_FILE=./secrets/data_key
export STELLARJAY_AUTH_FILE=./secrets/auth.json
stellarjay-server serveIt listens on 127.0.0.1:8080.
3. Record a fact. In another terminal:
export STELLARJAY_URL=http://127.0.0.1:8080
export STELLARJAY_TOKEN='the-writer-token'
curl -fsS -X POST "$STELLARJAY_URL/v1/events" \
-H "Authorization: Bearer $STELLARJAY_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: first-fact" \
--data '{
"type": "business.fact",
"entity_id": "customer-42",
"command": "fact assert",
"payload": {"predicate": "primary_contact", "value": "Ada Lovelace"},
"expected_root": ""
}'expected_root is empty only for the first entry in a new store. After that,
read the current value from GET /v1/root and send it with each write, so a
write based on stale information is refused. The API guide
covers reading history, pagination, named checkpoints, and snapshots.
4. Connect an agent. Point the MCP server at your store:
go install github.com/kyle-visner/stellarjay/cmd/stellarjay-mcp@latestThen use the self-hosted config from Agent setup with
STELLARJAY_URL=http://127.0.0.1:8080 and your writer token. Agents that don't
use MCP can read $STELLARJAY_URL/llm.txt, which explains how to work with the
store.
Deploy to a server
For a production store with HTTPS, you need a Linux host with Docker Compose, ports 80 and 443 open, and a DNS record pointing a domain at the host.
git clone https://github.com/kyle-visner/stellarjay.git
cd stellarjay
cp .env.example .env
# Edit .env and set STELLARJAY_DOMAIN.
go run ./cmd/stellarjay-server init ./secrets
docker compose up -d --build
curl https://stellarjay.example.com/health/readyinit will not replace existing secrets. Read the
operations runbook for backups, token rotation, and
upgrades, and the security model before storing sensitive
data. Or skip all of this and use AvianSuite.
Documentation
llm.md: the guide agents follow to read and write safely
docs/mcp.md: MCP server setup and tools
docs/how-it-works.md: storage model, design limits, and the embedded Go library
docs/api.md: HTTP API reference (OpenAPI)
docs/architecture.md, docs/security.md, docs/operations.md: running it in production
Development
GOCACHE=/tmp/stellarjay-gocache go test -race ./...
GOCACHE=/tmp/stellarjay-gocache go vet ./...
docker compose config
docker build -t stellarjay:test .License
AGPL-3.0-or-later. See LICENSE. Hosted Stellar Jay is available from
AvianSuite.
Available Tools
8 toolscorrect_factCorrect a factAInspect
Use when a recorded fact is wrong and you know the right value. Give the hash of the fact to replace and a reason. The old fact stays in history, marked as superseded, so the correction can itself be undone. When the result includes a receipt link, include it when you tell a person about the change, so they can check it and undo it.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | The corrected value. Any JSON. | |
| reason | Yes | Why the earlier fact was wrong. | |
| evidence | No | Where the fact came from, so a person can check it. | |
| entity_id | Yes | Stable ID of the thing the fact is about, such as customer:42 or ticket:T-1009. Not a display name. | |
| predicate | No | Predicate of the corrected fact. Defaults to the predicate of the fact being corrected. | |
| supersedes | Yes | Hash of the fact being corrected, from get_entity or an earlier write. | |
| observed_at | No | When the corrected value was observed, RFC 3339. | |
| operation_id | Yes | Your ID for this one write, 8-200 characters. Reuse it when retrying the same write so it is not recorded twice. Never reuse it for a different write. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety hints (non-read-only, non-destructive, closed-world), while the description adds the non-obvious behavior: the old fact remains in history marked superseded and the correction itself can be undone. It also tells the agent to surface the receipt link to a human so the change is verifiable and reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the usage trigger front-loaded, then the mutation mechanics, then the human-facing receipt guidance. No filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an eight-parameter mutation with a nested evidence object and no output schema, the description covers the essential safety and reversibility context and hints at the receipt link in the result. It leaves the return shape and idempotency mechanics to the schema, which is acceptable but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all eight parameters including supersedes, reason, and operation_id. The description only restates the supersedes hash and reason at a high level, adding no format or edge-case detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (correct/replace) and resource (a recorded fact), plus the corrective scope: the old fact is superseded rather than deleted. This cleanly separates it from siblings like record_fact (new fact) and retract_fact (removal), without needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use when a recorded fact is wrong and you know the right value" gives a clear trigger condition that implicitly excludes retraction (when a fact is no longer true) and fresh recording. It stops short of explicitly naming those sibling alternatives, so it is strong but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_entityGet an entityARead-onlyInspect
Use to read what is currently true about one entity, and how it got that way: every fact, correction and retraction, who made it and when.
| Name | Required | Description | Default |
|---|---|---|---|
| entity_id | Yes | Stable ID of the thing the fact is about, such as customer:42 or ticket:T-1009. Not a display name. | |
| include_history | No | Include every fact event, not only current facts. Default true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), lowering the bar. The description adds useful content-level context (facts, corrections, retractions, who and when), but says nothing about pagination, volume limits, auth requirements, or behavior when the entity_id is unknown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence, front-loaded with the use case and ending in a colon-delimited enumeration of what is returned. Nothing is wasted, though the trailing enumeration is doing descriptive work the schema partly duplicates.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with annotations and no output schema, the description is nearly sufficient: it conveys purpose and the shape of the returned data. Minor gaps remain around unknown-ID behavior and result size/pagination, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are fully documented in the schema, so the baseline is 3. The phrase "every fact, correction and retraction" loosely signals the include_history dimension, but the description adds no format or default detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (read) and resource (one entity) plus the scope of what is read: current facts, corrections, retractions and provenance. This implicitly contrasts with the mutation siblings (record_fact, correct_fact, retract_fact) and with list_changes, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use to read what is currently true about one entity" gives an implied trigger, and "and how it got that way" hints at the history variant. However, there is no explicit when-not guidance, no prerequisite statement, and no routing to sibling tools such as list_changes for multi-entity queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_changesList changesARead-onlyInspect
Use to see what changed in a time window, optionally only one agent's changes or one entity's. Use it to review an agent's work, and before undo_changes to find the actor name and window.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | Only changes made by this actor (the agent credential name shown in results). | |
| limit | No | Most changes to return, newest first. Default 50. | |
| since | Yes | Start of the window: an RFC 3339 time, or a duration back from now such as 30m, 2h or 24h. | |
| until | No | End of the window, RFC 3339. Defaults to now. | |
| entity_id | No | Only changes to this entity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the workflow context (results carry an actor credential name usable by undo_changes), but says nothing about result shape, ordering guarantees beyond the schema's 'newest first', or whether the window is capped. Adds modest value over annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core capability and followed immediately by the workflow guidance. No filler, no restatement of the title or parameter list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, 5-parameter listing tool with no output schema, the description covers purpose, filtering scope and workflow placement. It is nearly complete; only the return payload's structure (what a 'change' record contains) is left implicit, which matters slightly given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (actor, limit, since, until, entity_id) are already documented with types, bounds and defaults in the schema. The description adds no syntax, format or interaction detail beyond what the schema provides — baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('see what changed') scoped to a time window, with optional narrowing to one actor or one entity. The scope is unambiguous, though it never contrasts itself against a genuinely competing sibling (e.g. status or get_entity) — it only names undo_changes as a downstream consumer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions ('review an agent's work') and an explicit workflow handoff ('before undo_changes to find the actor name and window'), which tells the agent what prerequisite data this tool supplies to a sibling. Nothing about invocation timing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_factRecord a factAInspect
Use when you learn something about a customer, ticket, order or other business record and want it kept. Use this before, and instead of, overwriting records in other systems: the fact is kept with its evidence and attributed to you, and it can be corrected or undone later. Recording a new value for the same entity and predicate makes it the current value; the old one stays in history. When the result includes a receipt link, include it when you tell a person about the change, so they can check it and undo it.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | The value. Any JSON: string, number, boolean, object or array. | |
| evidence | No | Where the fact came from, so a person can check it. | |
| entity_id | Yes | Stable ID of the thing the fact is about, such as customer:42 or ticket:T-1009. Not a display name. | |
| predicate | Yes | What the fact is about, such as status, email, owner or amount_due. | |
| confidence | No | Optional confidence from 0 to 1. Leave unset when you are sure. | |
| observed_at | No | When the fact was observed at the source, RFC 3339. Defaults to unset. | |
| operation_id | Yes | Your ID for this one write, 8-200 characters. Reuse it when retrying the same write so it is not recorded twice. Never reuse it for a different write. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly=false, destructive=false, openWorld=false). The description adds substantive behavior the annotations do not: facts are kept with evidence and attributed, re-recording the same entity+predicate makes it current while the old value stays in history, writes are idempotent via operation_id, and results may carry a receipt link enabling undo. That is meaningful disclosure beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the usage trigger and organized around use/behavior/receipt. Every sentence carries information, though the middle sentences are dense and pack versioning, attribution, and idempotency together; a small amount of trimming would help.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema, the description covers the essential return behavior (a receipt link to relay and use for undo), the history/undo semantics, and idempotent retry. It does not address failure modes or required permissions, leaving a small gap for a mutation tool with 7 parameters and a nested evidence object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents each parameter (evidence, entity_id, predicate, value, confidence, observed_at, operation_id) with examples and constraints. The description adds semantics the schema does not: the entity+predicate pair defines identity and re-recording updates the predicate's current value. That is a real increment above the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: recording a fact learned about a customer, ticket, order or other business record, with the entity/predicate model making it distinguishable from correct_fact and retract_fact. It does not name those siblings explicitly, so the agent must infer the boundary, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ("Use when you learn something...") and an explicit alternative to avoid ("before, and instead of, overwriting records in other systems"). It stops short of naming the sibling tools (correct_fact, retract_fact) that an agent would pick between, so the routing guidance is partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retract_factRetract a factAInspect
Use when a recorded fact should no longer count and there is no replacement value. Give its hash and a reason. The fact stays in history, marked as retracted. Retracting a retraction restores the original fact. When the result includes a receipt link, include it when you tell a person about the change, so they can check it and undo it.
| Name | Required | Description | Default |
|---|---|---|---|
| hash | Yes | Hash of the fact or retraction to withdraw. | |
| reason | Yes | Why it should no longer count. | |
| entity_id | Yes | Stable ID of the thing the fact is about, such as customer:42 or ticket:T-1009. Not a display name. | |
| operation_id | Yes | Your ID for this one write, 8-200 characters. Reuse it when retrying the same write so it is not recorded twice. Never reuse it for a different write. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses the fact stays in history marked as retracted, that retracting a retraction restores the original, and how to handle the receipt link. None of this is derivable from readOnlyHint=false / destructiveHint=false, and it materially changes how an agent should report and retry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the usage condition, and each carries distinct information (when, what to pass, lifecycle behavior, receipt handling). Slight redundancy in the 'give its hash and a reason' instruction keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description supplies the missing lifecycle and reversal semantics plus the receipt-link handling, which is exactly the gap an agent would otherwise have to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents hash, reason, entity_id and operation_id fully. 'Give its hash and a reason' restates the schema without adding syntax or edge cases, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('a recorded fact should no longer count') and scopes it with 'there is no replacement value', which cleanly separates it from correct_fact. An agent can pick this tool over its siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The opening clause gives a clear condition for use (fact should no longer count, no replacement). The contrast with correct_fact is strong but implied rather than naming the sibling explicitly, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_checkpointSave a checkpointAInspect
Use before a risky job, such as an import, to name the current state (for example before-import). Saving an existing name moves it to the current state. Checkpoints never change any facts.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Checkpoint name: letters, digits, dot, dash or underscore, up to 100 characters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a non-read-only, non-destructive, closed-world write. The description adds real behavioral context beyond that: re-saving an existing name moves the checkpoint to the current state (overwrite semantics), and 'Checkpoints never change any facts' reassures the agent that no domain data is mutated. That is meaningful disclosure the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: usage cue first, then overwrite semantics, then a side-effect disclaimer. Every sentence carries information and nothing is repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema is needed since the save operation returns little, annotations cover the safety profile, and the single parameter is fully documented in the schema. The description fills the remaining gaps (when to call, overwrite behavior, no fact mutation), so an agent has everything required to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already defines the name format and length limit, so the baseline is 3. The description adds a concrete example name ('before-import') that conveys naming conventions beyond the raw schema constraint, lifting it slightly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (naming/saving the current state as a checkpoint) with a concrete example ('before-import'), so an agent knows what the tool produces. It doesn't explicitly contrast with siblings like undo_changes or list_changes, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger: 'Use before a risky job, such as an import.' That is a clear when-to-use context. It does not state when-not to use it or name an alternative sibling, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusStatusARead-onlyInspect
Use to check that the store is reachable and healthy, and to get its current root.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the description adds that the call verifies reachability/health and returns the store's current root, which is the only signal about return content since there is no output schema. It does not explain what 'healthy' or 'root' concretely mean.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action verb first and zero filler. Every clause earns its place by conveying a distinct facet (reachability, health, root).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-arg, read-only tool with annotations covering safety, the description supplies enough to call it correctly. The absence of an output schema means the exact return shape is only hinted at by 'current root,' leaving a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline of 4 applies; no parameter semantics are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'check that the store is reachable and healthy,' plus a secondary function ('get its current root'). An agent can tell this apart from the fact-mutation siblings (record_fact, retract_fact) since none of them are health checks, though no sibling is named explicitly to sharpen the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The imperative 'Use to check that the store is reachable and healthy' gives a clear use context for a pre-flight/health check scenario. It offers no explicit when-not conditions or named alternatives, but the sibling set makes the alternative path obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
undo_changesUndo changesAInspect
Use to reverse everything one agent did in a time window, for example after a bad import or a runaway loop. By default this is a dry run that lists what would be reversed. Call again with confirm: true to write the reversals. Each reversal is a new retraction, so an undo can itself be undone. Only facts can be undone; other events are listed as skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | Yes | The actor whose changes to reverse, as shown by list_changes. | |
| since | Yes | Start of the window: an RFC 3339 time, or a duration back from now such as 30m, 2h or 24h. | |
| until | No | End of the window, RFC 3339. Defaults to now. | |
| reason | No | Why the changes are being undone. Recorded on every reversal. | |
| confirm | No | Set true to write the reversals. Leave unset for a dry run. | |
| entity_id | No | Only reverse changes to this entity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: default dry-run mode, the confirm flag as the write switch, that each reversal is itself a new retraction (so undos are undoable), and that non-fact events are listed as skipped. The reversible, non-destructive nature aligns with destructiveHint=false rather than contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences with zero filler, front-loaded with purpose then immediately the dry-run safety contract. Every sentence carries distinct information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description explains what a dry run returns ('lists what would be reversed'), what gets skipped, and how to commit — enough for an agent to plan the two-call workflow correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters including confirm's dry-run semantics and the duration shorthand for since. The description reinforces the confirm/dry-run contract but adds no syntax or format detail the schema lacks; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reverse/undo) and scoped resource (everything one agent did in a time window), with concrete motivating examples (bad import, runaway loop). It is clearly distinguishable from sibling point-fixes like retract_fact or correct_fact by its bulk, actor-and-window scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context ('after a bad import or a runaway loop') and an internal call sequence (dry run first, then confirm: true). It stops short of naming an alternative tool for narrower single-item corrections, which would have made the routing explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.6.0- First observed
correct_fact - First observed
get_entity - First observed
list_changes - First observed
record_fact - First observed
retract_fact - First observed
save_checkpoint - First observed
status - First observed
undo_changes
TDQS
Scored across 8 tools
Each tool targets a distinct operation: status for health, record/correct/retract for the three fact-mutation paths, get_entity for reads, list_changes for auditing, undo_changes for reversal, and save_checkpoint for snapshots. The record/correct/retract trio is cleanly separated by intent (new value vs. replacement vs. no replacement), so there is no realistic misselection risk.
Most tools follow a clear verb_noun pattern (record_fact, correct_fact, retract_fact, get_entity, list_changes, undo_changes, save_checkpoint). The lone outlier is 'status,' a bare noun, but the deviation is minor and does not impair readability.
Eight tools map neatly onto the fact-store lifecycle with no redundant or filler entries. The count is well within the ideal range for a focused append-only fact/audit service.
The surface covers the full fact lifecycle (record, correct, retract), reading current state, auditing changes, bulk undo, checkpointing, and health. The only soft gap is the absence of a way to list or restore named checkpoints explicitly, which agents can largely work around by re-saving by name.
Maintenance
Related MCP Connectors
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Open-source CRM your AI agents can write to: companies, people, deals, tasks, notes, pipeline.
An agent-native database over MCP: shared, validated, structured records in every AI chat.
MCP-first toolbox for agents: KV storage, auth, queue, and utility tools. Free in early access.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAudit-grade memory backbone for agent teams. Bi-temporal facts (event time + transaction time, with recall(as_of=...) replay), 6-step deterministic retrieval (no LLM in the critical path), conversation ingest with speaker-locked dual-pass extraction, per-tenant Postgres row-level security, and Ed25519-signed provenance. Postgres + pgvector + Neo4j defaults.37 PyPI14MIT
- AlicenseNot gradedqualityAmaintenanceLocal-first MCP server that lets AI agents query their own LLM call history as a branchable DAG and offload conversation context into immutable, AES-256-GCM-encrypted capsules — restorable in full or per segment, crypto-shreddable, with RAID-style replication. 12 tools, no API keys, no cloud.44 npm3MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to maintain persistent, inspectable understanding through typed, revisable updates, and to coordinate multi-agent work via shared graph-based stigmergy.81 npm1MIT
- AlicenseNot gradedqualityAmaintenanceUniversal verifiable recovery for long-running AI agents with semantic checkpoints, idempotent action ledger and hash chained log as a deny by default MCP server. Framework agnostic with adapters for LangGraph, LangChain and OpenAI plus gateway and OTel.28Apache 2.0