Arroway
Server Details
Shared memory for people and their AIs: read what was decided before acting, log what was done.
- Status
- Healthy
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
- Repository
- oaleviola/arroway-plugin
- GitHub Stars
- 1
TDQS
Scored across 19 tools
The set is large and includes several read-adjacent tools (catch_up, read, norms, search_log) plus lifecycle tools (claim, close, pass), but descriptions carefully delineate when to use each. Minor risk of confusion between session-start digest and task-specific read, but not enough to impair selection.
Every tool uses the arroway_ prefix followed by a clear snake_case verb or verb_noun phrase. The convention is consistent throughout, with no mixing of styles.
The broad memory-governance domain justifies many operations, but 19 tools is on the heavy side and some read or handoff variants could potentially be consolidated. It is borderline rather than clearly over-scoped.
The surface covers memory creation, read, update via supersedes, retirement, withdrawal, movement, export/import, verification, and handoff lifecycle. Gaps include no explicit project update/delete or standalone memory search beyond handles/norms, but core workflows are well supported.
Available Tools
19 toolsarroway_catch_upCatch up on what happened, across the person's projectsARead-onlyIdempotentInspect
Call this ONCE at the START of a session, before there is a task: it opens a rolling, expandable digest of what happened recently across every project this person belongs to, using each entry's authored essence, AND opens with the roster of their projects — every slug and what each one is for. Open handoffs are a separate block from completed work and are never displaced by the log. When the digest spans responses, continue with the read_id and receipt returned by the last part: they are the cursor over one frozen window. Use arroway_read include:["#handle"] when the full body matters. This is the 'where do things stand' read; arroway_read is the task-specific one.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How many days back, counting today. Default 2 — yesterday and today, which is what a session normally needs to know where things stand. Ask for more only when you are picking up work after being away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds valuable behavioral context beyond annotations: the digest is rolling and expandable, open handoffs are kept in a separate block and never displaced, and read_id/receipt act as a cursor over a frozen window during paginated reads. This meaningfully enriches the annotation-only picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every sentence carries distinct operational value: the start-of-session instruction, the digest contents, handoff separation, pagination mechanics, and the sibling distinction. It is front-loaded with the most critical call-time instruction, and no sentence is filler. Slightly dense, which prevents a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers what an agent needs to invoke it correctly: when to call it, what it returns conceptually, how to continue multi-part digest reads with the cursor, and when to delegate to arroway_read. It also accounts for the session-start workflow and the non-displacement of open handoffs, making it complete in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the single optional days parameter is already well documented in the schema with default, bounds, and a usage cue ('Ask for more only when you are picking up work after being away'). The description does not need to repeat parameter details, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it opens a rolling, expandable digest of recent activity across every project the person belongs to, plus a roster of projects. It explicitly distinguishes itself from the sibling arroway_read ('this is the where-do-things-stand read; arroway_read is the task-specific one'), so an agent can select it correctly without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit usage rule: 'Call this ONCE at the START of a session, before there is a task.' It also names the alternative arroway_read for full-body content and contrasts session-level catch-up with task-specific reading, providing both when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_claimTake an open handoffAIdempotentInspect
Mark that you are taking an open handoff, BEFORE working on it, so parallel sessions see it is taken and do not duplicate the work. A claim is a mark, never a lock: it stops nobody, a stale claim is information, and claiming again overwrites the mark while saying whose it was. Addressing rule: addressed to YOUR user — take it; addressed to someone else — only a human in this conversation can decide to take it anyway, and an autonomous session must leave it. When the work lands, close it with arroway_close.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Who is taking it, in plain words — a routine names itself here (e.g. 'nightly build routine'). | |
| handoff | Yes | The #handle from the read (8 chars), or the full id. | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavioral nuance beyond annotations: a claim is a mark, never a lock, stops nobody, overwrites on re-claim, and a stale claim is information. This supplements idempotentHint=true and clarifies non-destructive semantics. The addressing rule also discloses how autonomous vs. human sessions should behave.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each carrying distinct information: the action and rationale, the mark-versus-lock semantics, the addressing rule, and the follow-up tool. The main action is front-loaded and every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers behavior, timing, ownership rules, and follow-up, but leaves the project parameter unexplained and doesn't mention what a successful claim returns. Given no output schema, these are minor omissions; the core behavior is well specified enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with project lacking a description. The description adds context for handoff and note (e.g., 'saying whose it was') but does not explain the project parameter at all. It provides some meaning beyond the schema but does not fully compensate for the missing project semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb and resource: 'Mark that you are taking an open handoff,' and clearly scopes the action with 'BEFORE working on it' and the parallel-session purpose. It also contrasts a claim with a lock and names arroway_close as the follow-up, making the tool's role distinct among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('BEFORE working on it') and why (prevent duplicate work). It provides clear when-not guidance via the addressing rule: autonomous sessions must leave handoffs addressed to someone else, while only a human may decide otherwise. It also routes to arroway_close when work lands.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_closeClose a handoff — the work landedAIdempotentInspect
Close an open handoff when the work it carried is DONE — call it alongside the arroway_log entry that records the residue. Closing without a completion log leaves the team an ending with no story. Work that removes what an open handoff exists to fix closes that handoff in the same pass — even when the handoff is someone else's. So this is also the call for someone ELSE's handoff, when your work emptied it: the file it wanted fixed left circulation, the card it pointed at was cancelled, the decision behind it was reverted. Close it in that same pass, with a note naming what emptied it — an open handoff whose object is gone keeps billing a person for work that no longer exists. What is NOT this: a handoff that still carries real work, which a passer-by must never close. And if the work turned out not to be worth finishing, that is not yours to decide alone: discarding a handoff is the human's gesture, in the panel.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | One line on what landed — it points at the residue, it does not replace it. | |
| handoff | Yes | The #handle from the read (8 chars), or the full id. | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish mutating, idempotent, and non-destructive traits. The description adds contextual behavior: closing without a completion log leaves 'an ending with no story,' and an open handoff whose object is gone 'keeps billing a person for work that no longer exists.' This enriches the safety profile without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary use case, but it is verbose with multiple conditional clauses and rationale. While every sentence adds context, the length is excessive for a tool description. It could be tightened without losing essential guidance, so it earns a middle score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nuanced edge cases (closing others' handoffs, the companion log requirement, and the exclusion of active handoffs), the description covers all relevant scenarios and constraints. No output schema exists, so return values are not required. It gives an agent enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema describes handoff and note; project lacks a description but has a pattern. The description gives guidance on the note ('a note naming what emptied it') and implies the required parameters through context. With schema coverage at 67%, the description adds some semantic value but does not fully compensate for the missing project description. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Close an open handoff' when the work is done. It clearly distinguishes this tool from siblings like arroway_log (which records residue) and arroway_retire/withdraw (discarding is human-only). The scope and intent are unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to call it alongside arroway_log, and clarifies when it applies to someone else's handoff ('when your work emptied it'). It also states what is NOT this tool: a handoff still carrying real work must never be closed by a passer-by, and discarding is the human's gesture in the panel. These clear conditions route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_complete_readConfirm that an Arroway read is completeARead-onlyIdempotentInspect
After the final part of a multi-part Arroway read arrives, acknowledge it with its read_id and receipt. This closes the sequence and records that the whole snapshot was received. It is not a precondition for working: delivery is already recorded when the first part arrives.
| Name | Required | Description | Default |
|---|---|---|---|
| read_id | Yes | Opaque id returned by the incomplete arroway_read response. | |
| receipt | Yes | Opaque receipt returned by the final part. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds value by explaining that the tool closes the sequence and records receipt of the whole snapshot, and clarifies that delivery is recorded earlier. This goes beyond the structured fields and doesn't contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The main action is front-loaded, and the clarification about preconditions is valuable. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema and annotations covering safety, the description is sufficient. It explains when to call, what it does, and a key non-requirement. No critical information is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters described in the schema. The description mentions 'read_id and receipt' but doesn't add new semantic detail beyond connecting them to the process. Baseline 3 is appropriate given the schema already documents the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: acknowledge a multi-part Arroway read with read_id and receipt to close the sequence. It uses a specific verb and resource, and the context of 'final part' distinguishes it from other read-related tools, though it doesn't explicitly name sibling alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance ('After the final part... arrives') and an important exclusion ('It is not a precondition for working'), which helps an agent decide when to invoke this tool. It doesn't mention alternatives but the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_configure_referrerConfigure a known acquisition referrerAIdempotentInspect
Configure how a known external referrer is counted in Arroway's acquisition view. This is a product-wide operational setting, available only to a product administrator: use it when that administrator explicitly asks to map a domain to a source and channel. It stores only the domain and two short labels, never a visitor address or browsing history. Calling it again for the same domain replaces the mapping, so a new source can be recognised without a deploy.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | External hostname only, such as linkedin.com — no URL, path or query string. | |
| source | Yes | Short source label, such as linkedin or anthropic-directory. | |
| channel | Yes | Short channel label, such as organic-social, directory or referral. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the annotations: it stores only domain and labels, never visitor addresses or browsing history, and calling it again for the same domain replaces the existing mapping. This complements the idempotentHint annotation and warns about the overwrite behavior, which is important for an operational setting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three focused sentences, each earning its place: purpose, usage context and permissions, privacy, and overwrite behavior. It is front-loaded with the core action and avoids redundant restatement of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter configuration tool, the description covers the core purpose, the trigger condition, the required permission level, the stored data scope, and the overwrite behavior. No output schema exists, but the description sufficiently explains the effect of the call for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already well documented in the schema. The description reinforces that source and channel are short labels and clarifies that only domain and labels are stored, but it does not add substantial new parameter-level meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: configuring how a known external referrer is counted in Arroway's acquisition view. It clearly states the mapping from domain to source and channel, making the tool's purpose unmistakable and distinct from all listed siblings, none of which concern referrer configuration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use it when a product administrator asks to map a domain to a source and channel. It also adds relevant context by noting it is a product-wide operational setting with administrator-only availability, though it does not explicitly mention when not to use it or name alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_continueReceive the next part of an Arroway readARead-onlyInspect
Fetch more of an Arroway read that did not fit in one response. Send exactly the read_id and receipt from the preceding part; each call acknowledges that part and returns the next one. Continuing is OPTIONAL and deliberate: the first part carries the project's norms in full and already counts as a delivered read, and it lists by name what the later parts hold. Call this when something in that list bears on your task — not as a formality.
| Name | Required | Description | Default |
|---|---|---|---|
| read_id | Yes | Opaque id returned by the incomplete arroway_read response. | |
| receipt | Yes | Opaque receipt returned by the immediately preceding part. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and destructiveHint, and the description adds meaningful behavior: each call acknowledges the previous part, returns the next one, and the first part already counts as a delivered read. This goes beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences, front-loaded with purpose, followed by mechanics and usage criteria. Each sentence contributes; the optionality point is slightly repeated in the final warning, but the overall structure is tight and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter continuation tool with annotations and no output schema, the description explains the return behavior, acknowledgement semantics, and the decision context well. It could mention how the end of the read is signaled, but that is not essential for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and both parameters already have descriptive text in the schema. The description reinforces that read_id and receipt must come from the preceding part, but it does not materially add new semantic information beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Fetch more of an Arroway read that did not fit in one response.' It clearly distinguishes continuation from the initial read and sibling tools by explaining that each call acknowledges the preceding part and returns the next one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-call condition: only when something in the first part's list of later contents bears on the task, and it warns not to call as a formality. It does not name alternative sibling tools, but the relevant alternative—not continuing—is directly addressed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_exportExport this project's norms as a memory packARead-onlyIdempotentInspect
Export what this project has DECIDED as a portable memory pack — a document you can hand to another team, keep as a file, or import into another project. Only sanctioned, live memories travel: a proposal is not a norm yet, an archived one stopped being one, and a FINDING never was one — findings are context nobody sanctioned, so they stay in the project and the answer tells you how many did. Nothing about people travels — no author, no who sanctioned it, no row identifier, and no memory about WHO SOMEONE IS. The pack is what the team decided, never who was there. Give it a name that says what it is for, because that name is what the person on the other side sees before deciding whether to trust any of it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | What this pack IS, in a few words — the vertical, the practice, the playbook. It travels with the document and is the first thing the other side reads. | |
| topic | No | Optional: export only the memories carrying this topic. Use it to hand over one practice instead of everything a project knows. | |
| project | Yes | Slug of the project whose norms you are exporting | |
| description | No | Optional: who this pack is for and what it assumes. Written for someone who has never seen this project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral detail beyond the readOnlyHint and idempotentHint annotations: only sanctioned live memories travel, proposals/archived memories/findings are excluded, findings remain in the project and the answer reports their count, and no people-related data is exported. It also explains that the pack name is the trust signal for the recipient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded and the filtering rules are grouped logically. However, the point about people-related data not traveling is stated twice ('Nothing about people travels...' and 'The pack is what the team decided, never who was there'), making the text slightly more redundant than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description still explains what the result is (a portable document), what it contains, what is filtered out, and what the answer reports (how many findings stayed). Together with the fully described input schema, nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description's only parameter-specific comment—'Give it a name that says what it is for'—largely duplicates the schema's `name` description, so it adds no significant new semantic value beyond baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Export what this project has DECIDED as a portable memory pack.' It clearly distinguishes the tool's scope by defining what counts as a norm (sanctioned, live memories) and what does not, and contrasts with the sibling arroway_import by mentioning import as a downstream use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear use context: hand it to another team, keep it as a file, or import it into another project. It does not explicitly name sibling tools as alternatives or state when not to use it, so it falls short of a 5, but the intended scenarios are concrete and recognizable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_importImport a memory pack as proposalsAInspect
Bring a memory pack into a project. Everything in it lands as a PROPOSAL for the human to sanction in review — never as an active rule, no matter what the pack says about itself. A pack has the authority to ASK, never to decide: what arrives is a stranger's opinion until someone here says otherwise. Nothing is pinned on import either, because pinning spends this project's reading budget in every session from now on, and that is the owner's call with their own shelf in view. If any part of the document is malformed, NOTHING is imported and the refusal says how many units were left out — half an import is worse than none, because the memory that did not make it is the one nobody will go looking for.
| Name | Required | Description | Default |
|---|---|---|---|
| pack | Yes | The pack document itself, as it was written by arroway_export. Read the file and pass its contents — this server does not fetch URLs on your behalf. | |
| project | Yes | Project slug. If it did not come from the person or from a read in this session, it is a guess: check it against the roster that opens arroway_catch_up before writing, because material from one circle must never land in another. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses several behaviors beyond the annotations: everything lands as a proposal, never as an active rule; pinning is deliberately not done because it spends reading budget; malformed documents cause a total refusal with a count of omitted units. This is rich behavioral disclosure that goes far beyond the simple annotations (readOnlyHint false, destructiveHint false). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the main action, then explains the proposal status, pinning, and atomicity. Each sentence adds information, though it is fairly long. It is well-structured, not padded, and every sentence contributes to understanding the tool's behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main behavioral aspects, parameter guidance, and edge cases (malformed documents). It does not mention the return value, but no output schema exists, so the description could have addressed that. Overall it is fairly complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes both parameters with 100% coverage. The description adds critical context: the pack must be read and passed as an object (not a URL), and the project slug must be verified against the roster. This adds meaning beyond the schema and helps the agent avoid common mistakes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Bring a memory pack into a project.' It then clarifies the outcome (everything lands as a proposal, never an active rule) and distinguishes the tool from siblings like arroway_export by stating what it does not do. This is not a tautology and is immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when you have a memory pack to bring into a project. It sets expectations that the pack is never activated, so it is for importing as proposals. It does not explicitly name alternatives, but the purpose is clear enough for an agent to infer when to choose it. It also hints at verification of the project slug, which is useful guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_logLog what a task left behindAInspect
Append the durable residue of a COMPLETED task to your daily log in Arroway: what someone arriving later would need in order not to redo it. You judge the length — say what carries, cut what does not. Dated state (numbers, statuses) belongs here; it ages out naturally. Team-visible. ALWAYS write the authored essence too: it is the one line the normal project read serves. An older, frozen catalog that truly has no essence field may still write the body; Arroway marks that entry as pending a reviewable essence proposal, never as if one had been authored. One thing that belongs here and nowhere else: if the person corrected you for insisting on something already decided, say so in the entry, so the memory that keeps being re-litigated can be flagged. AND CLOSE WHAT THIS WORK REALIZED: every memory in the read carries 'dies when: …'. If this task made that condition true — the idea was built, the intention was carried out, the dated status was replaced — name it in fulfills and Arroway retires it in the same act, with this entry as the reason; the human never has to archive the obvious by hand. Only decisions, facts and references can be fulfilled, and only ones this connection actually received in a read.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | Optional free tag (e.g. 'marketing') | |
| source | No | Optional pointer to a living source (URL, doc, CRM id) | |
| content | Yes | The full durable residue: what was decided, found or changed, and what it cost. It remains recoverable by handle. | |
| essence | No | One authored line with the operative result. This is what the normal project read serves; never copy the opening of the body. | |
| project | Yes | Project slug. If it did not come from the person or from a read in this session, it is a guess: check it against the roster that opens arroway_catch_up before writing, because material from one circle must never land in another. | |
| fulfills | No | ARROW-290. The memories this completed work REALIZED — their 'dies when' condition is now true because of what you did (a routine that an idea asked for is running; a planned move was made; a dated status has been replaced by what this entry records). Each one is retired in the same act as this log, with your 'how' as the reason on its archived line, reversible from the panel. Guards: only decision/fact/reference (a rule is never 'fulfilled'), and only memories this connection received in full in a read — if any item fails, the whole call is refused and nothing is written. Never list a memory because it merely looks old or unused: that is not evidence. | |
| backfilled | No | True when the entry describes something that happened BEFORE Arroway existed (history you are bringing in), not today's work. | |
| keep_whole | No | Set true only after you considered cutting an indivisible oversized residue. Arroway records this exception for calibration; it never truncates your text automatically. | |
| operational | No | Set true when this is the record of HOW THE WORK WENT — what failed, what you measured, what you corrected about your own earlier reading. It is kept in full for catching up at the start of a session, and stays OUT of the task-scoped read, so a long account never pushes a teammate's residue out of their block. Leave it out (the default) for what the team needs to know HAPPENED: that is what reaches other people while they work. This is a record of work, never a transcript of the conversation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations: it discloses that fulfilled memories are retired in the same act, that the retirement is reversible 'from the panel', that entries without an essence are marked as pending a reviewable proposal, and that the call is refused if any fulfills item fails. These are meaningful side-effect behaviors. No contradiction with annotations exists; readOnlyHint=false and destructiveHint=false are consistent with an append-plus-retire operation that is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and long, but almost every sentence adds operational guidance. It is front-loaded with the primary purpose and then layers correctness rules. Some repetition and capitalization make it harder to scan, but the length is largely justified given the tool's complexity and the important side-effect behavior that must not be missed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter write tool with no output schema and partially informative annotations, the description is thorough. It explains the write obligation, the essence rule, the fulfills/retirement behavior, and the exception for frozen catalogs. The remaining gaps — such as explicit guidance for operational, backfilled, topic, and source — are adequately covered in the input schema's own descriptions, so the description is not incomplete enough to cause misinvocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds substantial guidance beyond the schema. It explains what should go into content ('say what carries, cut what does not'), that essence is mandatory for normal entries, which dated states belong in the log, and what 'fulfills' means in terms of killing conditions. This materially improves an agent's ability to populate the most nuanced fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Append the durable residue of a COMPLETED task to your daily log in Arroway'. It clarifies the tool is for durable, team-visible task residue and explicitly distinguishes it from other log/read tools by emphasizing this is what remains after completion. This is specific enough to differentiate it from sibling tools like arroway_remember or arroway_pass.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use it for completed task residue, not for conversation transcripts or open work. It also says what belongs here ('Dated state', 'the person corrected you...') and what does not, and explains the fulfills mechanism. It does not explicitly name alternative tools or give explicit 'when not to use' exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_movePropose moving a memory to another projectAIdempotentInspect
PROPOSE that a memory belongs in a different project — you never move it yourself. Moving changes WHO SEES IT: leaving a team project removes it from the team, and going from a personal project to a shared one exposes it. So this only queues the change for the human's review. Use the #handle from the read.
| Name | Required | Description | Default |
|---|---|---|---|
| why | Yes | Why it belongs there and not here — the human decides on this sentence | |
| memory | Yes | #handle of the memory, from a read | |
| target | Yes | Slug of the project you believe it belongs to | |
| project | Yes | Slug of the project the memory is in TODAY |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond annotations by explaining the queuing behavior ('only queues the change for the human's review') and the visibility consequences of a move (team vs personal exposure). Annotations already cover idempotency and non-destructiveness, so the description adds useful context without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences front-load the core purpose, then behavior, then parameter source. Each sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage context, behavioral consequences, and parameter source. It doesn't describe the return value, but there is no output schema, and the human-review outcome is explicitly stated. Some missing details like failure conditions are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all four parameters at 100% with clear descriptions. The description adds the source for memory ('Use the #handle from the read') but that is already present in the schema. No additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('PROPOSE'), resource ('a memory belongs in a different project'), and distinguishes itself from actually moving by saying 'you never move it yourself' and 'only queues the change for the human's review.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (when you believe a memory belongs elsewhere) and instructs to use the #handle from the read. It stops short of naming alternative sibling tools, but 'you never move it yourself' establishes that this is the proposal path, not the execution path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_normsWhat the team already decided, and a map of what else it knowsARead-onlyIdempotentInspect
Call this in the seconds BEFORE you tell the person something or propose a course of action, to see whether it contradicts a decision the team already made. It returns the standing norms — the project's pinned rules and decisions, plus this person's own standing rules — AND THEN A MAP: the titles and handles of everything else the project remembers, so you can see whether anything there touches what you are doing. Read that map: if a title looks relevant, ask for it by handle with arroway_read include:["#handle"], which GUARANTEES those memories come back in full and first, where the ranking could otherwise have left them out — it does not make that read any cheaper than a read without it, so name a handle to be sure of getting something, never to pay less for it. It is meant to be called often: many times in one session, whenever you are about to commit to a claim. The map gives you names, never content, so this still does not answer 'how do I do this' — no recent context, no daily log, no ranked material. When you are starting work on a task, that is arroway_read, and this does not replace it. Because it is meant to be repeated, it ends by printing a session checkpoint: pass it back as since on your next call for this project and the standing norms are referenced instead of reprinted, so the repetition costs the map and what is in flight rather than the whole fixed block again.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Optional: the session checkpoint printed at the end of a previous arroway_norms call for THIS project. Presenting it lets the standing norms be referenced instead of reprinted, which is what makes calling this tool many times in one session cheap. It is this tool's own checkpoint — the one arroway_read prints is not interchangeable with it, because the two surfaces render the same memories at different fidelity. Omit it — or present one this connection does not hold — and the norms are written out in full; the server decides, never your local state. If your context was compacted or this is a new conversation, omit it. | |
| project | Yes | Project slug the claim is about. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds substantial context: it returns a map of titles/handles rather than content, it does not answer 'how do I do this', it prints a session checkpoint that can be passed back, and the server decides whether to reprint norms. These behavioral details go well beyond the annotations and clarify the tool's side effects and repetition semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured: it front-loads the primary use, then explains the map, the `since` mechanism, and the distinction from arroway_read. Every sentence serves a purpose, though some could be tightened. It earns a 4 because the length is justified by the tool's complexity and the need to convey nuanced behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return values, which it does thoroughly: standing norms, a map of titles/handles, and a session checkpoint. It also covers the cost implications of passing `since` and the relationship with arroway_read. An agent has all the information needed to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, yet the description enriches both parameters: it explains `since` as the checkpoint from a previous call and clarifies that omitting it forces full output, while `project` is tied to the slug of the claim. This adds operational meaning beyond the schema's definitions, making the parameters actionable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action—checking team decisions and norms before committing to a claim—and differentiates itself from arroway_read, which is for starting work. It clearly identifies the resource (project norms and memory map) and the intended trigger ('in the seconds BEFORE you tell the person something'). This distinguishes it from siblings and leaves no ambiguity about its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use (before telling someone something or proposing a course of action), how often ('many times in one session'), and when not to use it ('When you are starting work on a task, that is arroway_read, and this does not replace it'). It also explains the `since` parameter's role in making repeated calls cheap, covering both usage frequency and cost optimization.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_passPass unfinished work forwardAInspect
Call this when you STOP with the task unfinished — out of time, blocked, or told to stop. It leaves a HANDOFF: the in-flight state of open work, served at the top of every read of this project until someone takes it and closes it. It does not replace the log: a COMPLETED task still ends in arroway_log; this exists for the one you could not complete. ALWAYS write the authored essence too: it is the default in-flight state served in reads. An older, frozen catalog that truly has no essence field may still pass the full state; Arroway marks it as pending a reviewable essence proposal, never as if one had been authored. Deliberate only — never a dump of the session: write exactly what the next session needs to continue without redoing or re-deriving anything. A handoff never expires by clock: it dies in exactly three ways — closed with arroway_close, replaced by another pass naming it as superseded, or discarded by a person in the panel. Age is reported in the read as information; it never removes anything.
| Name | Required | Description | Default |
|---|---|---|---|
| card | No | Pointer to the living tracker for this work (card id, issue, PR), when one exists. | |
| essence | No | One authored line that states the operative in-flight state. This is the default read; the full handoff remains recoverable by handle. | |
| project | Yes | ||
| next_step | Yes | The single concrete next step. Not a plan: the step. | |
| open_risk | No | The risk left open — what could bite whoever continues. | |
| supersedes | No | #handle of the OPEN handoff this one replaces, from a read. Replacement is explicit by reference, never guessed from titles. | |
| do_not_redo | No | What is already done, verified or decided that the next session must NOT redo or re-litigate. | |
| addressed_to | No | Optional name of a person or routine this is meant for. A visible convention, never a lock: everyone still sees it, and the read says who may take it. | |
| read_handles | No | The reading list the next session needs: #handles of the memories, log entries and handoffs — of THIS project — that whoever continues must read in full. You know what you had to read to get here; naming it spares them rediscovering it. Order is kept as you write it. At most 10 unique references, and going over is refused rather than trimmed, because each one comes back in full body. Only what is genuinely required to continue: a handle merely cited in do_not_redo is history, not required reading, and is not picked up from the text — it counts only if you name it here. | |
| where_stopped | Yes | Where the work stands right now — what is done and verified, what is half-done. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals significant non-obvious behavior beyond the annotations: the handoff is served at the top of every read until closed, it never expires by clock, it dies in exactly three ways, and age is only informational and never auto-removes. It also mandates writing the essence and warns against session dumps. This adds real behavioral context beyond the readOnly/destructive hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the trigger condition and is densely informative. It is longer than minimal, but almost every sentence carries a distinct operational rule or constraint needed for correct use. A little tightening is possible, but the length is largely justified by the tool's role as a handoff mechanism with a precise lifecycle.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 parameters, no output schema, and nuanced lifecycle semantics, the description is remarkably complete. It covers when to call, what the handoff does, how it differs from the log, what the essence requirement is, how handoffs end, and how age is treated. It also gives per-field guidance through the schema such that an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 90%, the structured schema already documents most parameters well. The description adds higher-level operational semantics, such as essence being the default read, supersedes needing explicit reference, read_handles being required reading (with order and refusal behavior), and do_not_redo preventing wasted rework. It does not redefine parameters but enriches how to choose their values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific invocation condition ('Call this when you STOP with the task unfinished') and clearly identifies the resource as a handoff of in-flight work. It distinguishes itself from arroway_log by explicitly stating that completed tasks go to the log, so an agent can immediately tell this tool apart from its closest sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is explicitly conditional: use it when out of time, blocked, or told to stop, and not for completed tasks, which belong in arroway_log. It also describes the handoff lifecycle—closed by arroway_close, replaced by a superseding pass, or discarded by a person—which helps the agent know when this tool is appropriate versus those siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_readRead ArrowayARead-onlyIdempotentInspect
Read Arroway for a project BEFORE acting on any task related to it. It serves the authored essence of memories, dated-log entries and handoffs by default; name a #handle in include to recover that object's full body.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Optional: the session checkpoint printed at the end of a previous read of THIS project. Presenting it lets the standing rules be referenced instead of reprinted, freeing that space for the task. Omit it — or present one this connection does not hold — and they are written out in full; the server decides, never your local state. If your context was compacted or this is a new conversation, omit it. | |
| handoff | No | Optional: the #handle of an OPEN handoff you are continuing. The read is then built from that handoff's reading list (the read_handles its author named): the standing rules in full, the handoff and every reference it names in full, and nothing ranked — much smaller than a normal read, and it counts as a delivered read. A handoff with no reading list falls back to the normal read, and says so. Cannot be combined with include; `since` is ignored. If a reference on the list is no longer available, the read is refused rather than served incomplete — read normally instead. | |
| include | No | #handles of memories, dated-log entries or handoffs whose full body you need. They are placed first and cannot be dropped by the budget — so other memories fall out instead: this chooses what you get, it does not get you more. | |
| project | Yes | Project slug, e.g. 'seven50' | |
| task_context | No | Optional: what you are about to do. It ranks what comes back, but the declaration alone is NOT evidence that a memory was used. A matching completion residue from the same session must corroborate it before Arroway updates the memory's observed-use signal. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, idempotentHint, and destructiveHint, so safety is covered. The description adds behavioral detail: include places handles first and cannot be dropped by budget, handoff builds a smaller read, and since's checkpoint behavior is noted. It does not describe the return format, which is a gap given no output schema, but the added context is valuable and consistent with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. It front-loads the primary purpose and then provides a compact conditional for full-body recovery. Every clause contributes to understanding, and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main use cases (default read, include, handoff, since) but omits what the return payload looks like. With no output schema, the agent cannot predict the structure of the response, which is a significant gap for a read tool. It also doesn't define 'Arroway' or the concept of 'authored essence', though these may be assumed from the tool family.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The description adds meaning beyond the schema by explaining the budget implications of include and the special behavior of handoff, which are not fully conveyed in the parameter descriptions. This supplements the schema without redundancy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads Arroway for a project, specifying the verb 'read' and the resource 'Arroway'. It also distinguishes the default behavior (serving authored essence) from the include option to recover full bodies, and the 'BEFORE acting' instruction sets it apart from sibling tools that perform actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use this tool before acting on any task related to a project, which is a strong when-to-use cue. It also explains how to request full bodies via include and how to continue a handoff, but it does not explicitly mention alternatives like arroway_search_log for searching, leaving some room for interpretation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_rememberPropose a durable memoryAInspect
Save a durable memory (decision/rule/fact/preference/reference/identity) as governed by default. For a source-backed fact or reference that is useful only as agent context, set treatment='finding' and explain why: it costs nobody a review, and is stored as CONTEXT, never a human-sanctioned norm, cannot be pinned, and is shown separately in Findings. A finding cannot replace a sanctioned memory; correcting another finding is in-place replacement with no human queue. A finding a person revoked cannot be written again unless they reconsidered in this conversation — related_checked does not unlock that. 'identity' is WHO someone is — a person, a relation, a background: it never expires, never pins by default, and comes back when the person is part of the task; someone from the user's own circle belongs in their personal project, people of a business in that business's project. If the human explicitly stated or sanctioned it in this conversation, set decided_by_human=true (memory becomes active). Otherwise it is saved as a PROPOSAL for the human's weekly review — never present a proposal as a decision. kill_condition is mandatory: what would kill or force a review of this memory. A RULE takes TWO calls: send it with no pin fields and no related_checked, read the neighbours and the block cost the server returns, then call again with the token it gave you and the pin decided against what you just saw. Any OTHER type stays one call unless you ask to pin it: sending pin_suggested or pin_requested_by_human starts the same comparison, and only the second call with the returned token can pin it. Any type that the server stops for strong related memories also takes TWO calls: read those memories, then return related_checked=true together with its related_check_token. A flag alone never proves a comparison happened.
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | ||
| title | Yes | Short, unique within the project | |
| topic | No | ||
| content | Yes | The norm ITSELF, written as if it were true today and nobody had to be told how it got here. Length is never free: this body is re-read by every assistant on this project, on every read, from now on — write the shortest version that still governs correctly. Keep only the reasoning WITHOUT WHICH the rule would be applied wrongly; if removing a sentence would not change anyone's next action, it is biography, not norm. Leave out the biography: who said it, on what date, their exact words, the episode that produced it, and numbers measured on a particular day — all of that goes in why_source. A body that opens with a quote makes the next assistant follow the quote instead of the rule, and a date in the body makes it discount the whole memory as possibly stale. If a GENERAL rule shows up while you are writing about a specific case, the general rule is the memory and the case is the example — not the other way round: the case goes in the example field, never here. Condense: a memory per sentence is how Arroway fills with near-duplicates. | |
| essence | Yes | The memory's essence in ONE line: the operative norm or fact, not the first sentence and not its biography. This is the compressed layer that always travels when the full body does not fit. | |
| example | No | A concrete case that ILLUSTRATES the rule — the human sees it in the panel next to the rule; it NEVER travels in reads, so no assistant can mistake the case for the rule's scope. If the rule cannot be applied correctly without this example, the example is smuggling a criterion: name that criterion in content instead. Quotes, dates and the episode that produced the rule still belong in why_source. | |
| expires | No | ARROW-78. Set true ONLY for a dated STATUS — a number, a rollout state, a temporary condition — where being read after review_at would assert something stale as current: it stops being returned once the date passes. Leave it out for a durable fact that merely needs re-checking (a market structure, how a tool works): that one keeps being returned, marked as due for re-check. If unsure, leave it out — a fact that vanishes leaves no trace, one that comes back marked cannot be mistaken for current. | |
| project | Yes | Project slug. If it did not come from the person or from a read in this session, it is a guess: check it against the roster that opens arroway_catch_up before writing, because material from one circle must never land in another. | |
| review_at | No | ISO date (YYYY-MM-DD) when this stops being true on its own — the end of a cycle, a deadline, a season. This is how a COMMITMENT is carried: an objective is a decision with a horizon. REQUIRED for type=fact, and there it also removes the memory from reads once the date passes. Leave it empty only for something with no expiry date at all. | |
| treatment | No | Which of two destinations this write takes, and they cost different people. `governed` enters the person's review queue, and that review is a real cost to them — pay it for what governs. `finding` costs nobody a review: it is served to agents as context, never becomes a sanctioned norm, cannot be pinned and cannot replace one. The line between them is not the topic, it is the force: does this GOVERN what someone may or must do, or only INFORM? Governs → governed. Informs, and points at a source another agent can reopen and compare — a file, a URL, a record, a query → finding, and only fact and reference qualify. | |
| backfilled | No | True when this is HISTORY you are bringing in — something that was already true before Arroway existed, not something that just happened. The moment to use it: you finished a task about topic X, the read came back thin on X, and the work took you to a durable SOURCE about X — a file in this workspace, a document, a ticket. Then write that history too. Only the topic you just worked on, never a project dump; only what you can point a source at, never your own recollection of past conversations; a few memories, not a batch. | |
| source_ptr | No | Pointer to the living source, when verifiable | |
| why_source | No | Where it came from, and this is the ONLY place biography belongs: who decided, when, in their words, what was measured, which episode produced it. This does NOT travel in reads — it is what a person sees when auditing the memory in the panel, and what tells them whether the rule still deserves to exist. Writing it here costs nothing to whoever reads Arroway while working; writing it in the body costs every assistant, every read. | |
| contradiction | No | Required with supersedes_id: name what the earlier memory says that this one contradicts or changes. This is the named contrast that proves it is the same point, not merely a neighbour on the same topic. Do not point at a memory that can be followed alongside this one. | |
| pin_suggested | No | Whether this memory sits in the fixed block of EVERY read, costing tokens in every session of this project from now on. There is no default: a rule requires this on its confirming call, and any other type carrying this field begins the same two-call comparison. The first call does not decide the pin; return this value only with pin_decision_token after the server showed the neighbours and cost. Three questions decide: does breaking this cause expensive or irreversible damage · does it apply to every task, or only when someone touches one area · would the next person have found it anyway when they opened the relevant file? A memory discoverable where it matters is a reference or a fact, not a pinned rule. What a pin actually costs: it does not displace another pin — it eats the SAMPLED space of every read, which is where the memories relevant to someone else's task come from. This field carries YOUR judgement, so above the block's ceiling the server turns it into a proposal of the pin alone and the memory itself still lands. The person's own order is pin_requested_by_human, and it is not this field. | |
| supersedes_id | No | If this replaces an earlier memory: its #handle from a read. Point only when the two memories are the SAME point: acting on either would already satisfy or violate the other. A shared topic is not enough — if both can be followed together, they are different memories. | |
| kill_condition | Yes | Mandatory: explicit revocation / named trigger / review date | |
| related_checked | No | Only after the server showed you related memories: set true with related_check_token to confirm you read them and this is a genuinely different point, not an update to one of them. Never set either blindly on a first call. | |
| decided_by_human | Yes | true ONLY when the person stated or sanctioned this in THIS conversation. Never because it looks obviously right to you. false makes it a PROPOSAL for their review, which is the normal case — and a proposal is never free: it costs that person a review, whether they end up approving it, editing it or turning it down. What it costs afterwards depends on where it sits: only a PINNED memory travels in every read of the project from now on; a normal memory competes for space in the block relevant to the task at hand. So the question that decides whether to write at all is not what the write costs you — it is whether this would change what someone does on a DIFFERENT task. Three things never pass that question, however true they are: an instruction to check, verify, be careful or look at the source; a step of a routine, which belongs in the prompt that runs it; and anything that is here only because something went wrong once. | |
| split_considered | No | Set true only after you looked for the seams and this genuinely has to stay whole. Splitting is the default: a long memory usually holds a norm, the reasoning behind it, and a checklist — and only the first belongs here. | |
| example_considered | No | Only after the server flagged the body as carrying its example: set true to confirm you re-read it and what looks like an example there is actually criterion the rule needs — an exact string it matches on, a deadline it carries. Never set it blindly on a first call. | |
| pin_decision_token | No | The token the server gave you on the first call for a pin decision. A rule always uses this two-call flow; every other type uses it whenever pin_suggested or pin_requested_by_human is present. The first call returns the neighbouring memories with the pin state of each and what the fixed block currently costs, and only the second one writes. Send the token back unchanged on that second call, together with the pin decision. It is tied to the exact title and body you compared, works once, and expires — a comparison you cannot prove is a comparison that did not happen, so a confirmation without it is refused and nothing is written. | |
| related_check_token | No | The one-use token the server returned with related memories. It proves this confirmation follows THAT comparison and is tied to this exact text. | |
| pin_requested_by_human | No | true ONLY when the person, in THIS conversation, told you to fix this memory in every session — 'pin this', 'this has to be in every session', 'always send this'. On any type, sending it begins the two-call comparison; only the return with pin_decision_token records the order. Sanctioning the CONTENT is not asking for the pin: they can agree a memory is right, and correct, and still never have asked for it to travel in every read. Those are two different facts and the server records them separately. Their order pins immediately, above any ceiling, and the block's hygiene is then theirs; your own judgement goes in pin_suggested and passes through the ceiling. Never set this because the pin looks obviously right to you — that is pin_suggested. | |
| revoked_reconsideration | No | Required when a related finding was revoked by a person: quote what they said in THIS conversation that reconsiders it. related_checked does not unlock rewriting a revoked finding. Do not promise that old prompts or past actions will be undone. | |
| treatment_justification | No | Required with treatment=finding: say why this is verifiable context rather than a policy, priority, procedure, or decision for a person. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses non-obvious side effects and server behavior well beyond the minimal annotations: findings are stored as context and never become human-sanctioned norms, revoked findings cannot be rewritten without reconsideration, pinning requires a two-call token comparison, and the server can convert a pin suggestion into a proposal. This gives an agent a realistic model of what the write does and what it costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but nearly every sentence carries a policy constraint or decision rule relevant to correct invocation, and the core purpose is front-loaded. Its main weakness is structural: it is a single wall of text without bullets or section breaks, making it comprehensive rather than elegantly scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with the 92%-covered schema and the annotations, the description covers the mandatory kill_condition, active-vs-proposal semantics, finding restrictions, identity routing, and the complete two-call flows for rules and pinning. Nothing essential to making the first or confirmation call correctly appears to be missing, even without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers 92% of parameters with rich descriptions, so the baseline is satisfied. The description adds orchestration semantics not present in individual field docs: the two-call rule protocol, the pin_suggested/pin_requested_by_human comparison with pin_decision_token, the related_checked/related_check_token requirement, and the treatment='finding' conditions. It does not enumerate all 26 parameters, but the schema compensates for that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence is a specific verb-plus-resource statement: 'Save a durable memory' with the enumerated types. It immediately distinguishes governed vs finding treatment and special identity semantics, so an agent can understand the operation and its variants without relying on the title alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear contextual rules for key branch decisions: when to use treatment='finding' vs governed, when decided_by_human should be true vs false, and where identity memories should be placed. It does not explicitly name sibling tools like arroway_log or arroway_claim as alternatives, so the guidance is internal to this tool rather than cross-tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_retireRetire a sanctioned memory, on the person's instructionADestructiveInspect
Retire a memory the human already sanctioned, when THEY told you in this conversation to remove it and there is nothing to put in its place. It is archived, never deleted: it leaves reads and stays in the history with their reason, or with an explicit note that no reason was declared. Use arroway_remember with supersedes_id instead when you have a corrected version — that is replacement, not retirement. Use arroway_withdraw instead for a proposal YOU wrote that is still pending.
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Optional: what the person said made it obsolete, in their terms. It stays on the archived line forever; if omitted, the history explicitly says no reason was declared. | |
| memory | Yes | The memory to retire — the #handle from a read is enough. | |
| project | Yes | Project slug. | |
| decided_by_human | Yes | Must be true, and only when the person explicitly told you to remove it IN THIS CONVERSATION. Never infer it, never set it because it looks obsolete to you: a sanctioned memory is theirs. Without their word, say it is a gesture for them to make in the panel. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description goes further to clarify that the memory is 'archived, never deleted' and that reads and history are affected, including how the reason is preserved. This adds meaningful nuance beyond the annotation without contradicting it. It does not fully describe all side effects (e.g., impact on dependent items), so 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states the core purpose and condition, the second describes the archival behavior, and the third routes to two alternatives. The critical condition is front-loaded, and there is zero redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderately complex tool with 4 parameters and no output schema, the description covers the essential decision points: when to use, when not to use, the archival effect, and the required human instruction. It does not mention error conditions or return values, but given the schema richness and the description's clarity, it is sufficiently complete for an agent to invoke correctly. A 5 would require explicit output expectations or error handling, which are not present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter has a detailed description (e.g., decided_by_human explains the 'must be true' constraint and the prohibition on inferring). The main tool description adds no extra parameter-specific meaning beyond the schema, so the baseline 3 is appropriate. The description's mention of 'nothing to put in its place' indirectly clarifies the why, but that is already implied in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('retire'), the resource ('a memory the human already sanctioned'), and the precise condition ('when THEY told you in this conversation to remove it'). It clearly distinguishes from siblings by naming arroway_remember (replacement) and arroway_withdraw (pending proposal), so an agent can disambiguate without reading other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the when-to-use condition (sanctioned memory, human instruction in this conversation, nothing to replace it) and the when-not-to-use alternatives with the exact tools to call instead (arroway_remember with supersedes_id for corrected versions, arroway_withdraw for pending proposals). This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_search_logFind where something appeared in the dated logARead-onlyIdempotentInspect
Search the recent DATED LOG for a literal term or phrase when you need to know WHERE something was said. It matches the authored one-line essence as well as the body, so a fixed term written as the entry's state line is found. By default it scans every project this person belongs to; pass project to narrow it. Returns compact excerpts anchored by project, date and entry — never full days, and each excerpt says when it came from the essence. This searches what happened, not curated memory, and does NOT replace arroway_read before acting on a task.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Literal term or phrase to find in recent dated-log entries. Matching is case-insensitive. | |
| project | No | Optional project slug. Omit it to search every project this person can access. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds meaningful behavioral detail beyond annotations: matching covers the one-line essence and body, results are compact excerpts anchored by project/date/entry, full days are never returned, and the default scope is all accessible projects. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then efficiently covers matching behavior, scope, output shape, and a critical caveat. Every sentence earns its place; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with no output schema, the description provides enough for an agent to call it correctly: what is searched, how matches work, how to narrow scope, and the nature of returned excerpts. It does not mention result limits or ordering, a minor gap given the description already explains the essential behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters have solid descriptions. The description adds extra semantic value by clarifying the default behavior when 'project' is omitted, that 'query' is a literal case-insensitive term, and that matching applies to both essence and body. This goes beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search') and resource ('recent DATED LOG') and the exact need ('WHERE something was said'). It also distinguishes itself from curated memory and explicitly names arroway_read as a tool it does not replace, making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use context ('when you need to know WHERE something was said'), explains the default scope and how to narrow it with 'project', and warns that it should not replace arroway_read before acting. It does not systematically enumerate sibling alternatives, but the guidance is sufficient for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_start_projectOpen a new projectAInspect
Open a NEW project in Arroway, when the work at hand has no home. Use it when a read refuses an unknown project, or when you notice you have been working on something for a while with nowhere to write it. Propose the name and scope to the human in the conversation FIRST and only call this once they agree — then say it exists and can be renamed or killed in the panel. Scope is mandatory: it is the sentence that stops one circle's material from being written into another.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | How people will call it | |
| slug | Yes | Short, lowercase, hyphenated, unique — e.g. 'trading-solana' | |
| scope | Yes | What belongs here and what does not. Written for the next assistant to read before writing — say the subject AND name what is out. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations (readOnlyHint=false, destructiveHint=false) are minimal, so the description carries the burden. It discloses that the tool creates a new project, that the name and scope must be proposed first, that scope is mandatory to prevent cross-contamination, and that the project can later be renamed or deleted. This adds valuable behavioral context beyond the annotations, though it doesn't describe potential side effects beyond creation (e.g., file system changes) — a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph but well-structured: it front-loads the core purpose, then gives usage conditions, then step-by-step instructions. It is a bit verbose, but every sentence adds value—no fluff. The key information (when to use, what to do first) is prominent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 required parameters, no output schema, and minimal annotations, the description covers the essential aspects: what it does, when to use it, how to proceed (propose first), and the outcome (project created, can be renamed/killed). It doesn't explain return values, but that's not expected without an output schema. It's sufficiently complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides descriptions for all three parameters (100% coverage), so the baseline is 3. However, the description enriches the meaning, especially for scope: 'Scope is mandatory: it is the sentence that stops one circle's material from being written into another.' It also clarifies that name is 'How people will call it' and instructs the agent to propose these to the human, adding practical semantic context beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Open a NEW project in Arroway, when the work at hand has no home.' It uses a specific verb and resource, and provides context that distinguishes it from siblings by emphasizing it's for creating new projects when no existing one fits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use conditions: 'Use it when a read refuses an unknown project, or when you notice you have been working on something for a while with nowhere to write it.' It also instructs the agent to propose the name and scope to the human first and only call after agreement, and even tells the agent to say the project can be renamed or killed. This is clear, actionable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_verifyRecord what the source saysADestructiveIdempotentInspect
Record what you found when your work ALREADY took you to a memory's source. This is bookkeeping, not an errand: never go verifying memories as a task — measured over a month, that never happens. But when you open a file, query, ticket or CRM that a memory points at, and you can see whether it still holds, say so here. Use the #handle shown in the read. Never guess an outcome you did not actually see. What the evidence decides, Arroway closes by itself (ARROW-290): a FACT whose source is proven gone is retired on the spot; a FACT whose source says something different is retired the moment you propose the correction with arroway_remember supersedes_id in this same session. On a FINDING, source_gone archives it the same way; conflicts is corrected in place with treatment=finding and supersedes_id, with no human queue. Rules and decisions are never retired this way — they only get marked.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Required for 'conflicts' (what the source actually says — it is what lets a human fix it without checking again) and for 'source_gone' (what you saw that proves removal — a note describing a timeout, an auth failure or a service being down is refused). | |
| memory | Yes | The #handle from the read (8 chars), or the full id | |
| outcome | Yes | matches = the source says what the memory says · source_gone = the source is PROVEN REMOVED (404, file deleted from the repository, record gone) — never 'unreachable': a timeout, a login wall or a service that is down is unavailability, and nothing is retired for it · conflicts = the source says something DIFFERENT (then also call arroway_remember with supersedes_id; on a fact that proposal retires the wrong fact at once; on a finding the correction replaces it in place, with no human queue) | |
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though annotations already declare destructiveHint=true and idempotentHint=true, the description adds rich behavioral detail: facts are retired automatically when sources are gone, findings are archived or corrected in place, and notes describing timeouts or auth failures are refused. This goes well beyond the structured annotations and clarifies real side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core 'bookkeeping, not an errand' rule, and every sentence adds a meaningful constraint or edge case. It is somewhat long and contains run-on sentences, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the conditional required-note schema, the destructive potential, and the absence of an output schema, this description covers everything an agent needs: when to invoke, what outcomes mean, what to record in notes, and what automatic consequences follow. The only minor omission is explaining the project parameter, which is likely shared tool context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents memory, outcome, and note well (75% coverage). The description adds meaning by requiring the #handle from the read, forbidding guessed outcomes, and specifying that source_gone notes must describe actual proof of removal, not mere unavailability. Project is not mentioned, but the schema coverage and overall context make this a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a precise verb-resource pair: 'Record what you found when your work ALREADY took you to a memory's source.' It also differentiates itself from siblings by drawing a line against 'verifying memories as a task' and pointing to related tools like arroway_remember and arroway_retire.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when/when-not guidance: only use when you are already at the source, never go verify as a standalone errand, use the read's #handle, and never guess an outcome. It also names the exact follow-up actions for conflicts and source_gone, including when to call arroway_remember with supersedes_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arroway_withdrawTake back your own pending proposalAIdempotentInspect
Take back a pending PROPOSAL that turned out to be wrong — typically because the human corrected you after it was written. Pending proposals only; a memory the human already sanctioned is theirs, and changing it goes through them. A proposal written by SOMEONE ELSE's AI can also be taken back — the archive belongs to the project — but Arroway stops first, names who wrote it and what would leave their review queue, and the next call with other_author_confirmed=true finishes it. This is not deletion: the proposal is kept, marked as withdrawn, with your reason. If the corrected version is durable, write it with arroway_remember too — leaving the right answer only in your own notes is how Arroway ends up holding the wrong one.
| Name | Required | Description | Default |
|---|---|---|---|
| why | Yes | What made it wrong — usually what the human said. This is the part worth keeping. | |
| memory | Yes | #handle of the proposal, from a read | |
| project | Yes | ||
| other_author_confirmed | No | Only after Arroway told you the proposal was written by someone else's AI: set true to go ahead. Never set it blindly on a first call — the point of the stop is that the person hears whose work is leaving the queue before it leaves. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations: it discloses the non-destructive nature ('kept, marked as withdrawn'), the two-call confirmation flow for another author's proposal, and the consequence of leaving corrections only in personal notes. This is consistent with destructiveHint=false and idempotentHint=true, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries essential decision-relevant information: scope, exclusions, cross-author behavior, non-deletion semantics, and the follow-up tool. It is front-loaded with the core purpose and then layers edge cases in a readable order with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters, no output schema, and a subtle non-destructive withdrawal flow, the description covers all important context: when withdrawal is allowed, how the other-author stop works, what happens to the proposal, and what to do with the corrected version. An agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so the schema already documents why, memory, and other_author_confirmed in detail. The description reinforces the flow around other_author_confirmed but adds little parameter-level meaning beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: take back a pending PROPOSAL. It sharply distinguishes this from deletion ('This is not deletion: the proposal is kept, marked as withdrawn') and from other tools like arroway_remember, so an agent can tell exactly what arroway_withdraw is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use it ('pending PROPOSAL that turned out to be wrong'), when not to use it ('a memory the human already sanctioned is theirs, and changing it goes through them'), and what to do instead for durable corrections ('write it with arroway_remember too'). This gives an agent both positive and negative selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
- First observed
arroway_catch_up - First observed
arroway_claim - First observed
arroway_close - First observed
arroway_complete_read - First observed
arroway_configure_referrer - First observed
arroway_continue - First observed
arroway_export - First observed
arroway_import - First observed
arroway_log - First observed
arroway_move - First observed
arroway_norms - First observed
arroway_pass - First observed
arroway_read - First observed
arroway_remember - First observed
arroway_retire - First observed
arroway_search_log - First observed
arroway_start_project - First observed
arroway_verify - First observed
arroway_withdraw
Publisher details
- Operator
- Arroway · Publisher source
- Operator website
- https://www.arroway.app · Publisher source
- Vendor relationship
- First-party
- Documentation
- https://www.arroway.app/en/install · Publisher source
- Trust center
- Not available
- Restrictions
- Requires an Arroway account. The first connection opens sign-in (Google or an email link) and OAuth authorization; there is nothing to paste. · Publisher source
Related MCP Connectors
Shared memory for AI agents. One address per fact, one signature you check. No key to read.
- ContexelOAuthai.contexel
Shared AI memory. ChatGPT, Claude, Cursor and any MCP app read and write the same memory.
- HeirmosOAuthcom.heirmos
Persistent memory shared across Claude, ChatGPT, Grok and other MCP clients.
Shared memory for AI tools — save notes and work records, recall them from any other tool.
Related MCP Servers
- AlicenseAqualityAmaintenanceShared memory and handoff hub for AI agents, enabling seamless context transfer between sessions with token-budgeted resumes and automatic handoffs.1010 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI subagents to share a persistent, versioned memory across sessions and machines, with tools to add, recall, inspect, merge, resolve conflicts, and report usage for rewards.MIT
- AlicenseNot gradedqualityDmaintenanceShared memory hub for LLMs to persist and share project context, enabling seamless handoffs between different AI agents.9 npm1MIT
- AlicenseNot gradedqualityBmaintenanceGoverned, self-hosted memory for AI agents: writes queue until an authorized approver signs off.Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.