Metis
OfficialMetis is an open-source toolkit for capturing fragments of expert practice and making them available to AI agents as memory, with human review and agreed conditions for use.
Tacit fragments: a fourth layer of agent memory
A tacit fragment records what an expert noticed, how they responded, and the circumstances of that response. After human review, it sits alongside procedures, facts, and past events in the agent's memory.
Related MCP server: Metatron
The gap between procedure and practice
Procedures describe what should happen, and logs record what happened. The cue behind an expert's decision, and the reason for it, often go unrecorded.
How a fragment reaches an agent
Every capture, confirmation, review decision, and retrieval is recorded through the
CHAP reference coordinator,
chap-coordinator, on a hash-linked evidence chain.
The capture loop
When a recorded action differs from the procedure, a capture agent asks the expert one short question, a whisper, and the expert confirms the account in their own words.
Seventeen kinds of know-how
Each fragment carries one of the paper's seventeen categories of tacit knowledge, K1 to K17. The atlas on the website gives an example of each and a way to capture it.
Quickstart: run the pump example
python -m pip install metis-memory
metis demo manufacturing-pump-vibration
metis fragment list
metis memory list
metis audit verifyThe demo uses supplied observations, needs no model server, and keeps its records in ./.metis.
To work from source:
git clone https://github.com/BrightbeamAI/metis && cd metis
pip install -e .from metis import MetisEngine
from metis.conditions.context import TacitContext
from metis.consent.model import ConsentRecord, ConsentStatus
eng = MetisEngine() # local and deterministic
eng.join_default_participants()
# Capture the operator's practice where it departs from the procedure.
frag = eng.capture_observation(
{
"observation_id": "OBS-1",
"work_as_imagined": "Reduce load only when the alarm threshold is crossed.",
"work_as_done": "Ease back earlier, when high load meets a dull sound.",
"context": TacitContext(equipment_family="centrifugal_pump", operating_mode="high_load"),
},
consent=ConsentRecord(consent_status=ConsentStatus.granted),
category="K7_sensory",
).fragment # Evidence layer: reviewers only
# Two named reviewers promote it to Advisory.
eng.tier2_review(
frag.fragment_id, "promoted_to_advisory", summary="advisory cue only",
decided_by=["human:quality-lead@metis.local", "human:process-engineer@metis.local"],
)
# The gate returns it only where its conditions hold.
pump = TacitContext(equipment_family="centrifugal_pump", operating_mode="high_load", risk_class="moderate")
other = TacitContext(equipment_family="gear_pump", operating_mode="low_load", risk_class="moderate")
print(len(eng.retrieve(pump).eligible)) # 1
print(eng.retrieve(other).blocked[0].reason) # conditions_do_not_matchConnect Metis to your application
Area | Metis provides | Your application supplies |
Capture | Fragment schemas and the whisper flow | Capture tools, consent workflows, and access control |
Review | Confirmation, review, and authority records | Reviewer identity and formal change control |
Retrieval | The condition-aware gate and its reasons | Current context, permissions, and domain policies |
Action | Guidance with its permitted uses | Action limits and human escalation |
Records | Local persistence and CHAP evidence | Storage, retention, and access policy |
metis mcp serves the same governed memory to MCP clients such as Claude Desktop and Claude Code,
and uvx metis-memory mcp runs it with nothing installed first. See the
MCP server guide. To run Metis for a team, the
server guide covers sign-in, workspace roles, the web app, and PostgreSQL;
deploy/ runs it with Docker or Kubernetes; and the
agent integrations guide connects agents through remote MCP, a
Python client, or LangChain. Connectors capture from workplace systems
and put whispers in Slack or Teams, and the operations guide covers running
it in production.
Learn more
Website: the interactive walkthrough, the atlas, and common questions.
Documentation: architecture, governance, retrieval, and agent use.
ABOUT.md: the repository map and how to develop.
CHAP: the Collaborative Human-Agent Protocol.
docs/demo.htmlanddocs/explainer.html: an interactive demo and an illustrated explainer that open in any browser.
Ethical use
Metis captures fragments of human work with the worker's knowledge and consent. Do not use it for covert monitoring. It records no audio, video, biometrics, screenshots, or keystrokes. Production use needs worker consultation, legal review, and domain validation; read ETHICAL_USE.md first.
License
Apache-2.0. See LICENSE.
Citation
Metis is the reference implementation of Tacit Fragments: Operationalising Tacit Knowledge as a Governed Memory Layer for Agentic AI.
@article{shahid2026tacitfragments,
title = {Tacit Fragments: Operationalising Tacit Knowledge as a Governed Memory Layer for Agentic AI},
author = {Shahid, Arsalan and Suttie, Gordon and Black, Philip and Garz{\'o}n-Vico, Antonio},
journal = {Preprints},
year = {2026},
doi = {10.20944/preprints202608.0927.v1},
url = {https://metis.brightbeam.works/resources/tacit-fragments-preprint.pdf}
}Available Tools
10 toolsagent_memory_contextAssemble memory for a taskA
Assemble a task's memory in one call: the workspace's procedures (SOPs), the facts and past cases that match the context, and the governed tacit guidance that applies. Use it when starting or planning a multi-step task. For a check before a single action, use retrieve_guidance instead; calling both for one step records two decisions. Tacit guidance passes the same condition-aware gate as retrieve_guidance, and task is a label recorded with the query. When required_human_actions lists anything (an escalation, or a use constraint that calls for a person's check), stop and hand the decision to a person. Each call records the query and the gate's decisions on the evidence chain, may open an escalation task, and changes no fragment.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | The role you act for, for example operator. A fragment whose conditions name other roles is withheld with reason role_not_authorised. Omit it to skip that check. | |
| task | Yes | A short name for the task, recorded with the query. | |
| context | Yes | The current work situation, with risk_class. |
Output Schema
| Name | Required | Description |
|---|---|---|
| task | Yes | The task the context was assembled for. |
| tacit | Yes | Governed tacit guidance that applies. Empty when none applies. |
| episodic | Yes | Past cases that match the context. |
| semantic | Yes | Facts about the equipment, materials, and process that match the context. |
| withheld | Yes | Fragments about this situation that were withheld, with the reason. |
| procedural | Yes | The workspace's procedures, such as SOPs, all of them. |
| governance_notes | Yes | The rules that shaped this context. |
| escalation_task_id | Yes | The escalation opened for a person, when the gate handed anything over. |
| not_yet_authorised | Yes | How many unreviewed or unauthorised fragments were left out. They stay unnamed until reviewers promote them. |
| required_human_actions | Yes | Actions a person must take before you act: escalations, and use constraints that call for a person's check. When any is listed, stop and hand the decision to a person. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond annotations: it discloses that each call records the query and the gate's decisions to an evidence chain, may open an escalation task, changes no fragment, that tacit guidance passes a condition-aware gate, and that a populated required_human_actions means stop and hand to a person. This complements readOnlyHint=false/destructiveHint=false by explaining exactly what side effects occur.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose before routing guidance and behavioral caveats. Every sentence carries information, though the passage is dense and could be trimmed slightly for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need no explanation. For a complex, non-idempotent assembly call, the description still covers the side effects, the human-handoff trigger, and the relationship to the sibling guidance tool, leaving little an agent needs to know beyond this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents task, context, and role. The description only restates that task is a label recorded with the query; it adds no matching-syntax or field-usage detail beyond the schema, so the baseline 3 fits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Assemble a task's memory in one call') and enumerates what is assembled: workspace SOPs, matching facts and past cases, and governed tacit guidance. It explicitly distinguishes itself from retrieve_guidance by scope (multi-step task vs single action).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('starting or planning a multi-step task'), names the alternative and its condition ('for a check before a single action, use retrieve_guidance'), and adds a concrete anti-pattern warning that calling both for one step records two decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
answer_whisperRelay a worker's answer to a whisperA
Relay a worker's own answer to a whisper, with the consent they stated. Use it only with the choice and words of the worker the whisper is addressed to, never with your own judgement. To start a capture, use submit_observation; to dispute a stored fragment, use contest_fragment. A fragment is stored only for confirm, or correct with the worker's corrected_text, given with consent granted: it enters the Evidence layer and reaches agents only after a quorum of Mission Group reviewers promotes it. Any other answer is recorded and stores nothing. Every answer, defer included, closes the whisper, and a second answer or one from anyone but its addressee is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| consent | Yes | The worker's own decision on keeping their account: granted keeps it, declined keeps nothing. Relay what the worker said. | |
| response | Yes | The worker's answer: confirm (the candidate is right), correct (right with the changes in corrected_text), dismiss (not a real practice), or defer (the worker will not answer now). Each closes the whisper. | |
| whisper_id | Yes | The whisper's id, from submit_observation or list_pending_whispers. | |
| answered_by | Yes | The worker who answered: the person the whisper is addressed to. | |
| corrected_text | No | Required with response correct: the worker's corrected account in their own words. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What happens next. |
| stored | Yes | True when the answer stored a fragment. |
| consent | No | The worker's recorded consent. |
| fragment_id | No | The stored fragment. |
| authority_layer | No | evidence: visible to agents only after reviewers promote it. |
| validation_state | No | Its validation state. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark this as non-read-only, non-destructive, non-idempotent. The description adds substantial beyond that: only 'confirm' or 'correct'+corrected_text stores a fragment, it enters an Evidence layer, and reaches agents only after Mission Group quorum promotion; every answer closes the whisper; a second answer is refused. This is exactly the kind of downstream consequence an agent cannot infer from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and constraint, then branches to alternatives and persistence behavior. Dense but every clause carries operational meaning; slight packing of multiple ideas into the final sentence costs a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the consent semantics, persistence rules, the quorum promotion pipeline, terminal nature of the call, and rejection behavior. With 5 params (4 required, 2 enums) and a mutation tool, this gives an agent everything needed to invoke correctly; the output schema handles the return shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description goes further by clarifying that 'defer is included' in the set that closes the whisper and that 'correct' requires the worker's own corrected_text — reinforcing enum semantics in a way that matters when the agent is relaying a worker's answer verbatim.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (relay) and resource (a worker's own answer to a whisper), then explicitly routes to siblings: submit_observation for starting a capture and contest_fragment for disputes. An agent can distinguish this from any other tool in the sibling list without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('only with the choice and words of the worker the whisper is addressed to, never with your own judgement'), when-not (any answer from anyone but the addressee is refused), and alternatives (submit_observation, contest_fragment). This is the strongest level of guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_tailRead the latest evidence-chain entriesARead-onlyIdempotent
Read the latest entries on the workspace's evidence chain, newest last: each gives its sequence number, method, and sender, without its content. Use it to confirm what was recorded; retrieve_guidance's recorded_at_seq appears here as an entry's seq. Raise limit to look further back; with no offset, the latest 200 entries are as far as this tool reaches. To check the chain's integrity, use audit_verify.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many of the latest entries to return. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes | The latest entries, newest last. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent/non-destructive, but the description adds genuinely non-obvious behavior: entries are returned newest last, entry content is deliberately withheld, and there is a hard reach of the latest 200 entries with no offset to go further back. That is a real scope limitation an agent could otherwise discover only by failing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the resource and return shape before the routing guidance. Slightly dense with the parenthetical cross-reference to retrieve_guidance's recorded_at_seq, but every clause carries information and nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be spelled out, and annotations carry the safety profile; the description still adds the return ordering, content exclusion, and reach limit. For a single-parameter read tool, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description earns an extra point by explaining the operational intent of limit ('Raise limit to look further back') and connecting it to the 200-entry ceiling, which tells the agent what raising the value can and cannot achieve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the latest entries on the workspace's evidence chain') and immediately specifies the shape of what comes back (sequence number, method, sender, without content). It also distinguishes itself from the sibling audit_verify, which handles integrity instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the use case ('confirm what was recorded') and names the alternative with the condition that selects it ('To check the chain's integrity, use audit_verify'). It further ties itself to retrieve_guidance by explaining that recorded_at_seq appears here as an entry's seq, which tells the agent when to reach for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_verifyVerify the evidence chainARead-onlyIdempotent
Verify the workspace's whole hash-linked evidence chain and, when the workspace keeps an append-only ledger, check that the ledger agrees entry for entry. Use it before relying on the record, or after an incident. For the verified flag alone, use describe_workspace; to list recent entries, use audit_tail. A failed check is a normal result, verified false with the errors found, and the check reads every entry, so it takes longer as the chain grows.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| errors | Yes | Each problem found. Empty when verified. |
| entries | Yes | Entries checked on the chain. |
| verified | Yes | True when every hash link verifies. |
| ledger_agrees | Yes | True when the ledger matches the chain entry for entry. |
| ledger_entries | Yes | Entries in the append-only ledger, when the workspace keeps one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), but the description adds genuinely useful behavior: a failed check is a normal result returned as verified=false with errors, and the check reads every entry so cost grows with chain length. Auth/permission requirements are not mentioned, but for a read-only verifier that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what it does, then routing, then failure semantics and cost. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, yet the description still supplies the key interpretive fact that verified=false is a normal outcome. Alternatives and performance characteristics are covered, so an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is trivially complete and the baseline is 4. The description correctly implies a whole-workspace, parameterless scan rather than a scoped one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (verify) plus resource (the workspace's whole hash-linked evidence chain), and it even names the sub-case of ledger agreement. An agent can distinguish it from describe_workspace and audit_tail from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States positive triggers ('before relying on the record, or after an incident') and explicitly routes to alternatives with their selecting conditions: describe_workspace for the verified flag alone, audit_tail for listing recent entries. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contest_fragmentRelay a contest of a fragmentADestructive
Relay a person's contest of a stored fragment: challenge it, correct it, supersede it, withdraw it, or ask the worker to give the account again. Use it when a worker or reviewer says a fragment is wrong or outdated, or its contributor takes it back; to answer a pending whisper, use answer_whisper. Only the fragment's contributor may withdraw it, which revokes it at once and for good: it leaves agent-visible memory, and its history stays on the evidence chain. Every other action opens or joins a Mission Group review, and the fragment stays in use until the reviewers decide; a fragment already out of review keeps the contest on record.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | challenge (question it), correct (propose new wording), supersede (replace it with a newer account), withdraw (the contributor takes it back), or request_re_elicitation (ask the worker to give the account again). | |
| raised_by | Yes | The person raising the contest. withdraw needs the fragment's contributor. | |
| rationale | Yes | Why, in the person's own words. It is recorded on the evidence chain. | |
| fragment_id | Yes | The fragment's id, from list_tacit_memory, a retrieval, or answer_whisper. | |
| proposed_correction | No | The replacement wording, for correct and supersede. |
Output Schema
| Name | Required | Description |
|---|---|---|
| action | Yes | The contest action recorded. |
| recorded | Yes | True once the contest is recorded. |
| revocation | No | The artefact recording the revocation, after withdraw. |
| mission_group_task | No | The Mission Group review the contest opened or joined. |
| contestability_record | No | The evidence-chain artefact recording the contest. |
| re_elicitation_request | No | The artefact asking the worker to give the account again. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say destructiveHint=true and idempotentHint=false; the description goes well beyond by disclosing what is actually destroyed ('revokes it at once and for good: it leaves agent-visible memory, and its history stays on the evidence chain') and the opposite behavior for all other actions ('opens or joins a Mission Group review, and the fragment stays in use until the reviewers decide'). It even covers the edge case of contesting a fragment already out of review. This is exactly the kind of consequence disclosure a destructive tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the action list, then behavior and the sibling disambiguation. The long final clause about contests on already-reviewed fragments is dense, but every sentence carries distinct operational information; only mild tightening is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with an output schema and annotations already present, the description supplies the missing pieces an agent needs: trigger conditions, the contributor-only restriction, the permanence of withdraw versus the review path for other actions, and the sibling alternative. Nothing needed to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the enum meanings, the 'human:' pattern, and the withdraw-requires-contributor rule are already documented in the schema, and the description largely restates them. It adds only slight framing (the actions as the person's intent), not new syntax or constraints, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Relay a person's contest of a stored fragment') and immediately enumerates the five concrete actions (challenge, correct, supersede, withdraw, request re-elicitation), so the agent knows exactly what the tool does. It also names the sibling it is not (answer_whisper), making it distinguishable from the other nine tools in the list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use it when a worker or reviewer says a fragment is wrong or outdated, or its contributor takes it back') and an explicit alternative-with-condition ('to answer a pending whisper, use answer_whisper'). It also states a precondition (only the contributor may withdraw), which routes the agent correctly before it even picks an action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_workspaceDescribe the workspaceARead-onlyIdempotent
Summarise the workspace this server serves: fragments per authority layer, agent-visible memory, pending whispers, Mission Group reviewers and their review rule, and whether the evidence chain verifies. Use it to orient yourself at the start of a session. It walks every hash link to report verified; for the ledger comparison and the errors found, use audit_verify, and for the memory itself, list_tacit_memory.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| name | Yes | The workspace name. |
| workspace | Yes | The workspace id. |
| review_rule | Yes | How many approvals a promotion needs, for example quorum:2. |
| evidence_chain | Yes | The evidence chain's state. |
| pending_whispers | Yes | Whispers awaiting a worker's answer. |
| agent_visible_memory | Yes | Memory objects agents may receive now. |
| mission_group_reviewers | Yes | The people who review and promote fragments. |
| fragments_by_authority_layer | Yes | Fragment counts per authority layer: evidence (awaiting review), advisory, controlled. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuine context beyond that: it discloses that the tool walks every hash link to determine verification, implying traversal cost, and it clarifies which concerns it deliberately does not handle. No cost/rate detail or latency warning is given, but the behavioral picture is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what is summarised, then when to use it, then boundaries and routing to alternatives. The enumeration is dense but each item maps to a distinct reported section, so nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers scope, invocation timing, and sibling boundaries. Nothing an agent needs in order to call this correctly at session start is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly implies a no-argument call by framing it as a whole-workspace summary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (summarise) and resource (the workspace this server serves), then enumerates the exact contents returned: fragments per authority layer, agent-visible memory, pending whispers, Mission Group reviewers, and evidence-chain verification. This distinguishes it from the narrower list_* siblings without the agent needing to open a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to call it ('use it to orient yourself at the start of a session') and names two alternatives with their selecting conditions: audit_verify for ledger comparison and errors found, list_tacit_memory for the memory itself. Both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pending_whispersList whispers awaiting an answerARead-onlyIdempotent
List the whispers still waiting for a worker's answer, oldest first, in one call. Use it to recover a whisper_id before relaying an answer with answer_whisper; submit_observation already returns the whisper_id, so call this only when you no longer have it. A whisper leaves the list once answered, and a deferred observation never appears because it raised none. worker filters by the exact URI given to submit_observation.
| Name | Required | Description | Default |
|---|---|---|---|
| worker | No | Show only the whispers addressed to this worker. Omit it for every worker. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes | Whispers awaiting an answer, oldest first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), but the description adds genuine lifecycle semantics beyond them: ordering is oldest-first, a whisper leaves the list once answered, and deferred observations never appear because they raised no whisper. No mention of pagination or result size limits, so not quite a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and scope, then usage routing, then edge-case behavior, then the parameter note. Every sentence carries information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a single-optional-param, read-only list tool the description covers purpose, routing, lifecycle edge cases, and filter semantics completely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the worker filter matches 'the exact URI given to submit_observation', clarifying exact-match semantics rather than a fuzzy name lookup.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the whispers still waiting for a worker's answer'), plus scope details (oldest first, one call). An agent can distinguish it from answer_whisper and submit_observation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (recover a whisper_id before relaying an answer) and when not to (submit_observation already returns the whisper_id, so call this only when you no longer have it). This is a textbook when/when-not plus alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tacit_memoryList agent-visible tacit memoryARead-onlyIdempotent
List the tacit memory agents may receive in this workspace, as metadata with each item's conditions of applicability and review date, without the guidance text. Use it to see which situations memory covers and which context fields to give retrieve_guidance; call retrieve_guidance to receive the guidance itself. Only memory in use appears (promoted, consented, inside its review date), all in one call. Listing records nothing, and a listed item can still be withheld at retrieval when the situation, role, or risk class does not allow it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes | Agent-visible memory: promoted, consented, and inside its review date. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, yet the description adds substantive traits: only promoted, consented, in-review-date memory appears; listing records nothing; and a listed item can still be withheld at retrieval based on situation, role, or risk class. That last point is a non-obvious behavioral caveat an agent would otherwise miss.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each doing work, with the core purpose front-loaded before the routing advice and the withholding caveat. Slightly dense and could trim the parenthetical enumeration, but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated; the description still characterizes what each item carries (conditions of applicability and review date) and explains the inclusion filter. Nothing needed to call the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no parameter vocabulary for the description to elaborate on, and it correctly abstains from inventing filtering syntax.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the tacit memory agents may receive in this workspace') and sharpens scope with 'as metadata ... without the guidance text'. It clearly delineates itself from the sibling retrieve_guidance, which returns the guidance itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use is given: 'see which situations memory covers and which context fields to give retrieve_guidance', with the alternative and its outcome named. No inference is required to route between list_tacit_memory and retrieve_guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retrieve_guidanceRetrieve governed guidanceA
Retrieve the reviewed tacit guidance that applies to the current work situation, with the use constraints reviewers attached. Call it before acting on equipment, a process, or a product. For a whole task's procedures, reference facts, and past cases as well, use agent_memory_context; to see what memory exists without its text, use list_tacit_memory. A context field left out never matches a condition, so give every field you know. role is the role you act for: a fragment restricted to other roles is withheld as not authorised, while context.role is matched like any other condition. When required_human_actions lists anything, stop and hand the decision to a person: high and critical risk and near misses (the same equipment in a different situation) always need one, and the handover opens an escalation task (escalation_task_id). Each call records the decision on the workspace's evidence chain and changes no fragment.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | The role you act for, for example operator. A fragment whose conditions name other roles is withheld with reason role_not_authorised. Omit it to skip that check. | |
| context | Yes | The current work situation, with risk_class. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | How to use the guidance. |
| guidance | Yes | Guidance whose recorded conditions match the context. Empty when none applies. |
| withheld | Yes | Fragments about this situation that were withheld, with the reason. |
| recorded_at_seq | Yes | The evidence-chain sequence number of the recorded decision. |
| escalation_task_id | Yes | The escalation opened for a person, when the gate handed anything over. |
| not_yet_authorised | Yes | How many unreviewed or unauthorised fragments were left out. They stay unnamed until reviewers promote them. |
| required_human_actions | Yes | Decisions a person must make before you act. When any is listed, stop and hand the decision to a person. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: each call writes to the workspace's evidence chain (explaining readOnlyHint=false and non-idempotency), does not alter fragments, withholds fragments for other roles as role_not_authorised, and opens an escalation task on handover. This clarifies rather than contradicts the annotation profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and the pre-action trigger, then layers alternatives and constraints. Dense but every sentence carries operational information; only the role/context.role clarification could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, yet the description still covers the consequential behaviors an agent must plan around: withholding, human handover, escalation_task_id, and evidence-chain recording. Complete for a governance-gated retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: omitted context fields never match a condition, so every known field should be supplied, and top-level role is an authorization filter while context.role is matched like an ordinary condition. It does not restate the size/entry limits already documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource and scope: 'Retrieve the reviewed tacit guidance that applies to the current work situation, with the use constraints reviewers attached.' It also distinguishes itself from siblings by name (agent_memory_context for full task context, list_tacit_memory for memory without text), so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger: 'Call it before acting on equipment, a process, or a product.' It names two alternatives with the exact conditions that select them, and gives a stop-and-escalate rule when required_human_actions is populated (high/critical risk, near misses). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_observationReport a divergence from procedureAIdempotent
Report where a worker's action differed from the written procedure; Metis drafts one short question (a whisper) for you to put to that worker. Use it when you see or are told that work was done differently from the SOP, and relay the worker's reply with answer_whisper; to dispute a stored fragment, use contest_fragment. Metis infers a candidate account (a hypothesis) and stores no fragment until the worker answers. A worker gets at most five whispers in eight hours by default; past that budget the call returns deferred, asks nothing, and spends the id, so report it again later under a new observation_id. A retry with the same observation_id, worker, and work_as_done returns the whisper while it awaits an answer; any other report under a used id is refused. Identities are not verified here, so give the worker's real URI.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | A short title for the candidate fragment. Metis drafts one when omitted. | |
| worker | Yes | The worker to ask, as their participant URI. Only a person (human:...) is asked. | |
| context | Yes | The situation the worker was in when the practice was observed. | |
| category | No | The kind of tacit knowledge, from the K1 to K17 taxonomy in the metis://taxonomy resource. Metis infers one when omitted. | |
| work_as_done | Yes | What the worker did, in plain words, as observed or reported. | |
| observation_id | Yes | Your stable id for this observation, such as a work-order number. Reuse it when you retry, so the worker is asked once; a deferral or an answer spends it. | |
| work_as_imagined | No | What the written procedure says should happen, when you know it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What to do next. |
| reason | No | Why the capture was deferred. |
| worker | No | The worker who is asked. |
| options | No | The answers the worker can give. |
| deferred | Yes | True when the worker had reached the whisper budget and nothing was asked. |
| question | No | The question to put to the worker. |
| repeated | No | True when this observation was reported before and its whisper is returned again. |
| candidate | No | The candidate account Metis inferred, for the worker to confirm or correct. |
| whisper_id | No | The whisper's id, for answer_whisper. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses the at-most-five-whispers-per-eight-hours budget, that exceeding it returns 'deferred' and spends the id, that no fragment is stored until the worker answers, that certain retries are refused, and that identities are not verified here. This is exactly the kind of GIGO/precondition detail annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and alternative, and every sentence carries operational weight (budget, deferral, retry, refusal). It is dense and long, but no sentence is filler; a small trim of the retry/refusal clauses could tighten it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, stateful, side-effecting tool with an output schema, the description covers the full lifecycle: what is stored, when, what is returned (a whisper or deferred), retry semantics, and failure modes. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics on top: observation_id must be reused on retry so the worker is asked once, and worker must be a real human URI because identity is unverified. The other five parameters' meaning is left entirely to the already-verbose schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('report where a worker's action differed from the written procedure') and immediately names the artifact produced (a drafted whisper). It also distinguishes itself from the sibling contest_fragment by naming that alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when you see or are told that work was done differently from the SOP'), the follow-up tool (answer_whisper), and the boundary case ('to dispute a stored fragment, use contest_fragment'). When-to-use, what-to-do-next, and the alternative are all stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.6- First observed
agent_memory_context - First observed
answer_whisper - First observed
audit_tail - First observed
audit_verify - First observed
contest_fragment - First observed
describe_workspace - First observed
list_pending_whispers - First observed
list_tacit_memory - First observed
retrieve_guidance - First observed
submit_observation
TDQS
Scored across 10 tools
Each tool has a clearly delineated role: observation submission, guidance retrieval, task memory assembly, metadata listing, workspace overview, whisper management, fragment contestation, and audit functions. The descriptions explicitly cross-reference and distinguish similar tools (e.g., retrieve_guidance vs agent_memory_context vs list_tacit_memory), preventing misselection.
Names are predominantly verb_noun in snake_case (submit_observation, retrieve_guidance, list_tacit_memory, etc.), but 'agent_memory_context' is a noun phrase and 'audit_verify'/'audit_tail' invert the verb-noun order, creating minor inconsistency.
10 tools appropriately cover the server's domain of tacit knowledge governance without redundancy or thinness; each tool has a distinct purpose and the count feels well-scoped.
The surface covers the full lifecycle: reporting observations, generating whispers, answering/contesting fragments, retrieving guidance, listing memory, auditing the evidence chain, and summarizing the workspace. No critical gaps are apparent.
Maintenance
Related MCP Connectors
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Give your AI agent persistent, governed memory for every project. At task start it recalls the approved decisions, conventions, risks and architecture (semantic search, ranked by importance); at close it proposes what was learned as typed memories that you review and approve — governance, not a notes dump. Agents propose, humans govern: edits go back to pending and deletion is human-only by design. Connect Claude Code, Cursor, Claude Desktop or any MCP client in two minutes with just your API key — hosted (nothing to install) or locally via `uvx solucortex-mcp`. Built by SoluAI and dogfooded daily: SoluCortex is developed using its own living memory.
Private, portable memory and reusable skills for AI agents.
Related MCP Servers
- AlicenseAqualityAmaintenanceSelf-hosted memory and governance layer for AI coding agents. 28 MCP tools with hybrid search, structured knowledge capture, behavioral nudges, and git-native storage. Zero cloud dependencies.306Business Source 1.1
- AlicenseAqualityAmaintenanceMetatron is a self-hosted system that captures a codebase's real implementation decisions — preferred patterns, rejected approaches, edge cases, internal conventions — as structured priors, and serves them to coding agents over MCP524MIT
- AlicenseNot gradedqualityAmaintenanceLocal-first, source-grounded memory for AI agents, with citations, bitemporal history, review-gated corrections, and MCP tools for search and recall.3Apache 2.0
- AlicenseAqualityAmaintenanceMCP server that gives coding agents persistent, verified memory of codebase decisions, conventions, and skills, with evidence-based claims that are re-checked via git hooks and human-gated review. Enables memory search, propose/approve, chat harvesting, and critique across MCP-compatible tools.21108 npm1MIT