Skip to main content
Glama

Metis is an open-source toolkit for capturing fragments of expert practice and making them available to AI agents as memory, with human review and agreed conditions for use.

Tacit fragments: a fourth layer of agent memory

A tacit fragment records what an expert noticed, how they responded, and the circumstances of that response. After human review, it sits alongside procedures, facts, and past events in the agent's memory.

Related MCP server: Metatron

The gap between procedure and practice

Procedures describe what should happen, and logs record what happened. The cue behind an expert's decision, and the reason for it, often go unrecorded.

How a fragment reaches an agent

Every capture, confirmation, review decision, and retrieval is recorded through the CHAP reference coordinator, chap-coordinator, on a hash-linked evidence chain.

The capture loop

When a recorded action differs from the procedure, a capture agent asks the expert one short question, a whisper, and the expert confirms the account in their own words.

Seventeen kinds of know-how

Each fragment carries one of the paper's seventeen categories of tacit knowledge, K1 to K17. The atlas on the website gives an example of each and a way to capture it.

Quickstart: run the pump example

python -m pip install metis-memory
metis demo manufacturing-pump-vibration
metis fragment list
metis memory list
metis audit verify

The demo uses supplied observations, needs no model server, and keeps its records in ./.metis. To work from source:

git clone https://github.com/BrightbeamAI/metis && cd metis
pip install -e .
from metis import MetisEngine
from metis.conditions.context import TacitContext
from metis.consent.model import ConsentRecord, ConsentStatus

eng = MetisEngine()  # local and deterministic
eng.join_default_participants()

# Capture the operator's practice where it departs from the procedure.
frag = eng.capture_observation(
    {
        "observation_id": "OBS-1",
        "work_as_imagined": "Reduce load only when the alarm threshold is crossed.",
        "work_as_done": "Ease back earlier, when high load meets a dull sound.",
        "context": TacitContext(equipment_family="centrifugal_pump", operating_mode="high_load"),
    },
    consent=ConsentRecord(consent_status=ConsentStatus.granted),
    category="K7_sensory",
).fragment  # Evidence layer: reviewers only

# Two named reviewers promote it to Advisory.
eng.tier2_review(
    frag.fragment_id, "promoted_to_advisory", summary="advisory cue only",
    decided_by=["human:quality-lead@metis.local", "human:process-engineer@metis.local"],
)

# The gate returns it only where its conditions hold.
pump = TacitContext(equipment_family="centrifugal_pump", operating_mode="high_load", risk_class="moderate")
other = TacitContext(equipment_family="gear_pump", operating_mode="low_load", risk_class="moderate")
print(len(eng.retrieve(pump).eligible))       # 1
print(eng.retrieve(other).blocked[0].reason)  # conditions_do_not_match

Connect Metis to your application

Area

Metis provides

Your application supplies

Capture

Fragment schemas and the whisper flow

Capture tools, consent workflows, and access control

Review

Confirmation, review, and authority records

Reviewer identity and formal change control

Retrieval

The condition-aware gate and its reasons

Current context, permissions, and domain policies

Action

Guidance with its permitted uses

Action limits and human escalation

Records

Local persistence and CHAP evidence

Storage, retention, and access policy

metis mcp serves the same governed memory to MCP clients such as Claude Desktop and Claude Code, and uvx metis-memory mcp runs it with nothing installed first. See the MCP server guide. To run Metis for a team, the server guide covers sign-in, workspace roles, the web app, and PostgreSQL; deploy/ runs it with Docker or Kubernetes; and the agent integrations guide connects agents through remote MCP, a Python client, or LangChain. Connectors capture from workplace systems and put whispers in Slack or Teams, and the operations guide covers running it in production.

Learn more

  • Website: the interactive walkthrough, the atlas, and common questions.

  • Documentation: architecture, governance, retrieval, and agent use.

  • ABOUT.md: the repository map and how to develop.

  • CHAP: the Collaborative Human-Agent Protocol.

  • docs/demo.html and docs/explainer.html: an interactive demo and an illustrated explainer that open in any browser.

Ethical use

Metis captures fragments of human work with the worker's knowledge and consent. Do not use it for covert monitoring. It records no audio, video, biometrics, screenshots, or keystrokes. Production use needs worker consultation, legal review, and domain validation; read ETHICAL_USE.md first.

License

Apache-2.0. See LICENSE.

Citation

Metis is the reference implementation of Tacit Fragments: Operationalising Tacit Knowledge as a Governed Memory Layer for Agentic AI.

@article{shahid2026tacitfragments,
  title   = {Tacit Fragments: Operationalising Tacit Knowledge as a Governed Memory Layer for Agentic AI},
  author  = {Shahid, Arsalan and Suttie, Gordon and Black, Philip and Garz{\'o}n-Vico, Antonio},
  journal = {Preprints},
  year    = {2026},
  doi     = {10.20944/preprints202608.0927.v1},
  url     = {https://metis.brightbeam.works/resources/tacit-fragments-preprint.pdf}
}

Available Tools

10 tools
agent_memory_contextAssemble memory for a taskA

Assemble a task's memory in one call: the workspace's procedures (SOPs), the facts and past cases that match the context, and the governed tacit guidance that applies. Use it when starting or planning a multi-step task. For a check before a single action, use retrieve_guidance instead; calling both for one step records two decisions. Tacit guidance passes the same condition-aware gate as retrieve_guidance, and task is a label recorded with the query. When required_human_actions lists anything (an escalation, or a use constraint that calls for a person's check), stop and hand the decision to a person. Each call records the query and the gate's decisions on the evidence chain, may open an escalation task, and changes no fragment.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoThe role you act for, for example operator. A fragment whose conditions name other roles is withheld with reason role_not_authorised. Omit it to skip that check.
taskYesA short name for the task, recorded with the query.
contextYesThe current work situation, with risk_class.

Output Schema

ParametersJSON Schema
NameRequiredDescription
taskYesThe task the context was assembled for.
tacitYesGoverned tacit guidance that applies. Empty when none applies.
episodicYesPast cases that match the context.
semanticYesFacts about the equipment, materials, and process that match the context.
withheldYesFragments about this situation that were withheld, with the reason.
proceduralYesThe workspace's procedures, such as SOPs, all of them.
governance_notesYesThe rules that shaped this context.
escalation_task_idYesThe escalation opened for a person, when the gate handed anything over.
not_yet_authorisedYesHow many unreviewed or unauthorised fragments were left out. They stay unnamed until reviewers promote them.
required_human_actionsYesActions a person must take before you act: escalations, and use constraints that call for a person's check. When any is listed, stop and hand the decision to a person.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond annotations: it discloses that each call records the query and the gate's decisions to an evidence chain, may open an escalation task, changes no fragment, that tacit guidance passes a condition-aware gate, and that a populated required_human_actions means stop and hand to a person. This complements readOnlyHint=false/destructiveHint=false by explaining exactly what side effects occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose before routing guidance and behavioral caveats. Every sentence carries information, though the passage is dense and could be trimmed slightly for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need no explanation. For a complex, non-idempotent assembly call, the description still covers the side effects, the human-handoff trigger, and the relationship to the sibling guidance tool, leaving little an agent needs to know beyond this.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents task, context, and role. The description only restates that task is a label recorded with the query; it adds no matching-syntax or field-usage detail beyond the schema, so the baseline 3 fits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Assemble a task's memory in one call') and enumerates what is assembled: workspace SOPs, matching facts and past cases, and governed tacit guidance. It explicitly distinguishes itself from retrieve_guidance by scope (multi-step task vs single action).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('starting or planning a multi-step task'), names the alternative and its condition ('for a check before a single action, use retrieve_guidance'), and adds a concrete anti-pattern warning that calling both for one step records two decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

answer_whisperRelay a worker's answer to a whisperA

Relay a worker's own answer to a whisper, with the consent they stated. Use it only with the choice and words of the worker the whisper is addressed to, never with your own judgement. To start a capture, use submit_observation; to dispute a stored fragment, use contest_fragment. A fragment is stored only for confirm, or correct with the worker's corrected_text, given with consent granted: it enters the Evidence layer and reaches agents only after a quorum of Mission Group reviewers promotes it. Any other answer is recorded and stores nothing. Every answer, defer included, closes the whisper, and a second answer or one from anyone but its addressee is refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
consentYesThe worker's own decision on keeping their account: granted keeps it, declined keeps nothing. Relay what the worker said.
responseYesThe worker's answer: confirm (the candidate is right), correct (right with the changes in corrected_text), dismiss (not a real practice), or defer (the worker will not answer now). Each closes the whisper.
whisper_idYesThe whisper's id, from submit_observation or list_pending_whispers.
answered_byYesThe worker who answered: the person the whisper is addressed to.
corrected_textNoRequired with response correct: the worker's corrected account in their own words.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYesWhat happens next.
storedYesTrue when the answer stored a fragment.
consentNoThe worker's recorded consent.
fragment_idNoThe stored fragment.
authority_layerNoevidence: visible to agents only after reviewers promote it.
validation_stateNoIts validation state.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark this as non-read-only, non-destructive, non-idempotent. The description adds substantial beyond that: only 'confirm' or 'correct'+corrected_text stores a fragment, it enters an Evidence layer, and reaches agents only after Mission Group quorum promotion; every answer closes the whisper; a second answer is refused. This is exactly the kind of downstream consequence an agent cannot infer from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and constraint, then branches to alternatives and persistence behavior. Dense but every clause carries operational meaning; slight packing of multiple ideas into the final sentence costs a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the consent semantics, persistence rules, the quorum promotion pipeline, terminal nature of the call, and rejection behavior. With 5 params (4 required, 2 enums) and a mutation tool, this gives an agent everything needed to invoke correctly; the output schema handles the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description goes further by clarifying that 'defer is included' in the set that closes the whisper and that 'correct' requires the worker's own corrected_text — reinforcing enum semantics in a way that matters when the agent is relaying a worker's answer verbatim.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (relay) and resource (a worker's own answer to a whisper), then explicitly routes to siblings: submit_observation for starting a capture and contest_fragment for disputes. An agent can distinguish this from any other tool in the sibling list without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('only with the choice and words of the worker the whisper is addressed to, never with your own judgement'), when-not (any answer from anyone but the addressee is refused), and alternatives (submit_observation, contest_fragment). This is the strongest level of guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_tailRead the latest evidence-chain entriesA
Read-onlyIdempotent

Read the latest entries on the workspace's evidence chain, newest last: each gives its sequence number, method, and sender, without its content. Use it to confirm what was recorded; retrieve_guidance's recorded_at_seq appears here as an entry's seq. Raise limit to look further back; with no offset, the latest 200 entries are as far as this tool reaches. To check the chain's integrity, use audit_verify.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many of the latest entries to return.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesThe latest entries, newest last.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, but the description adds genuinely non-obvious behavior: entries are returned newest last, entry content is deliberately withheld, and there is a hard reach of the latest 200 entries with no offset to go further back. That is a real scope limitation an agent could otherwise discover only by failing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the resource and return shape before the routing guidance. Slightly dense with the parenthetical cross-reference to retrieve_guidance's recorded_at_seq, but every clause carries information and nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be spelled out, and annotations carry the safety profile; the description still adds the return ordering, content exclusion, and reach limit. For a single-parameter read tool, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description earns an extra point by explaining the operational intent of limit ('Raise limit to look further back') and connecting it to the 200-entry ceiling, which tells the agent what raising the value can and cannot achieve.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the latest entries on the workspace's evidence chain') and immediately specifies the shape of what comes back (sequence number, method, sender, without content). It also distinguishes itself from the sibling audit_verify, which handles integrity instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the use case ('confirm what was recorded') and names the alternative with the condition that selects it ('To check the chain's integrity, use audit_verify'). It further ties itself to retrieve_guidance by explaining that recorded_at_seq appears here as an entry's seq, which tells the agent when to reach for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_verifyVerify the evidence chainA
Read-onlyIdempotent

Verify the workspace's whole hash-linked evidence chain and, when the workspace keeps an append-only ledger, check that the ledger agrees entry for entry. Use it before relying on the record, or after an incident. For the verified flag alone, use describe_workspace; to list recent entries, use audit_tail. A failed check is a normal result, verified false with the errors found, and the check reads every entry, so it takes longer as the chain grows.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorsYesEach problem found. Empty when verified.
entriesYesEntries checked on the chain.
verifiedYesTrue when every hash link verifies.
ledger_agreesYesTrue when the ledger matches the chain entry for entry.
ledger_entriesYesEntries in the append-only ledger, when the workspace keeps one.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), but the description adds genuinely useful behavior: a failed check is a normal result returned as verified=false with errors, and the check reads every entry so cost grows with chain length. Auth/permission requirements are not mentioned, but for a read-only verifier that is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with what it does, then routing, then failure semantics and cost. Every sentence carries distinct information with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, yet the description still supplies the key interpretive fact that verified=false is a normal outcome. Alternatives and performance characteristics are covered, so an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema is trivially complete and the baseline is 4. The description correctly implies a whole-workspace, parameterless scan rather than a scoped one.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (verify) plus resource (the workspace's whole hash-linked evidence chain), and it even names the sub-case of ledger agreement. An agent can distinguish it from describe_workspace and audit_tail from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States positive triggers ('before relying on the record, or after an incident') and explicitly routes to alternatives with their selecting conditions: describe_workspace for the verified flag alone, audit_tail for listing recent entries. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contest_fragmentRelay a contest of a fragmentA
Destructive

Relay a person's contest of a stored fragment: challenge it, correct it, supersede it, withdraw it, or ask the worker to give the account again. Use it when a worker or reviewer says a fragment is wrong or outdated, or its contributor takes it back; to answer a pending whisper, use answer_whisper. Only the fragment's contributor may withdraw it, which revokes it at once and for good: it leaves agent-visible memory, and its history stays on the evidence chain. Every other action opens or joins a Mission Group review, and the fragment stays in use until the reviewers decide; a fragment already out of review keeps the contest on record.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYeschallenge (question it), correct (propose new wording), supersede (replace it with a newer account), withdraw (the contributor takes it back), or request_re_elicitation (ask the worker to give the account again).
raised_byYesThe person raising the contest. withdraw needs the fragment's contributor.
rationaleYesWhy, in the person's own words. It is recorded on the evidence chain.
fragment_idYesThe fragment's id, from list_tacit_memory, a retrieval, or answer_whisper.
proposed_correctionNoThe replacement wording, for correct and supersede.

Output Schema

ParametersJSON Schema
NameRequiredDescription
actionYesThe contest action recorded.
recordedYesTrue once the contest is recorded.
revocationNoThe artefact recording the revocation, after withdraw.
mission_group_taskNoThe Mission Group review the contest opened or joined.
contestability_recordNoThe evidence-chain artefact recording the contest.
re_elicitation_requestNoThe artefact asking the worker to give the account again.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say destructiveHint=true and idempotentHint=false; the description goes well beyond by disclosing what is actually destroyed ('revokes it at once and for good: it leaves agent-visible memory, and its history stays on the evidence chain') and the opposite behavior for all other actions ('opens or joins a Mission Group review, and the fragment stays in use until the reviewers decide'). It even covers the edge case of contesting a fragment already out of review. This is exactly the kind of consequence disclosure a destructive tool needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the action list, then behavior and the sibling disambiguation. The long final clause about contests on already-reviewed fragments is dense, but every sentence carries distinct operational information; only mild tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema and annotations already present, the description supplies the missing pieces an agent needs: trigger conditions, the contributor-only restriction, the permanence of withdraw versus the review path for other actions, and the sibling alternative. Nothing needed to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the enum meanings, the 'human:' pattern, and the withdraw-requires-contributor rule are already documented in the schema, and the description largely restates them. It adds only slight framing (the actions as the person's intent), not new syntax or constraints, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Relay a person's contest of a stored fragment') and immediately enumerates the five concrete actions (challenge, correct, supersede, withdraw, request re-elicitation), so the agent knows exactly what the tool does. It also names the sibling it is not (answer_whisper), making it distinguishable from the other nine tools in the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use it when a worker or reviewer says a fragment is wrong or outdated, or its contributor takes it back') and an explicit alternative-with-condition ('to answer a pending whisper, use answer_whisper'). It also states a precondition (only the contributor may withdraw), which routes the agent correctly before it even picks an action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_workspaceDescribe the workspaceA
Read-onlyIdempotent

Summarise the workspace this server serves: fragments per authority layer, agent-visible memory, pending whispers, Mission Group reviewers and their review rule, and whether the evidence chain verifies. Use it to orient yourself at the start of a session. It walks every hash link to report verified; for the ledger comparison and the errors found, use audit_verify, and for the memory itself, list_tacit_memory.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYesThe workspace name.
workspaceYesThe workspace id.
review_ruleYesHow many approvals a promotion needs, for example quorum:2.
evidence_chainYesThe evidence chain's state.
pending_whispersYesWhispers awaiting a worker's answer.
agent_visible_memoryYesMemory objects agents may receive now.
mission_group_reviewersYesThe people who review and promote fragments.
fragments_by_authority_layerYesFragment counts per authority layer: evidence (awaiting review), advisory, controlled.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuine context beyond that: it discloses that the tool walks every hash link to determine verification, implying traversal cost, and it clarifies which concerns it deliberately does not handle. No cost/rate detail or latency warning is given, but the behavioral picture is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with what is summarised, then when to use it, then boundaries and routing to alternatives. The enumeration is dense but each item maps to a distinct reported section, so nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers scope, invocation timing, and sibling boundaries. Nothing an agent needs in order to call this correctly at session start is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly implies a no-argument call by framing it as a whole-workspace summary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (summarise) and resource (the workspace this server serves), then enumerates the exact contents returned: fragments per authority layer, agent-visible memory, pending whispers, Mission Group reviewers, and evidence-chain verification. This distinguishes it from the narrower list_* siblings without the agent needing to open a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call it ('use it to orient yourself at the start of a session') and names two alternatives with their selecting conditions: audit_verify for ledger comparison and errors found, list_tacit_memory for the memory itself. Both when-to-use and when-not-to-use are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pending_whispersList whispers awaiting an answerA
Read-onlyIdempotent

List the whispers still waiting for a worker's answer, oldest first, in one call. Use it to recover a whisper_id before relaying an answer with answer_whisper; submit_observation already returns the whisper_id, so call this only when you no longer have it. A whisper leaves the list once answered, and a deferred observation never appears because it raised none. worker filters by the exact URI given to submit_observation.

ParametersJSON Schema
NameRequiredDescriptionDefault
workerNoShow only the whispers addressed to this worker. Omit it for every worker.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesWhispers awaiting an answer, oldest first.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), but the description adds genuine lifecycle semantics beyond them: ordering is oldest-first, a whisper leaves the list once answered, and deferred observations never appear because they raised no whisper. No mention of pagination or result size limits, so not quite a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and scope, then usage routing, then edge-case behavior, then the parameter note. Every sentence carries information; there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a single-optional-param, read-only list tool the description covers purpose, routing, lifecycle edge cases, and filter semantics completely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the worker filter matches 'the exact URI given to submit_observation', clarifying exact-match semantics rather than a fuzzy name lookup.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the whispers still waiting for a worker's answer'), plus scope details (oldest first, one call). An agent can distinguish it from answer_whisper and submit_observation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (recover a whisper_id before relaying an answer) and when not to (submit_observation already returns the whisper_id, so call this only when you no longer have it). This is a textbook when/when-not plus alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tacit_memoryList agent-visible tacit memoryA
Read-onlyIdempotent

List the tacit memory agents may receive in this workspace, as metadata with each item's conditions of applicability and review date, without the guidance text. Use it to see which situations memory covers and which context fields to give retrieve_guidance; call retrieve_guidance to receive the guidance itself. Only memory in use appears (promoted, consented, inside its review date), all in one call. Listing records nothing, and a listed item can still be withheld at retrieval when the situation, role, or risk class does not allow it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesAgent-visible memory: promoted, consented, and inside its review date.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, yet the description adds substantive traits: only promoted, consented, in-review-date memory appears; listing records nothing; and a listed item can still be withheld at retrieval based on situation, role, or risk class. That last point is a non-obvious behavioral caveat an agent would otherwise miss.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each doing work, with the core purpose front-loaded before the routing advice and the withholding caveat. Slightly dense and could trim the parenthetical enumeration, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be restated; the description still characterizes what each item carries (conditions of applicability and review date) and explains the inclusion filter. Nothing needed to call the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is no parameter vocabulary for the description to elaborate on, and it correctly abstains from inventing filtering syntax.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the tacit memory agents may receive in this workspace') and sharpens scope with 'as metadata ... without the guidance text'. It clearly delineates itself from the sibling retrieve_guidance, which returns the guidance itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use is given: 'see which situations memory covers and which context fields to give retrieve_guidance', with the alternative and its outcome named. No inference is required to route between list_tacit_memory and retrieve_guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retrieve_guidanceRetrieve governed guidanceA

Retrieve the reviewed tacit guidance that applies to the current work situation, with the use constraints reviewers attached. Call it before acting on equipment, a process, or a product. For a whole task's procedures, reference facts, and past cases as well, use agent_memory_context; to see what memory exists without its text, use list_tacit_memory. A context field left out never matches a condition, so give every field you know. role is the role you act for: a fragment restricted to other roles is withheld as not authorised, while context.role is matched like any other condition. When required_human_actions lists anything, stop and hand the decision to a person: high and critical risk and near misses (the same equipment in a different situation) always need one, and the handover opens an escalation task (escalation_task_id). Each call records the decision on the workspace's evidence chain and changes no fragment.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoThe role you act for, for example operator. A fragment whose conditions name other roles is withheld with reason role_not_authorised. Omit it to skip that check.
contextYesThe current work situation, with risk_class.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYesHow to use the guidance.
guidanceYesGuidance whose recorded conditions match the context. Empty when none applies.
withheldYesFragments about this situation that were withheld, with the reason.
recorded_at_seqYesThe evidence-chain sequence number of the recorded decision.
escalation_task_idYesThe escalation opened for a person, when the gate handed anything over.
not_yet_authorisedYesHow many unreviewed or unauthorised fragments were left out. They stay unnamed until reviewers promote them.
required_human_actionsYesDecisions a person must make before you act. When any is listed, stop and hand the decision to a person.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: each call writes to the workspace's evidence chain (explaining readOnlyHint=false and non-idempotency), does not alter fragments, withholds fragments for other roles as role_not_authorised, and opens an escalation task on handover. This clarifies rather than contradicts the annotation profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the pre-action trigger, then layers alternatives and constraints. Dense but every sentence carries operational information; only the role/context.role clarification could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, yet the description still covers the consequential behaviors an agent must plan around: withholding, human handover, escalation_task_id, and evidence-chain recording. Complete for a governance-gated retrieval tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: omitted context fields never match a condition, so every known field should be supplied, and top-level role is an authorization filter while context.role is matched like an ordinary condition. It does not restate the size/entry limits already documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource and scope: 'Retrieve the reviewed tacit guidance that applies to the current work situation, with the use constraints reviewers attached.' It also distinguishes itself from siblings by name (agent_memory_context for full task context, list_tacit_memory for memory without text), so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger: 'Call it before acting on equipment, a process, or a product.' It names two alternatives with the exact conditions that select them, and gives a stop-and-escalate rule when required_human_actions is populated (high/critical risk, near misses). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_observationReport a divergence from procedureA
Idempotent

Report where a worker's action differed from the written procedure; Metis drafts one short question (a whisper) for you to put to that worker. Use it when you see or are told that work was done differently from the SOP, and relay the worker's reply with answer_whisper; to dispute a stored fragment, use contest_fragment. Metis infers a candidate account (a hypothesis) and stores no fragment until the worker answers. A worker gets at most five whispers in eight hours by default; past that budget the call returns deferred, asks nothing, and spends the id, so report it again later under a new observation_id. A retry with the same observation_id, worker, and work_as_done returns the whisper while it awaits an answer; any other report under a used id is refused. Identities are not verified here, so give the worker's real URI.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoA short title for the candidate fragment. Metis drafts one when omitted.
workerYesThe worker to ask, as their participant URI. Only a person (human:...) is asked.
contextYesThe situation the worker was in when the practice was observed.
categoryNoThe kind of tacit knowledge, from the K1 to K17 taxonomy in the metis://taxonomy resource. Metis infers one when omitted.
work_as_doneYesWhat the worker did, in plain words, as observed or reported.
observation_idYesYour stable id for this observation, such as a work-order number. Reuse it when you retry, so the worker is asked once; a deferral or an answer spends it.
work_as_imaginedNoWhat the written procedure says should happen, when you know it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYesWhat to do next.
reasonNoWhy the capture was deferred.
workerNoThe worker who is asked.
optionsNoThe answers the worker can give.
deferredYesTrue when the worker had reached the whisper budget and nothing was asked.
questionNoThe question to put to the worker.
repeatedNoTrue when this observation was reported before and its whisper is returned again.
candidateNoThe candidate account Metis inferred, for the worker to confirm or correct.
whisper_idNoThe whisper's id, for answer_whisper.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses the at-most-five-whispers-per-eight-hours budget, that exceeding it returns 'deferred' and spends the id, that no fragment is stored until the worker answers, that certain retries are refused, and that identities are not verified here. This is exactly the kind of GIGO/precondition detail annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and alternative, and every sentence carries operational weight (budget, deferral, retry, refusal). It is dense and long, but no sentence is filler; a small trim of the retry/refusal clauses could tighten it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, stateful, side-effecting tool with an output schema, the description covers the full lifecycle: what is stored, when, what is returned (a whisper or deferred), retry semantics, and failure modes. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics on top: observation_id must be reused on retry so the worker is asked once, and worker must be a real human URI because identity is unverified. The other five parameters' meaning is left entirely to the already-verbose schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('report where a worker's action differed from the written procedure') and immediately names the artifact produced (a drafted whisper). It also distinguishes itself from the sibling contest_fragment by naming that alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('when you see or are told that work was done differently from the SOP'), the follow-up tool (answer_whisper), and the boundary case ('to dispute a stored fragment, use contest_fragment'). When-to-use, what-to-do-next, and the alternative are all stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.6
    • First observedagent_memory_context
    • First observedanswer_whisper
    • First observedaudit_tail
    • First observedaudit_verify
    • First observedcontest_fragment
    • First observeddescribe_workspace
    • First observedlist_pending_whispers
    • First observedlist_tacit_memory
    • First observedretrieve_guidance
    • First observedsubmit_observation

TDQS

A4.7/5.0

Scored across 10 tools

Disambiguation5/5

Each tool has a clearly delineated role: observation submission, guidance retrieval, task memory assembly, metadata listing, workspace overview, whisper management, fragment contestation, and audit functions. The descriptions explicitly cross-reference and distinguish similar tools (e.g., retrieve_guidance vs agent_memory_context vs list_tacit_memory), preventing misselection.

Naming Consistency4/5

Names are predominantly verb_noun in snake_case (submit_observation, retrieve_guidance, list_tacit_memory, etc.), but 'agent_memory_context' is a noun phrase and 'audit_verify'/'audit_tail' invert the verb-noun order, creating minor inconsistency.

Tool Count5/5

10 tools appropriately cover the server's domain of tacit knowledge governance without redundancy or thinness; each tool has a distinct purpose and the count feels well-scoped.

Completeness5/5

The surface covers the full lifecycle: reporting observations, generating whispers, answering/contesting fragments, retrieving guidance, listing memory, auditing the evidence chain, and summarizing the workspace. No critical gaps are apparent.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Self-hosted memory and governance layer for AI coding agents. 28 MCP tools with hybrid search, structured knowledge capture, behavioral nudges, and git-native storage. Zero cloud dependencies.
    30
    6
    Business Source 1.1
  • A
    license
    A
    quality
    A
    maintenance
    Metatron is a self-hosted system that captures a codebase's real implementation decisions — preferred patterns, rejected approaches, edge cases, internal conventions — as structured priors, and serves them to coding agents over MCP
    5
    24
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first, source-grounded memory for AI agents, with citations, bitemporal history, review-gated corrections, and MCP tools for search and recall.
    3
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    MCP server that gives coding agents persistent, verified memory of codebase decisions, conventions, and skills, with evidence-based claims that are re-checked via git hooks and human-gated review. Enables memory search, propose/approve, chat harvesting, and critique across MCP-compatible tools.
    21
    108 npm
    1
    MIT