MIHAD
Summary: This MIHAD memory server lets a coding agent recall verified project memory, propose new items for verification, correct wrong ones, and report whether used items helped — so knowledge is adopted only with independent evidence.
Recall memory (
memory_recall) — Search persistent memory before starting work; pass aquerydescribing the task and optionallyk(default 5). Returns verified skills, facts and lessons pluspreferences(standing user working preferences you should follow). Items flagged provisional/unverified are hints only.Propose memory (
memory_propose) — Add something worth remembering, withkind:skill— reusable technique/pattern, verified by running existing project tests (verify.command,verify.test_files).fact— claim about the codebase, verified by quoting a project file (verify.file,verify.quote).preference— how the user wants future work done;verify.quotemust be the user's exact words.lesson— general lesson; stays provisional until a human approves it.Also supports
title,content,tags,supersedes(replaces an older item), andderived_from(item IDs it builds on).
Correct memory (
memory_correct) — Mark an item (byid) as wrong with areason; it and everything derived from it are suspended.Report outcomes (
memory_report_outcome) — At task end, reportsuccessand optionallytask_id,used_ids,helpful_ids,harmful_ids, andnotes, so the system learns which memory actually helped or caused problems.Overall — Only adopted items count as verified knowledge; verification is kind-specific (project tests, code quotes, or user's own words), keeping unproven conclusions out of the agent's memory.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MIHADremember that I always want a regression test with every fix"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
💡 Why MIHAD
Coding agents start every session like a new hire on day one. They forget yesterday's corrections and repeat the same mistakes. Saving everything they "learn" is not the answer either: an agent that remembers its own wrong conclusions repeats them with confidence.
MIHAD gives an agent a memory and a body of experience that grow with use, and it adopts nothing without independent evidence. It is built on a research programme with pinned protocols, real repository history and a blind quality judge. Its measured gains are largest for cheaper models: with a strong model as the agent, success was the same with and without it (51/63 vs 50/63; on the sound tasks of that benchmark, 45/45 either way). See the paper.
Related MCP server: Code Project Brain
✨ What it does
🛡️ Verified memory
A skill must pass your project's own tests. A fact must quote the code as it was before the session. A preference must be your own words. Everything else stays provisional, and corrections reach everything built on an item.
🗣️ Your preferences, kept
Say "Always add a regression test" once, in any message. It is captured verbatim, adopted, and followed in every later session. One-off instructions are never stored.
🧭 Brief, live warnings, review
At the start of each request the agent gets a short brief. While it works, known pitfalls are flagged the moment they recur. Before it finishes, its change is reviewed and executed: lines no test covers, lost formatting, broken preferences. Findings send it back to work.
🔁 Learns from your corrections
Every session is recorded. When you commit, your commit is the final version, and what you changed after the agent becomes experience. There is nothing extra to do.
🌙 Dreaming with A/B tests
Between sessions it practises on slightly broken copies of past fixes, with and without its lessons. A lesson is kept only if the test shows it helps.
🧩 Works where you work
OMP, Claude Code (terminal and desktop) and Codex. Ten programming languages. Standard-library Python with no dependencies. One command installs it into a project.
📊 What the experiments showed
Finding | Evidence | |
✅ | Captured preferences carried into every later task | 15/15 vs 0/15 without capture |
✅ | Wrong items planted in memory were never followed | Two contaminated-memory runs on real commits |
✅ | Operational experience transferred to new tasks | Same success (11/14), 14.4% fewer tokens, quality 3.93 vs 3.79 |
❌ | Knowledge about the code did not transfer | Same success, +16–17% cost (memory and notes) |
❌ | Generic advice in every session is noise | 10/14, more cost, lower quality: now off by default |
✅ | A conscience: a stronger model that speaks up after repeated mistakes | 44/63 vs 37/63, at ~37% of an always-on advisor's cost |
❌ | Generic property checks, contract ontology, rules learned from history | Precise but caught almost no failures on unseen repositories |
✅ | Executable review checks: facts about the change, not advice | 35/42 vs 32/42, 0 format regressions vs 10, +2% cost (3 repetitions) |
Exploratory results. The later experiments have 3 repetitions per arm and include repositories the designs had not seen; the earlier ones one run per cell. Methods and limits are in the paper.
🚀 Quick start
Requirements: Python 3.12+, git, and at least one of OMP, Claude Code or Codex.
1 · Install the tool
git clone https://github.com/abdullahbalabel/mihad.git
cd mihad
python -m pip install -e .2 · Add it to a project. Preview first, then install for your agent:
mihad-install D:/path/to/project --agent claude --dry-run
mihad-install D:/path/to/project --agent claude # or: --agent codex, --agent omp, or several3 · Work as usual. Open the project in your agent. To see what was learned and kept:
mihad report # lessons, verifier scripts, skills, competence, memory, dream cycles
mihad-memory list # what the memory holdsCodex loads a project's hooks only after you trust the project once in Codex. Claude Code needs no approval: the installer pre-approves the memory server for the project. See docs/AGENTS.md.
🤖 Agents
OMP | Claude Code | Codex | |
Verified memory (MCP) | ✅ | ✅ | ✅ |
Brief at each request | ✅ | ✅ | ✅ |
Live warnings after tools | ✅ | ✅ | ✅ |
Mandatory review before finishing | ✅ | ✅ | ✅ |
Learning from your commits | ✅ | ✅ | ✅ |
Dreaming with this agent | ✅ | ✅ | ✅ |
Wired through | extension | hooks | hooks |
Verification status for each check is in docs/AGENTS.md.
🧠 How it works
flowchart LR
U([You]) -->|request| A[Coding agent]
A -->|brief| B[(Memory and lessons)]
A -->|each tool result| D{Live detection}
D -->|warning| A
A -->|wants to finish| R{Review}
R -->|findings: one more turn| A
A -->|session log| E[Episodes]
U -->|your commit| E
E --> M[Mining: pitfalls, corrections, skills]
M --> V[Verifier scripts]
V --> P[Dreaming: practice with and without lessons]
P -->|A/B result| BEvery lesson has a gate before it is used:
Part | Kept only if |
Memory item | It passes the check for its kind: project tests, a quote from the starting commit, or your own words |
Failure-path lesson | Seen in two independent tasks, with a resolution that worked |
Co-change rule | Two independent tasks support it |
Verifier script | It fails before the fix and passes after it |
Skill | It succeeds on your final version of two past fixes |
Promotion | It wins an A/B test in practice without losing success |
Language | Function lookup | Test runners | Verified live |
Python | AST | pytest, unittest | ✅ |
JavaScript / TypeScript | Declaration scanner | node --test, Jest, Vitest, Mocha | ✅ |
Java / Kotlin | Declaration scanner | Maven, Gradle | analysis |
C# | Declaration scanner | dotnet test | analysis |
Go | Declaration scanner | go test | analysis |
Rust | Declaration scanner | cargo test | analysis |
PHP / Ruby | Declaration scanner | PHPUnit, RSpec, Minitest | analysis |
C / C++ | Declaration scanner | CTest, make test | analysis |
📚 Documentation
Guide | What's inside |
Requirements, install, uninstall, what is written where, safety | |
Daily workflow, commands, configuration, dreaming and its cost | |
OMP, Claude Code and Codex: setup, wiring, what has been verified | |
Every part, how it decides what to trust, and the evidence behind it | |
Modules, data files and the flow of a session | |
Findings at a glance, with links to the paper |
📄 Research
Better Coding Agents Without Retraining the Model: Verified Memory, Operational Experience and a Conscience. What Learns Is the System Around the Model, version 2.7 (Markdown · Word).
@techreport{balabel2026mihad,
author = {Balabel, Abdullah Mohammed},
title = {Adopting Reasoning Outputs Only After Verification: The {MIHAD} Architecture,
a Verified Memory and an Experience Engine for Coding Agents},
year = {2026},
month = {10},
note = {Version 2.1},
url = {https://github.com/abdullahbalabel/mihad}
}🌙 بالعربية
ذاكرة مهاد النمائية أداة تعطي وكيل البرمجة ذاكرة وخبرة تنموان مع الاستعمال، دون أن يتعلم أخطاءه ويكررها.
🛡️ ذاكرة لا تعتمد إلا ما له دليل مستقل: المهارة تُعتمد إن نجحت في اختبارات المشروع، والحقيقة إن اقتبست نص الكود كما كان قبل الجلسة، والتفضيل إن كان كلامك الحرفي.
🗣️ تفضيلاتك تُحفظ: تقولها مرة في أي رسالة، فتُلتقط بنصها وتُتبع في كل جلسة لاحقة.
🧭 موجز في بداية كل طلب، وتنبيه حي، ومراجعة إجبارية قبل الإنهاء: المراجعة تشغّل التغيير فعلًا، فتكشف الأسطر التي لا يغطيها أي اختبار، وضياع التنسيق، ومخالفة تفضيلاتك. إن وجدت مشكلة، يعود الوكيل للعمل.
🔁 يتعلم من تصحيحاتك: حين تعمل commit، تصبح نسختك هي النهائية، ويتعلم المحرك مما غيّرته بعد الوكيل.
🌙 الحلم: بين الجلسات يتدرب على نسخ معدلة من إصلاحات سابقة، بالدروس وبدونها، ولا يبقي درسًا إلا إن أثبتت التجربة فائدته.
🧩 يعمل مع OMP وClaude Code وCodex، وبعشر لغات برمجة، ويُثبَّت على أي مشروع بأمر واحد.
البحث الكامل في مجلد paper، والأدلة في مجلد docs.
⚖️ License
Copyright © 2026 Abdullah Mohammed Balabel.
Licensed under the PolyForm Noncommercial License 1.0.0. It is free for research, personal and other non-commercial use. Commercial use requires a separate licence from the author.
Available Tools
4 toolsmemory_correctB
Mark a memory item as wrong. It and everything derived from it is suspended.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| reason | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the important cascading effect that derived items are also suspended. However, it leaves key behavioral questions open: whether suspension is reversible, whether it requires special authorization, and whether 'suspended' means deleted or merely flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no wasted words; the consequential cascade behavior is stated immediately after the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no annotations and no output schema, the description covers the core action and the important derived-item consequence, but omits reversibility, permission requirements, and any indication of what the caller gets back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both 'id' and 'reason' are undocumented in structured fields and the description adds nothing about them. The parameter names are largely self-explanatory, but the description does not clarify what identifier format is expected or how the reason is used.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Mark ... as wrong') and resource ('a memory item'), which is clearly distinct from the recall/propose/report_outcome siblings. It does not explicitly name how it differs from those siblings, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus memory_report_outcome or the other siblings, and no prerequisites or conditions are stated. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_proposeA
Propose something worth remembering for future sessions. kind=skill: a reusable technique or code pattern, verified by running existing project tests. kind=fact: a fact about this codebase, verified by quoting a project file. kind=preference: how the user wants work done in future sessions; verify.quote must be the user's exact words. kind=lesson: a general lesson; stays provisional until a human approves it. Only adopted items count as verified knowledge.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| tags | No | ||
| title | Yes | ||
| verify | No | skill: {command, test_files} - tests must already exist unchanged in the git baseline. fact: {file, quote} - exact text in a project file that supports the fact. preference: {quote} - the user's exact words stating the preference. | |
| content | Yes | The knowledge itself, self-contained. | |
| supersedes | No | ID of an older item this replaces. | |
| derived_from | No | IDs of memory items this builds on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses lifecycle behavior ('kind=lesson... stays provisional until a human approves it', 'Only adopted items count as verified knowledge'), which goes beyond the schema. It says nothing about side effects, persistence timing, permissions, or what a proposal returns, leaving significant behavioral gaps for an un-annotated write tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, and the remaining semicolon-separated clauses are dense but each carries distinct kind semantics. There is some redundancy with the schema's verify description, which prevents a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with a nested object and no output schema, the description covers the essential model: the four kinds, how each is verified, and the adoption/provisional lifecycle. It omits relationship semantics (supersedes/derived_from) and any routing versus sibling memory tools, so it is solid but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 57% and the nested verify object already documents the per-kind shape ({command, test_files}, {file, quote}, {quote}). The description's verification sentences largely restate that schema text rather than adding syntax or format detail. Tags, title, supersedes, and derived_from get no elaboration in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Propose something worth remembering for future sessions') and then enumerates the four kinds (skill, fact, preference, lesson), so the agent knows exactly what the tool produces. It does not explicitly contrast itself with the siblings memory_recall, memory_correct, or memory_report_outcome, which keeps it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The per-kind clauses give clear conditions for choosing each kind ('kind=skill: a reusable technique or code pattern... kind=fact: a fact about this codebase...'), effectively telling the agent when each mode applies. However, there is no guidance on when to use this tool rather than memory_correct or memory_recall, and no exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_recallA
Search the project's persistent memory before starting work. Returns verified skills, facts and lessons from earlier sessions. Items marked provisional or unverified have NOT passed independent checks - treat them as hints only. 'preferences' are the user's standing working preferences: follow them.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| query | Yes | What you are about to work on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses the trust hierarchy (verified vs. provisional/unverified treated as hints only) and instructs that 'preferences' be followed. It does not state permission/auth requirements or rate limits, but for a read tool the output-semantics disclosure is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the action and timing, then return semantics and caveats. Dense with useful information and no filler, though the 'preferences' instruction could be folded more tightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description adequately explains what comes back (verified skills, facts, lessons, preferences) and how to treat unverified items. The gap is the unexplained 'k' parameter and any limits on result volume.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only 'query' has schema documentation (coverage 50%); 'k' is undocumented in both schema and description, and the description never mentions result count or how to control it. The phrase 'what you are about to work on' loosely aligns with query but adds little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Search') and resource ('the project's persistent memory') and clarifies the return contents (verified skills, facts, lessons). It does not explicitly contrast itself with the write-oriented siblings (memory_propose, memory_correct, memory_report_outcome), though the read/search nature is inherently distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Search the project's persistent memory before starting work' gives a clear triggering context. No when-not conditions or named alternatives are given, so it falls short of full 5-level routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_report_outcomeB
At the end of a task, report which memory items you used and whether they helped or caused problems.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| success | Yes | ||
| task_id | No | ||
| used_ids | No | ||
| harmful_ids | No | ||
| helpful_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It mentions reporting helpful or harmful items, but doesn't explain what happens with that data (e.g., is it used to update memory? are there consequences?). It doesn't state permissions, reversibility, or side effects. For a tool that likely influences a memory system, this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the timing and action. Every word earns its place; there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters with 0% schema description coverage, no annotations, and no output schema, the description is far too sparse. It doesn't clarify parameter usage, expected data formats, or behavioral consequences. It leaves the agent with significant ambiguity about how to correctly populate the fields (e.g., what constitutes 'used' vs 'helpful' vs 'harmful', how task_id relates, what notes should contain).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only vaguely alludes to a subset of parameters ('used memory items,' 'helped or caused problems'). With 6 parameters including success, notes, task_id, used_ids, helpful_ids, harmful_ids, the description does not explain the meaning, format, or relationships of these parameters. An agent would need to infer from parameter names alone, which is insufficient given the zero schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('report which memory items you used') and its timing ('at the end of a task'). It is reasonably distinct from siblings like memory_recall and memory_correct, though it doesn't explicitly name them. The purpose is clear: a feedback mechanism for memory usage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage condition: 'At the end of a task.' This tells the agent when to call it. It doesn't explicitly exclude other times or name alternatives, but the timing constraint is a strong usage guideline. No misleading guidance is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.2.0- First observed
memory_correct - First observed
memory_propose - First observed
memory_recall - First observed
memory_report_outcome
TDQS
Scored across 4 tools
Each tool targets a distinct action on memory items: recall (read), propose (create), correct (invalidate), and report_outcome (feedback). Their purposes do not overlap, and the descriptions clearly delineate when to use each.
All tools follow a consistent memory_<verb> snake_case pattern: memory_recall, memory_propose, memory_correct, memory_report_outcome. The convention is predictable and readable.
Four tools is well-scoped for a persistent memory interface, covering the essential agent interactions without redundancy. Each tool earns its place in the memory lifecycle.
The set covers read, create, invalidate, and feedback, but lacks explicit update or delete operations. Agents can work around this by proposing new items and marking old ones as wrong, so the gap is minor.
Maintenance
Related MCP Connectors
The project brain for AI coding agents — memory, decisions, sprints, knowledge base via MCP.
Long-term memory for AI coding agents: durable project facts, recalled by every MCP client.
Hosted MCP memory for coding agents: persistent across sessions, editable markdown, team sharing.
Give your AI agent persistent, governed memory for every project. At task start it recalls the approved decisions, conventions, risks and architecture (semantic search, ranked by importance); at close it proposes what was learned as typed memories that you review and approve — governance, not a notes dump. Agents propose, humans govern: edits go back to pending and deletion is human-only by design. Connect Claude Code, Cursor, Claude Desktop or any MCP client in two minutes with just your API key — hosted (nothing to install) or locally via `uvx solucortex-mcp`. Built by SoluAI and dogfooded daily: SoluCortex is developed using its own living memory.
Related MCP Servers
- AlicenseAqualityAmaintenanceMCP server that gives coding agents persistent, verified memory of codebase decisions, conventions, and skills, with evidence-based claims that are re-checked via git hooks and human-gated review. Enables memory search, propose/approve, chat harvesting, and critique across MCP-compatible tools.21108 npm1MIT
- FlicenseNot gradedqualityCmaintenanceProvides AI agents with a governed, three-layer project memory (guide, code facts, and knowledge) through namespaced MCP tools for code search, context compilation, impact analysis, and proposal-driven documentation updates.2 npm2-
- AlicenseNot gradedqualityCmaintenanceProvides coding agents with structured planning, persistent project memory, automated verification, and safety permission controls through MCP tools, enabling better planning, context retention, self-checking, and guarded execution.5 npmMIT
- FlicenseBqualityBmaintenanceProvides AI coding agents with a hybrid cognitive memory engine that reduces token usage, prevents amnesia, detects repetitive error loops, and retrieves relevant code context through MCP.6-