Skip to main content
Glama

💡 Why MIHAD

Coding agents start every session like a new hire on day one. They forget yesterday's corrections and repeat the same mistakes. Saving everything they "learn" is not the answer either: an agent that remembers its own wrong conclusions repeats them with confidence.

MIHAD gives an agent a memory and a body of experience that grow with use, and it adopts nothing without independent evidence. It is built on a research programme with pinned protocols, real repository history and a blind quality judge. Its measured gains are largest for cheaper models: with a strong model as the agent, success was the same with and without it (51/63 vs 50/63; on the sound tasks of that benchmark, 45/45 either way). See the paper.

Related MCP server: Code Project Brain

✨ What it does

🛡️ Verified memory

A skill must pass your project's own tests. A fact must quote the code as it was before the session. A preference must be your own words. Everything else stays provisional, and corrections reach everything built on an item.

🗣️ Your preferences, kept

Say "Always add a regression test" once, in any message. It is captured verbatim, adopted, and followed in every later session. One-off instructions are never stored.

🧭 Brief, live warnings, review

At the start of each request the agent gets a short brief. While it works, known pitfalls are flagged the moment they recur. Before it finishes, its change is reviewed and executed: lines no test covers, lost formatting, broken preferences. Findings send it back to work.

🔁 Learns from your corrections

Every session is recorded. When you commit, your commit is the final version, and what you changed after the agent becomes experience. There is nothing extra to do.

🌙 Dreaming with A/B tests

Between sessions it practises on slightly broken copies of past fixes, with and without its lessons. A lesson is kept only if the test shows it helps.

🧩 Works where you work

OMP, Claude Code (terminal and desktop) and Codex. Ten programming languages. Standard-library Python with no dependencies. One command installs it into a project.

📊 What the experiments showed

Finding

Evidence

✅

Captured preferences carried into every later task

15/15 vs 0/15 without capture

✅

Wrong items planted in memory were never followed

Two contaminated-memory runs on real commits

✅

Operational experience transferred to new tasks

Same success (11/14), 14.4% fewer tokens, quality 3.93 vs 3.79

❌

Knowledge about the code did not transfer

Same success, +16–17% cost (memory and notes)

❌

Generic advice in every session is noise

10/14, more cost, lower quality: now off by default

✅

A conscience: a stronger model that speaks up after repeated mistakes

44/63 vs 37/63, at ~37% of an always-on advisor's cost

❌

Generic property checks, contract ontology, rules learned from history

Precise but caught almost no failures on unseen repositories

✅

Executable review checks: facts about the change, not advice

35/42 vs 32/42, 0 format regressions vs 10, +2% cost (3 repetitions)

Exploratory results. The later experiments have 3 repetitions per arm and include repositories the designs had not seen; the earlier ones one run per cell. Methods and limits are in the paper.

🚀 Quick start

Requirements: Python 3.12+, git, and at least one of OMP, Claude Code or Codex.

1 · Install the tool

git clone https://github.com/abdullahbalabel/mihad.git
cd mihad
python -m pip install -e .

2 · Add it to a project. Preview first, then install for your agent:

mihad-install D:/path/to/project --agent claude --dry-run
mihad-install D:/path/to/project --agent claude      # or: --agent codex, --agent omp, or several

3 · Work as usual. Open the project in your agent. To see what was learned and kept:

mihad report          # lessons, verifier scripts, skills, competence, memory, dream cycles
mihad-memory list     # what the memory holds
TIP

Codex loads a project's hooks only after you trust the project once in Codex. Claude Code needs no approval: the installer pre-approves the memory server for the project. See docs/AGENTS.md.

🤖 Agents

OMP

Claude Code

Codex

Verified memory (MCP)

✅

✅

✅

Brief at each request

✅

✅

✅

Live warnings after tools

✅

✅

✅

Mandatory review before finishing

✅

✅

✅

Learning from your commits

✅

✅

✅

Dreaming with this agent

✅

✅

✅

Wired through

extension

hooks

hooks

Verification status for each check is in docs/AGENTS.md.

🧠 How it works

flowchart LR
    U([You]) -->|request| A[Coding agent]
    A -->|brief| B[(Memory and lessons)]
    A -->|each tool result| D{Live detection}
    D -->|warning| A
    A -->|wants to finish| R{Review}
    R -->|findings: one more turn| A
    A -->|session log| E[Episodes]
    U -->|your commit| E
    E --> M[Mining: pitfalls, corrections, skills]
    M --> V[Verifier scripts]
    V --> P[Dreaming: practice with and without lessons]
    P -->|A/B result| B

Every lesson has a gate before it is used:

Part

Kept only if

Memory item

It passes the check for its kind: project tests, a quote from the starting commit, or your own words

Failure-path lesson

Seen in two independent tasks, with a resolution that worked

Co-change rule

Two independent tasks support it

Verifier script

It fails before the fix and passes after it

Skill

It succeeds on your final version of two past fixes

Promotion

It wins an A/B test in practice without losing success

Language

Function lookup

Test runners

Verified live

Python

AST

pytest, unittest

✅

JavaScript / TypeScript

Declaration scanner

node --test, Jest, Vitest, Mocha

✅

Java / Kotlin

Declaration scanner

Maven, Gradle

analysis

C#

Declaration scanner

dotnet test

analysis

Go

Declaration scanner

go test

analysis

Rust

Declaration scanner

cargo test

analysis

PHP / Ruby

Declaration scanner

PHPUnit, RSpec, Minitest

analysis

C / C++

Declaration scanner

CTest, make test

analysis

📚 Documentation

Guide

What's inside

Installation

Requirements, install, uninstall, what is written where, safety

Usage

Daily workflow, commands, configuration, dreaming and its cost

Agents

OMP, Claude Code and Codex: setup, wiring, what has been verified

Capabilities

Every part, how it decides what to trust, and the evidence behind it

Architecture

Modules, data files and the flow of a session

Research

Findings at a glance, with links to the paper

📄 Research

Better Coding Agents Without Retraining the Model: Verified Memory, Operational Experience and a Conscience. What Learns Is the System Around the Model, version 2.7 (Markdown · Word).

@techreport{balabel2026mihad,
  author  = {Balabel, Abdullah Mohammed},
  title   = {Adopting Reasoning Outputs Only After Verification: The {MIHAD} Architecture,
             a Verified Memory and an Experience Engine for Coding Agents},
  year    = {2026},
  month   = {10},
  note    = {Version 2.1},
  url     = {https://github.com/abdullahbalabel/mihad}
}

🌙 بالعربية

ذاكرة مهاد النمائية أداة تعطي وكيل البرمجة ذاكرة وخبرة تنموان مع الاستعمال، دون أن يتعلم أخطاءه ويكررها.

  • 🛡️ ذاكرة لا تعتمد إلا ما له دليل مستقل: المهارة تُعتمد إن نجحت في اختبارات المشروع، والحقيقة إن اقتبست نص الكود كما كان قبل الجلسة، والتفضيل إن كان كلامك الحرفي.

  • 🗣️ تفضيلاتك تُحفظ: تقولها مرة في أي رسالة، فتُلتقط بنصها وتُتبع في كل جلسة لاحقة.

  • 🧭 موجز في بداية كل طلب، وتنبيه حي، ومراجعة إجبارية قبل الإنهاء: المراجعة تشغّل التغيير فعلًا، فتكشف الأسطر التي لا يغطيها أي اختبار، وضياع التنسيق، ومخالفة تفضيلاتك. إن وجدت مشكلة، يعود الوكيل للعمل.

  • 🔁 يتعلم من تصحيحاتك: حين تعمل commit، تصبح نسختك هي النهائية، ويتعلم المحرك مما غيّرته بعد الوكيل.

  • 🌙 الحلم: بين الجلسات يتدرب على نسخ معدلة من إصلاحات سابقة، بالدروس وبدونها، ولا يبقي درسًا إلا إن أثبتت التجربة فائدته.

  • 🧩 يعمل مع OMP وClaude Code وCodex، وبعشر لغات برمجة، ويُثبَّت على أي مشروع بأمر واحد.

البحث الكامل في مجلد paper، والأدلة في مجلد docs.

⚖️ License

Copyright © 2026 Abdullah Mohammed Balabel.

Licensed under the PolyForm Noncommercial License 1.0.0. It is free for research, personal and other non-commercial use. Commercial use requires a separate licence from the author.

Available Tools

4 tools
memory_correctB

Mark a memory item as wrong. It and everything derived from it is suspended.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
reasonYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the important cascading effect that derived items are also suspended. However, it leaves key behavioral questions open: whether suspension is reversible, whether it requires special authorization, and whether 'suspended' means deleted or merely flagged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no wasted words; the consequential cascade behavior is stated immediately after the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no annotations and no output schema, the description covers the core action and the important derived-item consequence, but omits reversibility, permission requirements, and any indication of what the caller gets back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so both 'id' and 'reason' are undocumented in structured fields and the description adds nothing about them. The parameter names are largely self-explanatory, but the description does not clarify what identifier format is expected or how the reason is used.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Mark ... as wrong') and resource ('a memory item'), which is clearly distinct from the recall/propose/report_outcome siblings. It does not explicitly name how it differs from those siblings, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus memory_report_outcome or the other siblings, and no prerequisites or conditions are stated. Usage is only implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_proposeA

Propose something worth remembering for future sessions. kind=skill: a reusable technique or code pattern, verified by running existing project tests. kind=fact: a fact about this codebase, verified by quoting a project file. kind=preference: how the user wants work done in future sessions; verify.quote must be the user's exact words. kind=lesson: a general lesson; stays provisional until a human approves it. Only adopted items count as verified knowledge.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
tagsNo
titleYes
verifyNoskill: {command, test_files} - tests must already exist unchanged in the git baseline. fact: {file, quote} - exact text in a project file that supports the fact. preference: {quote} - the user's exact words stating the preference.
contentYesThe knowledge itself, self-contained.
supersedesNoID of an older item this replaces.
derived_fromNoIDs of memory items this builds on.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It usefully discloses lifecycle behavior ('kind=lesson... stays provisional until a human approves it', 'Only adopted items count as verified knowledge'), which goes beyond the schema. It says nothing about side effects, persistence timing, permissions, or what a proposal returns, leaving significant behavioral gaps for an un-annotated write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, and the remaining semicolon-separated clauses are dense but each carries distinct kind semantics. There is some redundancy with the schema's verify description, which prevents a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with a nested object and no output schema, the description covers the essential model: the four kinds, how each is verified, and the adoption/provisional lifecycle. It omits relationship semantics (supersedes/derived_from) and any routing versus sibling memory tools, so it is solid but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 57% and the nested verify object already documents the per-kind shape ({command, test_files}, {file, quote}, {quote}). The description's verification sentences largely restate that schema text rather than adding syntax or format detail. Tags, title, supersedes, and derived_from get no elaboration in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Propose something worth remembering for future sessions') and then enumerates the four kinds (skill, fact, preference, lesson), so the agent knows exactly what the tool produces. It does not explicitly contrast itself with the siblings memory_recall, memory_correct, or memory_report_outcome, which keeps it just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The per-kind clauses give clear conditions for choosing each kind ('kind=skill: a reusable technique or code pattern... kind=fact: a fact about this codebase...'), effectively telling the agent when each mode applies. However, there is no guidance on when to use this tool rather than memory_correct or memory_recall, and no exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_recallA

Search the project's persistent memory before starting work. Returns verified skills, facts and lessons from earlier sessions. Items marked provisional or unverified have NOT passed independent checks - treat them as hints only. 'preferences' are the user's standing working preferences: follow them.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNo
queryYesWhat you are about to work on.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses the trust hierarchy (verified vs. provisional/unverified treated as hints only) and instructs that 'preferences' be followed. It does not state permission/auth requirements or rate limits, but for a read tool the output-semantics disclosure is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the action and timing, then return semantics and caveats. Dense with useful information and no filler, though the 'preferences' instruction could be folded more tightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description adequately explains what comes back (verified skills, facts, lessons, preferences) and how to treat unverified items. The gap is the unexplained 'k' parameter and any limits on result volume.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only 'query' has schema documentation (coverage 50%); 'k' is undocumented in both schema and description, and the description never mentions result count or how to control it. The phrase 'what you are about to work on' loosely aligns with query but adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Search') and resource ('the project's persistent memory') and clarifies the return contents (verified skills, facts, lessons). It does not explicitly contrast itself with the write-oriented siblings (memory_propose, memory_correct, memory_report_outcome), though the read/search nature is inherently distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Search the project's persistent memory before starting work' gives a clear triggering context. No when-not conditions or named alternatives are given, so it falls short of full 5-level routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_report_outcomeB

At the end of a task, report which memory items you used and whether they helped or caused problems.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNo
successYes
task_idNo
used_idsNo
harmful_idsNo
helpful_idsNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It mentions reporting helpful or harmful items, but doesn't explain what happens with that data (e.g., is it used to update memory? are there consequences?). It doesn't state permissions, reversibility, or side effects. For a tool that likely influences a memory system, this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the timing and action. Every word earns its place; there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters with 0% schema description coverage, no annotations, and no output schema, the description is far too sparse. It doesn't clarify parameter usage, expected data formats, or behavioral consequences. It leaves the agent with significant ambiguity about how to correctly populate the fields (e.g., what constitutes 'used' vs 'helpful' vs 'harmful', how task_id relates, what notes should contain).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only vaguely alludes to a subset of parameters ('used memory items,' 'helped or caused problems'). With 6 parameters including success, notes, task_id, used_ids, helpful_ids, harmful_ids, the description does not explain the meaning, format, or relationships of these parameters. An agent would need to infer from parameter names alone, which is insufficient given the zero schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('report which memory items you used') and its timing ('at the end of a task'). It is reasonably distinct from siblings like memory_recall and memory_correct, though it doesn't explicitly name them. The purpose is clear: a feedback mechanism for memory usage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear usage condition: 'At the end of a task.' This tells the agent when to call it. It doesn't explicitly exclude other times or name alternatives, but the timing constraint is a strong usage guideline. No misleading guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.2.0
    • First observedmemory_correct
    • First observedmemory_propose
    • First observedmemory_recall
    • First observedmemory_report_outcome

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct action on memory items: recall (read), propose (create), correct (invalidate), and report_outcome (feedback). Their purposes do not overlap, and the descriptions clearly delineate when to use each.

Naming Consistency5/5

All tools follow a consistent memory_<verb> snake_case pattern: memory_recall, memory_propose, memory_correct, memory_report_outcome. The convention is predictable and readable.

Tool Count5/5

Four tools is well-scoped for a persistent memory interface, covering the essential agent interactions without redundancy. Each tool earns its place in the memory lifecycle.

Completeness4/5

The set covers read, create, invalidate, and feedback, but lacks explicit update or delete operations. Agents can work around this by proposing new items and marking old ones as wrong, so the gap is minor.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    MCP server that gives coding agents persistent, verified memory of codebase decisions, conventions, and skills, with evidence-based claims that are re-checked via git hooks and human-gated review. Enables memory search, propose/approve, chat harvesting, and critique across MCP-compatible tools.
    21
    108 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides coding agents with structured planning, persistent project memory, automated verification, and safety permission controls through MCP tools, enabling better planning, context retention, self-checking, and guarded execution.
    5 npm
    MIT
  • F
    license
    B
    quality
    B
    maintenance
    Provides AI coding agents with a hybrid cognitive memory engine that reduces token usage, prevents amnesia, detects repetitive error loops, and retrieves relevant code context through MCP.
    6
    -