Skip to main content
Glama

Agent Guardian

AI coding agents forget.

After context compaction they can repeat rejected ideas, redo completed work, lose the original goal and declare unfinished work complete.

Agent Guardian gives them persistent project memory.

Problem

LLM-based coding agents operate in ephemeral contexts. When context is compacted or a session ends, critical project state is lost:

  • Rejected Ideas: The agent may propose a solution that was previously rejected for good reason.

  • Completed Work: The agent may attempt to redo tasks that are already finished.

  • Goal Drift: The agent may lose sight of the original objective, leading to scope creep or tangential work.

  • False Completion: The agent may declare a task complete without verifying that acceptance criteria are met.

Related MCP server: ITHZ MCP

Solution

Agent Guardian is a lightweight, local-first persistence layer for AI coding agents. It provides a "guardian" process that agents can query before taking actions. It tracks:

  • Rejection Memory: A log of rejected ideas with reasons, preventing repetition.

  • State Ledger: A structured log of project state (goals, plans, completed items, blockers).

  • Assumption Ledger: Tracks assumptions made by the agent, their verification status, and staleness.

  • Drift Detector: Monitors agent actions to detect deviation from the stated goal.

  • Loop Detector: Identifies repetitive failure patterns (loops).

  • Completion Verifier: Validates that tasks are actually complete based on defined criteria and evidence.

Architecture

+----------------+       +-------------------------------------------------------+       +----------------+
|                |       |                   Agent Guardian                      |       |                |
|   Coding       |------>|  [Rejection Memory] [State Ledger] [Assumption Ledger]|------>|   ALLOW        |
|   Agent        |       |  [Drift Detector]   [Loop Detector] [Completion Verifier]|    |   WARN         |
|                |       |                                                       |       |   BLOCK        |
+----------------+       +-------------------------------------------------------+       +----------------+

Storage layout:

.guardian/
├── guardian.db   # SQLite database
├── config.yaml   # Configuration
└── project.md    # Project notes

Installation

Requires Python 3.11+. No cloud, no API keys.

# From a clone
git clone https://github.com/niallsemple/agent-guardian.git
cd agent-guardian
pip install .

# With PyYAML support (optional, for better config parsing)
pip install ".[yaml]"

Quick start

# Initialize a new project
guardian init

# Set the project goal
guardian goal "Build a REST API for user management"

# Record a rejected idea
guardian reject "Use MongoDB" --reason "We need ACID transactions"

# Check a proposal
guardian check "Use PostgreSQL for the database"

# Get context for the agent
guardian context

CLI examples

State

guardian goal "Build a REST API"
guardian completed "Design the user schema"
guardian next "Implement the /users endpoint"
guardian status
guardian history --field goal

Reject/Reopen

guardian reject "Use GraphQL" --reason "REST is simpler for this use case"
guardian reopen rej_001 --reason "Actually, GraphQL is better"

Check

guardian check --json "Use MongoDB"
{
  "decision": "BLOCK",
  "risk": 0.92,
  "reasons": [
    {
      "module": "rejection_memory",
      "decision": "BLOCK",
      "risk": 0.92,
      "message": "Similar to rejected idea: Use MongoDB"
    }
  ],
  "proposal": "Use MongoDB"
}

Loops

guardian record test_fail src/app.ts "TypeError: x is undefined"
guardian loops

Assumptions

guardian assumption add "The API will be stateless" --confidence 0.8
guardian assumption verify ass_001
guardian assumption stale

Drift

guardian drift "Implementing user authentication"

Example (goal: "Build a REST API for user management"):

$ guardian drift "Use MongoDB for the database"
ON_TASK — activity matches the goal.
$ guardian drift "Rewrite the frontend in Svelte"
POSSIBLE_DRIFT — ... DRIFT SCORE: 1.00. Evidence: different layer than the goal: frontend, svelte
$ guardian drift "Build a mobile game UI"
HIGH_DRIFT — ... DRIFT SCORE: 0.87. Evidence: out of scope for this goal: game; different layer than the goal: mobile, ui
$ guardian check "Use MongoDB for the database"
ALLOW (risk 0.00)

With a goal of "Add rate limiting to the payments API (backend only)" and criterion "No frontend changes", guardian check "Rewrite the frontend in Svelte" gives WARN with HIGH_DRIFT ... contradicts a scope constraint: 'frontend' is excluded by the goal.

Contract/Evidence/Verify

guardian contract "Build user API" --tests "pytest -q" --path src/users.py
guardian evidence behaviour_verified --pass --evidence "manual curl of /users returned 200"
guardian verify

Override

guardian override --reason "Need to use MongoDB for this specific feature" --proposal "Use MongoDB"

Audit

guardian audit --limit 10

Export

guardian export --output guardian-export.json

Adapter

guardian adapter claude

MCP

guardian mcp

Examples

Real output from the CLI (a session with goal "Build a REST API for user management").

BLOCK — a rejected idea proposed again after context was cleared (guardian check "Use MongoDB for the database", exit 2)

BLOCK (risk 1.00)
[rejection_memory] BLOCK — This proposal appears substantially similar to a previously rejected approach.
Previous idea: Use MongoDB for the database
Reason rejected: We need ACID transactions
Rejected: 2026-10-04
Similarity: 1.00
[state] State ledger: 'Use MongoDB for the database' is marked rejected.

LOOP DETECTED (guardian loops, exit 2, after the same fix failed four times)

LOOP DETECTED — The agent has attempted substantially similar fixes 4 times.
Repeated failure: src/session.ts: TypeError: Cannot read properties of undefined (reading 'id') line 44
Previous attempts:
1. Add a null check before reading session.user.id
2. Add a null check before reading session.user.id
3. Add a null check before reading session.user.id
4. Add a null check before reading session.user.id
The next attempt must be materially different.
Suggested action: Reassess the underlying assumption.

Stale assumption warning (guardian check "Increase the API requests per minute in src/api/client.ts", exit 1)

WARN (risk 0.73)
[assumptions] WARNING — This action relies on a stale assumption.
Claim: API allows 100 requests per minute
Last verified: 64 days ago
Recommendation: verify before continuing.

Completion checklist (guardian verify while a test is failing, exit 2)

Completion check: Add an add() helper
✓ feature_exists — All paths exist: calc.py
✗ tests_pass — pytest: exit 1: 1 failed in 0.01s
✓ build_passes — python: exit 0
✓ no_placeholder_code — no placeholder markers found
✓ requested_output_exists — All paths exist: calc.py (optional)
✗ behaviour_verified — no evidence provided (optional)
TASK STATUS: NOT COMPLETE

Once the test passes, the last line becomes TASK VERIFIED COMPLETE and the exit code is 0.

Context block for an agent after compaction (guardian context)

AGENT GUARDIAN PROJECT CONTEXT
Goal: Build a REST API for user management
Completed:
- Design the user schema
Rejected:
- Use MongoDB for the database (rej_001) — We need ACID transactions
Important assumptions:
- API allows 100 requests per minute [STALE] (verified 64d ago)
Next action: Implement the /users endpoint
Do not repeat rejected approaches without new evidence.

Configuration

Default config.yaml:

guardian:
  matcher: keyword
  rejection_memory:
    warn_threshold: 0.70
    block_threshold: 0.85
  loop_detection:
    warning_after: 2
    block_after: 4
  drift:
    warning_threshold: 0.65
    high_threshold: 0.85
    block_on_high: false
  assumptions:
    stale_after_days: 30
  completion:
    require_tests: true
    require_build: true

How drift scoring works

Drift scoring is offline and deterministic (no models, no new dependencies):

  1. Goal context: goal, success criteria, plan, next action and the last few completed items.

  2. Lexical relevance: the configured matcher compares the activity with that context, ignoring generic engineering words ("build", "fix", "test", "refactor") so they cannot make unrelated work look relevant.

  3. Domain evidence (agent_guardian/drift_domains.py): each content word is mapped to a domain (api, storage, auth, data, frontend, mobile, game, ml, devops, ...). Its relation to the goal's domains is same, supports (a database or auth layer for an API), adjacent (a frontend for an API) or foreign (a game for an API). Files recently touched for the goal count as support only when the activity names them.

  4. Scope constraints: "backend only", "only X", "no frontend changes", "without X" in the goal/criteria are parsed; contradicting them is always HIGH drift.

  5. Classification: HIGH_DRIFT needs strong evidence (foreign or constraint-violating work, or several unrelated words with no support); POSSIBLE_DRIFT needs the score above warning_threshold plus some evidence; short or vague activities default to ON_TASK. Results include level (LOW/MEDIUM/HIGH) and an evidence breakdown that also appears in the message.

guardian check reports drift only at MEDIUM (POSSIBLE_DRIFT) or above; on-task work adds nothing. Add your own domains with register_domain(name, terms, supports=(), adjacent=(), names=()) and generic vocabulary with register_neutral(words).

The labelled evaluation set lives in tests/drift_eval_data.py (45 cases over 4 goals, plus 18 held-out cases). Run python scripts/drift_eval.py -v to see precision/recall and per-case results.

Supported agents

  • Generic: Shell/JSON adapter for any agent (Kimi, Gemini CLI, OpenCode, OpenClaw)

  • Claude Code: PreToolUse hook

  • Codex: AGENTS.md integration

  • Cursor: Rules integration

  • MCP: Optional Model Context Protocol server

Exit codes

  • 0: ALLOW / OK

  • 1: WARN

  • 2: BLOCK

  • 3: Error / Not initialised

Roadmap

Phase 5:

  • Advanced semantic matching via pluggable local embeddings (register_matcher() in agent_guardian/similarity.py is the extension point)

  • Grow the drift domain vocabulary and evaluation set (register_domain() in agent_guardian/drift_domains.py)

  • Richer adapters

  • Auto-capture of agent actions

How this was built

Agent Guardian was built step by step by a local Qwen model (via aider), running offline against a local OpenAI-compatible server, supervised by an AI agent that reviewed every step:

  • SPEC.md is the original specification.

  • steps/ holds the step-by-step prompts (01 to 26) given to the model, one feature per step.

  • run_build.sh is the build driver: for each step it runs aider, then the tests, and feeds failures back for a bounded number of fix attempts. launch_detached.sh starts it in the background.

  • Each step was committed only after its tests passed; commits note where a step was hand-fixed. The drift detection rework (drift.py, drift_domains.py, evaluation set) was written by hand, not by the model.

  • scripts/dod_check.sh is the definition-of-done check run at the end.

Contributing

  • Use pytest for testing

  • Keep it deterministic and offline

  • Small PRs preferred

Licence

MIT © 2026 Niall Semple

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Local-first deterministic project memory for AI coding agents, with context packs, decisions, gates, risks, scoped claims and explicit checkpoints in project-owned files.
    -
  • A
    license
    A
    quality
    C
    maintenance
    Local-first project memory for AI coding agents. Records failed attempts, fragile files, and decisions per repo, and warns the agent via hooks before it repeats a recorded mistake.
    6
    59 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Self-hosted decision memory for AI coding agents. Captures decisions with the alternatives you rejected, and warns before an agent re-proposes a rejected approach.
    4
    81
    Apache 2.0