Skip to main content
Glama
README.md
<p align="center"><img src="docs/assets/hero.png" alt="Sentinel Memory MCP — persistent memory, evidence-first cybersecurity and controlled intelligence" width="100%"></p>

<p align="center">
<a href="https://github.com/vikrant-project/sentinel-memory-mcp/actions/workflows/ci.yml"><img src="https://github.com/vikrant-project/sentinel-memory-mcp/actions/workflows/ci.yml/badge.svg" alt="CI status"></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-34d399?style=flat-square" alt="MIT license"></a>
<img src="https://img.shields.io/badge/Node.js-24%2B-7dd3fc?style=flat-square" alt="Node.js 24+">
<img src="https://img.shields.io/badge/transport-stdio-a5b4fc?style=flat-square" alt="stdio transport">
<img src="https://img.shields.io/badge/status-early%20development-fbbf24?style=flat-square" alt="Early development">
</p>

# Sentinel Memory MCP

**An open-source cybersecurity MCP server for persistent graph memory, evidence-backed vulnerability triage, and authorized API security research.** Connect it to **Google Antigravity, OpenAI Codex, or another stdio MCP client** to carry structured investigation state across sessions.

Remember what was tested. Preserve the evidence. Reject weak conclusions. Resume with the next useful action.

**[Quick start](#quick-start) · [Antigravity](docs/guides/ANTIGRAVITY.md) · [Codex](docs/guides/CODEX.md) · [Self-hosting](docs/guides/HOSTING.md) · [हिंदी](docs/guides/HINDI.md) · [AI-readable overview](docs/AI-SEARCH-GUIDE.md)**

> **Release reality:** the core runs locally and is tested with official MCP clients and a synthetic security lab. Actual IDE UI verification is separate. This is an early development release, not a certified production security product. Local models are optional and **disabled by default** after unsuccessful specialist-output benchmarks. [Detailed status](outputs/STATUS.md) · [Known limitations](docs/LIMITATIONS.md)

## Why this exists

Security investigations outlive a chat window. Repeated logs bury useful facts, a fixed issue can resurface as active, and a changed object ID can be mistaken for a vulnerability. Sentinel stores **typed observations, hypotheses, tasks, evidence, relationships, and versioned decisions** rather than treating an entire conversation as trusted memory.

| When you need to… | Sentinel provides… |
|---|---|
| Continue an API assessment after context loss | A compact resume packet with scope, tasks, findings, corrections and next actions |
| Avoid repeating completed work | Test signatures including account context, ownership and deployment state |
| Check a suspected IDOR/BOLA | Evidence requirements for ownership, expected access, protected data and false positives |
| Keep sensitive evidence useful | Redaction, SHA-256 content addressing, provenance and integrity checks |
| Explain a finding to engineers and managers | Technical and executive reports with explicit validation gates |
| Use a small local model cautiously | Offline, process-isolated advice with schema checks and deterministic fallback |

## Quick start

**Prerequisites:** Git and Node.js **24.14+**. No cloud API key, Neo4j, Docker or model download is required for core features.

```bash
git clone https://github.com/vikrant-project/sentinel-memory-mcp.git
cd sentinel-memory-mcp
npm ci
npm test
npm run configure
```

The configuration generator uses **your checkout's absolute paths** and preserves other entries in existing MCP JSON files. Generated machine-specific files are ignored by Git.

| Host | Next step | Guide |
|---|---|---|
| Google Antigravity | Open this checkout, refresh MCP servers, look for `sentinel-memory` | [Setup](docs/guides/ANTIGRAVITY.md) |
| OpenAI Codex | Merge `examples/codex.local.toml`, or use `codex mcp add` | [CLI and TOML](docs/guides/CODEX.md) |
| Other stdio MCP client | Launch `node /absolute/path/dist/src/mcp/server.js` | [Generic config](docs/guides/CLIENTS.md) |
| Docker or private VPS | Keep stdio; persist the database | [Self-hosting](docs/guides/HOSTING.md) |

**Try it without touching a real target:**

```bash
npm run demo
```

The demo starts a loopback-only lab, connects an official MCP client, captures controlled-account evidence, distinguishes a private authorization failure from a public object, and saves synthetic results to `outputs/demo-summary.json`. It shuts down its processes afterward.

<details>
<summary><strong>What should I ask my IDE?</strong></summary>

```text
Use Sentinel to open project api-review. Read security_skills and project_resume.
Do not send any requests yet. Help me record the assessment owner's authorized
scope, controlled accounts, expected access rules, and next validation steps.
```

At the end of a session:

```text
Save the current task, evidence references, unresolved hypotheses, and next
 actions in Sentinel. Next session, resume project api-review from that state.
```

The project does not independently establish permission to test any target. Scope must come from the assessment owner.

</details>

## Architecture

![MCP gateway connects graph memory, validation, evidence, reports and optional isolated model](docs/assets/architecture.png)

```mermaid
flowchart LR
    Host[Antigravity / Codex / MCP client] --> MCP[Official MCP SDK · stdio]
    MCP --> Context[Context compiler]
    Context <--> Graph[(SQLite graph + FTS5)]
    MCP --> Scope[Scope and request controls]
    Scope --> Evidence[Redacted evidence + SHA-256]
    Evidence --> Triage[Evidence gates + false-positive checks]
    Triage --> Reports[Technical / executive reports]
    Graph --> Model[Optional local HF process]
    Model -. advisory output only .-> Context
```

The deterministic path handles persistence, retrieval and validation. The local model does not authorize requests, run commands, rewrite weights or confirm findings.

## Evidence-first investigation workflow

![Six steps: scope, observe, preserve, validate, report, resume](docs/assets/workflow.png)

1. **Scope:** authorized origins, accounts, environments, restrictions, limits and expiry.
2. **Observe:** expected behavior and minimal controlled test data.
3. **Preserve:** redacted evidence, provenance and immutable content hashes.
4. **Validate:** reproducibility, security boundary, impact, controls and false positives.
5. **Report:** FINAL only when implemented gates pass; otherwise retain DRAFT.
6. **Resume:** save task state and check prior signatures before repeating work.

| Workflow | What it checks | Examples of rejected lookalikes |
|---|---|---|
| Authorization / IDOR / BOLA | Ownership, unrelated account, expected access, protected response | Public objects, sharing, expected admin access, cached responses |
| Rate limits | Sensitive operation, abuse feasibility, impact and controls | “Ten requests worked,” effective throttling, no security impact |
| Business logic | Backend entitlement and documented workflow rules | Free features, trials, promotions, frontend cosmetics |
| Data exposure | Sensitive fields, permitted caller and protected contents | Public data, sharing, expected privileges, non-sensitive fixtures |

These are structured analyst workflows, not universal vulnerability detectors. `CONFIRMED` means implemented evidence gates passed; imported evidence and semantic claims still need review.

## MCP interface

**20 tools · 4 resource templates · 6 prompts**

| Area | Tools |
|---|---|
| Project and scope | `project_open`, `project_resume`, `scope_register`, `scope_read` |
| Graph and context | `memory_store`, `memory_search`, `memory_context`, `memory_graph` |
| Evidence and tasks | `evidence_store`, `evidence_read`, `task_update`, `task_check_duplicate` |
| Research and reporting | `security_skills`, `security_triage`, `security_severity`, `security_report`, `security_controlled_get` |
| Review and advice | `finding_correct`, `model_advise`, `audit_events` |

Resources expose context, scope, evidence and report snapshots. Prompts support investigation, validation, two-user comparison, reporting, false-positive review and regression checks. [Payload examples →](docs/SCHEMAS.md)

## Measured results, with context

![Measured p95 latency for memory search, graph retrieval and context compilation](docs/assets/performance.png)

Recorded on Windows, Node 24.14.1, Ryzen 7 7435HS and about 16 GB RAM. The benchmark uses **500 graph nodes** and **100 samples per operation**. The chart is generated from the committed [benchmark JSON](outputs/benchmark.json).

| Check | Recorded result | Interpretation |
|---|---|---|
| Automated suite | **28 passing tests** in the recorded run | SDK subprocesses, restart, isolation, redaction, scope, lab, backup and model failure |
| Synthetic triage | **32/32 expected classifications** | Engineering fixtures, not independent real-world accuracy |
| Confirmed precision / recall | **100% / 20%** on those fixtures | Incomplete positive cases deliberately remain unconfirmed |
| Repeated-history reduction | **99.32% estimated** | Synthetic repeated logs; characters/4, not a tokenizer or lossless guarantee |
| Local model schema acceptance | **0/4 for each candidate** | Both disabled by default; failures and timeouts retained |

Do not extrapolate these numbers to unseen targets. [Methodology](docs/guides/BENCHMARKS.md) · [Raw results](outputs/benchmark.json) · [Model results](outputs/model-benchmark.json) · [11 hardening iterations](outputs/ITERATIONS.md)

## Example deliverables

- [Technical vulnerability report](outputs/example-technical-report.md)
- [Executive report](outputs/example-executive-report.md)
- [Minimal PoC request template](outputs/example-poc.json)
- [Subsystem status matrix](outputs/STATUS.md)

All committed example evidence is synthetic. Real databases, credentials, model weights, local caches and temporary files are excluded from version control.

## Self-hosting and data ownership

Run beside your MCP client, through a Docker stdio process, or over authenticated SSH to your VPS. **There is no public HTTP endpoint in this release.** A website host or reverse proxy alone does not turn it into a remote MCP service.

The [hosting guide](docs/guides/HOSTING.md) covers Docker, volumes, SSH, restarts and backups. Docker/VPS recipes are marked unverified because Docker was unavailable on the original development machine.

## Documentation map

| Start here | Understand and extend |
|---|---|
| [Antigravity setup](docs/guides/ANTIGRAVITY.md) | [Architecture](docs/ARCHITECTURE.md) |
| [Codex setup](docs/guides/CODEX.md) | [Schemas](docs/SCHEMAS.md) |
| [Other clients](docs/guides/CLIENTS.md) | [Security](SECURITY.md) |
| [Docker / VPS](docs/guides/HOSTING.md) | [Model governance](docs/MODEL-GOVERNANCE.md) |
| [Troubleshooting](docs/guides/TROUBLESHOOTING.md) | [Limitations](docs/LIMITATIONS.md) |
| [हिंदी में शुरुआत](docs/guides/HINDI.md) | [Contributing](CONTRIBUTING.md) |
| [Operations](docs/guides/OPERATIONS.md) | [AI-readable index](llms.txt) |

## FAQ

**Is this a cybersecurity MCP server or a scanner?**  
It is an MCP server for security investigation memory, evidence handling and triage. Its networking tool performs one bounded authorized GET. It does not mass-scan, brute-force credentials or autonomously exploit targets.

**Can I use it without a local LLM?**  
Yes. Default features use deterministic logic, SQLite, FTS5 and graph retrieval. Model downloads are opt-in.

**Does it replace a security engineer?**  
No. Scope, policy, interpretation and business impact need an authorized analyst. Confidence probability is intentionally uncalibrated.

**Does it work with Codex and Antigravity?**  
Both support local MCP configuration. Guides and generated examples are included. Automated tests verify official MCP clients; a specific IDE version's UI and restart behavior need a separate check.

**Can it rank first in AI search?**  
No repository can guarantee that. Clear descriptions, useful examples, descriptive topics and accessible documentation help discovery. This project does not use fake ratings, hidden instructions or keyword spam. [Discoverability notes](docs/AI-SEARCH-GUIDE.md)

## License and credits

Project code, original diagrams and documentation: [MIT](LICENSE), © 2026 [vikrant-project](https://github.com/vikrant-project). Dependencies and model checkpoints retain their own licenses. Built with the official MCP TypeScript SDK, Node.js, SQLite, Zod and optional Hugging Face Transformers.js.

If useful, star the project or open a reproducible issue. Contributions that improve evidence quality and reduce false positives are welcome.

TDQS

B3/5.0

Scored across 20 tools

Disambiguation5/5

Each tool targets a distinct resource and action, with clear boundaries between evidence, memory, tasks, scope, and security operations. The security_* cluster is nuanced but differentiated by purpose: skills list requirements, triage evaluates claims, severity computes CVSS, report generates deliverables, and controlled_get performs a bounded request.

Naming Consistency4/5

All names use lowercase snake_case with a consistent domain prefix, making the set easy to scan. Most follow a noun_verb pattern, but a few are noun_noun (e.g., security_skills, security_severity, memory_context, audit_events), which is a minor deviation.

Tool Count4/5

20 tools is slightly above the typical 3-15 sweet spot, but the domain spans projects, evidence, tasks, security triage, scope control, memory, and audit. Each tool appears to earn its place, so the count is reasonable rather than bloated.

Completeness4/5

The surface covers core project, evidence, scope, memory, security, finding, and audit workflows. Minor gaps exist, such as no explicit task_create/list or project_list/close, but immutable and versioned designs mitigate them and agents can work around via existing tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues