Skip to main content
Glama
README.md
# chaos-core-mcp

An MCP server where **the AI is the decision-making kernel, not a tool it exposes.** A calling
client (Claude, ChatGPT, Codex, whatever) doesn't enumerate low-level endpoints — it hands Chaos
Core an objective and lets the Cognitive Core reason about it, discover capabilities, plan,
check deterministic policy, execute, evaluate, and remember.

As of v0.2 the cognitive core is **transport-agnostic**. The same core, tools, policies, memory,
and capability registry are reachable two ways: over **stdio** for local MCP clients, and over
**Streamable HTTP** at `/mcp` for remote MCP clients such as Claude custom connectors.

```
                     CHAOS CORE
                         │
                  Cognitive Core
                         │
        ┌────────────────┴────────────────┐
        │                                 │
     stdio                         Streamable HTTP
        │                                 │
        ▼                                 ▼
 Local MCP clients                Remote MCP clients
                                     /mcp
```

There is no HTTP variant of the cognition. `src/transport/stdio.ts` and `src/transport/http.ts`
both call the single server factory `createChaosCoreServer()` — the transport is invisible to
the cognitive layer, and there are no `http_reason` / `remote_plan` duplicates.

## The Cognitive Core loop

```
objective
   ↓
context
   ↓
AI planning
   ↓
policy
   ↓
capability execution
   ↓
evaluation
   ↓
result
```

V1 exposes each stage as its own MCP tool, so every step stays inspectable and the calling AI
stays in control between stages:

| Tool | Purpose |
|---|---|
| `chaoscore_reason` | Analyze an objective + context before any plan exists (Intent Analyzer) |
| `chaoscore_plan` | Convert an objective into an ordered, capability-grounded plan |
| `chaoscore_execute` | Run a plan: policy check → capability selection → execution → evaluation |
| `chaoscore_inspect` | Read-only introspection: capabilities, policy, providers, memory, audit trail, session |
| `chaoscore_remember` | Persist a fact to durable Semantic Memory |
| `chaoscore_recall` | Retrieve from Semantic Memory |

Both transports serve this identical list — enforced by a test that lists tools over a real MCP
client on each transport and compares the definitions.

`core/brain.ts` also implements the *full* loop as one composable function (`runCognitiveCore`)
— objective straight through to result, with automatic replanning on step failure and an
immediate halt on `REQUIRE_APPROVAL`. It is **not** registered as an MCP tool in V1 (see
[V1 boundary](#v1-boundary)) but exists fully wired, ready to back a future `chaoscore_achieve`
tool without a rewrite.

## Architecture

```
src/
  index.ts                    transport dispatcher (stdio by default)
  config.ts                   the only file that reads process.env

  server/                     ← composition root; transport-independent
    create-server.ts          createRuntime() + createChaosCoreServer()
    register-tools.ts         the single definition of the V1 tool surface
    types.ts                  RuntimeServices / ChaosCoreDependencies
    schemas.ts                shared Zod schemas
    tools/                    reason plan execute inspect remember recall

  transport/                  ← the ONLY transport-aware code
    stdio.ts                  local subprocess transport (stdout reserved for JSON-RPC)
    http.ts                   Streamable HTTP at /mcp (stateful sessions)

  core/                       brain intent planner evaluator context types
  capabilities/               registry executor types + built-in/
  memory/                     store (factory) sqlite (impl) types (MemoryStore interface)
  policy/                     engine permissions approvals types
  providers/                  ai-provider (AIProvider interface) openai index
  state/                      session (Working Memory) execution (trace assembly)
  observability/              logger events audit
  util/                       to-structured
```

### Dependency injection, and what has which lifetime

`createRuntime()` builds the process-wide services **once**: config, capability registry, policy
engine, memory store, provider registry, audit log, logger. `createChaosCoreServer()` builds one
`McpServer` per MCP session on top of that runtime, adds a per-session `SessionState`, and
registers the tools with the combined container injected.

| Component | Lifetime | Consequence |
|---|---|---|
| memory, policy, capabilities, providers, audit | per **process** | A remote HTTP client and a local stdio client hitting the same process see the same state |
| `SessionState` (working memory: last plan/reasoning/trace) | per **MCP session** | A `plan_id` from one client can't be executed by another |

No core module imports the dependency container. `core/intent.ts`, `core/planner.ts`, and
`capabilities/executor.ts` each declare a narrow structural interface (`IntentDeps`,
`PlannerDeps`, `ExecutorDeps`) that the container happens to satisfy — so the core is testable
in isolation and genuinely unaware of the server and transport layers.

## Policy sits outside the AI

```
AI proposes action
      ↓
deterministic policy engine
      ↓
ALLOW / DENY / REQUIRE_APPROVAL
```

The model may propose any capability; `policy/engine.ts` decides, as a pure function of the
capability name and the operator-controlled policy file. No model is consulted. Split into:

- `policy/permissions.ts` — allow/deny lists (`allowedCapabilities`, `deniedCapabilities`)
- `policy/approvals.ts` — which allowed capabilities still need a human (`requireConfirmationFor`)
- `policy/engine.ts` — composes them, plus bounded resources (`httpAllowedDomains`)

`data/policy.json` is auto-created with safe defaults on first run:

```json
{
  "allowedCapabilities": [],
  "deniedCapabilities": [],
  "requireConfirmationFor": ["http.request"],
  "httpAllowedDomains": []
}
```

**Transport cannot bypass policy.** `capabilities/executor.ts` is the only path from a plan step
to a capability handler, it calls `policy.check()` first, and it contains no transport-conditional
branch. Steps that resolve to `REQUIRE_APPROVAL` are skipped unless the caller passes
`confirmed: true`; steps that resolve to `DENY` never run at all. Every decision is written to
the audit trail with its session id.

## The AI model is replaceable — by design

Nothing outside `src/providers/openai.ts` imports an AI vendor SDK. Everything goes through one
interface:

```ts
// src/providers/ai-provider.ts
interface AIProvider {
  id: string;
  displayName: string;
  generateText(instructions, input, options?): Promise<{ text, model, providerId }>;
  generateJson(instructions, input, jsonShapeDescription, options?): Promise<{ raw, model, providerId }>;
  isConfigured(): boolean;
}
```

The cognitive stages map onto it as **reason → `generateJson`**, **plan → `generateJson`**, and
**evaluate → deterministic code in `core/evaluator.ts`**. Evaluation is deliberately *not* a
provider call, so a model can never grade its own failed execution into a success.

**To add a model/vendor:** write `src/providers/<name>.ts` implementing `AIProvider`, register it
in `providers/index.ts`, set `CHAOS_CORE_PROVIDER=<name>`. The model name itself is configured
once, via `OPENAI_MODEL` — it appears in no other file.

## Capability registry — the extension seam

`Capability` objects are `{ name, description, risk, inputSchema (Zod), annotations, handler }`.
Two ship in V1:

- `cognition.generate_text` — general-purpose text generation via the active provider
- `http.request` — GET-only, gated by `policy.httpAllowedDomains`

To add one — an external API, a database, another MCP server, or one of your own apps: create a
file in `src/capabilities/built-in/` exporting a `Capability`, register it in
`src/capabilities/index.ts`. Nothing in `core/`, `policy/`, `server/`, or `transport/` changes,
and it becomes visible to local and remote clients simultaneously. The AI reasons over the
registry's descriptions to discover what solves a plan step — you never hardcode
`if (task === "email") ...`.

**Future direction:** the registry is the growth path — capability *packs* (registered groups),
per-capability policy keyed on `risk` rather than on names one at a time, an adapter capability
that wraps a remote MCP client so Chaos Core can federate other MCP servers, and durable
procedural memory that learns which capability sequences succeed for recurring objectives.

## Memory

V1 implements the durable **Semantic Memory** layer, behind a `MemoryStore` interface
(`src/memory/types.ts`) with a SQLite implementation (`src/memory/sqlite.ts`) chosen by a factory
(`src/memory/store.ts`). Backed by `node:sqlite` — built into Node 22.5+, zero native deps:
key/value with tags, TTL, substring search, pagination.

Swapping SQLite for Postgres or a vector store means adding one file next to `sqlite.ts` and
changing the factory. The MCP tools, planner, cognitive core, and policy engine don't change,
because none of them reference SQLite.

The same database is used regardless of how a request arrived — a fact written over stdio is
recallable over HTTP, and survives a restart.

**Working Memory** (current session context) is `src/state/session.ts`. **Episodic Memory** (what
happened during past tasks) and **Procedural Memory** (learned successful step sequences) are
named in the architecture but not implemented in V1.

## Setup

```bash
npm install
cp .env.example .env    # then fill in OPENAI_API_KEY
npm run build
```

### Run over stdio (local clients, development)

```bash
npm start
```

`npm run start:stdio` is the explicit equivalent; `npm start` remains stdio so existing local
setups are unaffected.

Under stdio, **stdout belongs to the MCP protocol**. Every diagnostic in the codebase goes
through `observability/logger.ts`, and the stdio transport forces that logger to stderr even if
`CHAOS_CORE_LOG_STREAM=stdout` is set.

### Run over Streamable HTTP (remote clients)

```bash
npm run start:http
```

Listens on `HOST:PORT` (default `127.0.0.1:3000`) and exposes:

| Method | Path | Purpose |
|---|---|---|
| `POST` | `/mcp` | client → server JSON-RPC (initialize, tools/list, tools/call, …) |
| `GET` | `/mcp` | server → client SSE notification stream for an existing session |
| `DELETE` | `/mcp` | explicit session termination |
| `GET` | `/health` | liveness + active session count (not part of MCP) |

Local endpoint: **`http://localhost:3000/mcp`**

The HTTP transport is **stateful**: each `initialize` mints an `Mcp-Session-Id`, and subsequent
requests must carry it. That is what lets `chaoscore_plan` hand a `plan_id` to
`chaoscore_execute` without leaking plans between remote clients. A request with an unknown
session id gets `404`; a non-initialize request with no session id gets `400`.

### Environment variables

| Variable | Default | Purpose |
|---|---|---|
| `OPENAI_API_KEY` | — | Required by the OpenAI provider. Read by the server only; never exposed to MCP clients |
| `OPENAI_MODEL` | `gpt-5.6` | Default model. The single place a model name is configured |
| `OPENAI_REASONING_EFFORT` | `medium` | `none`\|`low`\|`medium`\|`high`\|`xhigh`\|`max` |
| `CHAOS_CORE_PROVIDER` | `openai` | Which registered `AIProvider` answers reason/plan calls |
| `PORT` | `3000` | HTTP transport port |
| `HOST` | `127.0.0.1` | HTTP transport bind address |
| `MCP_HTTP_PATH` | `/mcp` | Path the MCP endpoint is mounted at |
| `MCP_ALLOWED_HOSTS` | — | Comma-separated; setting it enables DNS-rebinding protection |
| `MCP_ALLOWED_ORIGINS` | — | Comma-separated; same |
| `MCP_HTTP_MAX_BODY` | `4mb` | Max JSON body accepted on `/mcp` |
| `CHAOS_CORE_DB_PATH` | `./data/chaos-core.db` | SQLite file for remember/recall |
| `CHAOS_CORE_POLICY_PATH` | `./data/policy.json` | Policy config file |
| `CHAOS_CORE_LOG_STREAM` | `stderr` | `stderr`\|`stdout`; stdio mode always forces `stderr` |
| `CHAOS_CORE_RESPONSE_LIMIT` | `25000` | Character ceiling per tool response |
| `MCP_TRANSPORT` | `stdio` | `stdio`\|`http`, overridden by `--stdio`/`--http` |

A `.env` in the working directory is loaded automatically (Node's built-in loader — no
dependency). `.env.example` contains placeholders only; never commit real credentials.

The pre-0.2 `COGNITION_*` variable names still work as fallbacks.

### Connecting a local MCP client

Claude Desktop / Claude Code / any stdio client:

```json
{
  "mcpServers": {
    "chaos-core": {
      "command": "node",
      "args": ["F:/Chaos-Origins/chaos-core-mcp/dist/index.js", "--stdio"],
      "env": { "OPENAI_API_KEY": "sk-..." }
    }
  }
}
```

Or with the MCP Inspector:

```bash
npm run inspector:stdio
```

### Connecting a remote MCP client

Start the HTTP transport, then point the client at the endpoint URL:

```
http://localhost:3000/mcp
```

For a Claude custom connector, add it as a remote MCP server with that URL (a public deployment
needs a public HTTPS URL — see the security warning below). To poke at it manually:

```bash
npm run inspector:http
```

then choose "Streamable HTTP" and enter the URL.

## ⚠️ Security warning for remote deployment

**V1 ships no authentication.** That is deliberate and is only safe because the HTTP transport
binds to `127.0.0.1` by default. The layer is structured so authentication middleware drops in
cleanly (`AuthMiddleware` in `src/transport/http.ts`, applied to the MCP route before any MCP
handling) — but nothing fake is provided: no stub OAuth, no hard-coded secrets, no bearer token
that only looks like security.

Before exposing this beyond localhost you **must** add:

- **Authentication** on the `/mcp` route (OAuth 2.1 resource server per the MCP auth spec, or a
  gateway that terminates identity)
- **TLS** — the server speaks plain HTTP; terminate TLS at a reverse proxy
- **Rate limiting and request-size limits** — every `reason`/`plan` call spends your OpenAI quota
- **DNS-rebinding protection** — set `MCP_ALLOWED_HOSTS` / `MCP_ALLOWED_ORIGINS`
- **A reviewed `policy.json`** — the default allows every registered capability except those
  requiring confirmation
- **Durable audit storage** — the V1 audit trail is an in-memory ring buffer

If you bind to a non-loopback address without middleware, the server logs a warning at startup
saying exactly this. See `docs/remote-deployment.md` for the full checklist.

The OpenAI API key is read from the server's environment inside `providers/openai.ts` and is
never returned in tool output, inspect payloads, audit entries, or HTTP responses.

## V1 capabilities and boundary

What's in:

- TypeScript/Node, MCP SDK, OpenAI Responses API as the default (swappable) provider
- Dual transport: stdio + Streamable HTTP at `/mcp`, one shared cognitive core
- Six-tool cognitive surface, identical on both transports
- Capability registry + deterministic policy engine + structured audit events
- SQLite Semantic Memory behind a swappable `MemoryStore` interface
- Zod validation on every tool input and every capability input

What's deliberately out:

- No UI
- No agent swarms / multi-agent architecture
- No autonomous background execution — `chaoscore_execute` runs exactly the steps it's given;
  `core/brain.ts`'s full-loop replanning exists but isn't exposed as a tool
- No OAuth implementation, no multi-tenancy, no marketplace
- No MCP-server federation (the registry could host an adapter capability; none ships)

## Build & test

```bash
npm run build
```

```bash
npm test
```

The suite runs against the built output and covers: policy determinism and non-bypassability,
memory persistence across a simulated restart, and a live MCP client connecting over **both**
transports to verify identical tool surfaces, shared memory, and that a denied capability is
blocked on each.

TDQS

A4.7/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct role in the cognitive core loop: reason analyzes before planning, plan produces an executable plan, execute runs it, inspect provides read-only introspection, and remember/recall handle persistent memory. There is no overlap in purpose; even reason and plan, which share similar arguments, are explicitly differentiated by what they produce.

Naming Consistency5/5

All tools follow a consistent naming pattern: the 'chaoscore_' prefix followed by a lowercase verb (reason, plan, execute, inspect, remember, recall). No mixing of camelCase or inconsistent verb styles; the pattern is uniform and predictable.

Tool Count5/5

With 6 tools, the server is well-scoped. Each tool corresponds to a necessary stage of the cognitive core workflow (reason, plan, execute, inspect) plus persistent memory operations (remember/recall). There are no redundant or extraneous tools, and the count is well within the ideal 3-15 range.

Completeness4/5

The tool surface covers the full lifecycle: analyze (reason), plan (plan), execute (execute), observe (inspect), and persist/retrieve knowledge (remember/recall). Minor gaps exist, such as no explicit delete tool for memory (though overwrite covers updates) and no dedicated cancel/abort for plans, but these are not critical to the core loop. Overall, the domain is well-covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues