Enzo
Enzo is an MCP server that turns big questions into small, falsifiable claims, tracks evidence and investigation state, and optionally uses Jev semantic sensing under explicit approval.
Atomize claims: validate one proposed claim and return ATOMIC, DECOMPOSE, or NEEDS_REFINEMENT, creating child atoms when needed.
Observe atoms: submit typed evidence and get honest outcomes: VERIFIED, CONTRADICTED, UNKNOWN, or INSUFFICIENT_EVIDENCE.
Use deterministic evidence first: tests, schema validation, AST inspection, static analysis, runtime measurements, and similar results remain authoritative.
Invoke Jev semantic sensing only when needed and only via two-phase dispatch: preview a manifest, approve its SHA-256 digest, then allow the external call.
Track investigation history: view atoms, derived statuses, current frontier, contradictions, unresolved atoms, branches, revisions, dependencies, and parent impacts.
Manage evidence requirements and provenance: enforce required evidence kinds, deterministic requirements, assumptions, and source provenance.
Cache approved Jev responses locally in a bounded, namespace-scoped SQLite ledger without storing request state; replay still requires fresh approval.
Keep privacy boundaries: external sends are opt-in, allowlisted, digest-bound, and limited in size; no request state is persisted in the response memory.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@EnzoDecompose the question: is our authentication flow secure?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Enzo is not another autonomous-agent framework. The LLM keeps responsibility for reasoning, strategy, and deciding what to investigate. Enzo makes each next question precise enough to test.
Why Enzo?
The name is inspired by the Japanese Zen ensō (円相), the hand-drawn circle. Enzo uses that image as a reasoning metaphor: begin with the large circle of a problem, then reduce it into smaller circles until each contains exactly one independently falsifiable claim.
large question → smaller question → atomic claim → evidence → semantic resultMost reasoning systems are comfortable producing an answer. Enzo is designed to apply pressure before that answer exists:
Common failure | Enzo's response |
One question hides several claims | Decompose it into independently testable atoms |
Missing information becomes vague confidence | Return |
Semantic judgment overrides a test | Deterministic evidence remains authoritative |
Conclusions lose their history | Preserve evidence, provenance, revisions, and dependencies |
A tool quietly sends context outside the process | Preview an allowlisted payload and bind the send to its SHA-256 digest |
Related MCP server: genpark-multihop-research-query-decomposer-skill
How it works
flowchart LR
A[Large question] --> B{One falsifiable claim?}
B -- No --> C[Independent child claims]
C --> B
B -- Yes --> D[Collect typed evidence]
D --> E{Deterministic evidence resolves it?}
E -- Yes --> G[Constrained result]
E -- No --> F[Jev semantic sensor]
F --> G
G --> H[LLM chooses what to ask next]The responsibilities stay deliberately separate:
Component | Responsibility |
LLM | Intelligence, strategy, interpretation, and choosing the next question |
Enzo | Atomicity, evidence contracts, state transitions, and decomposition pressure |
Jev | Typed semantic sensing against supplied context and evidence |
Jev response memory | Bounded reuse of exact approved responses without storing request state |
Deterministic tools | Tests, schemas, AST inspection, type checking, and runtime measurements |
Quick start
Requirements: Python 3.11+ and uv.
git clone https://github.com/mahawi1992/enzo-mcp.git
cd enzo-mcp
uv sync --group dev
uv run enzo-mcpJev is optional for deterministic workflows. To enable it, create a local .env
file containing:
TYPESAFE_API_KEY=your-keyThe file is ignored by Git. Configuring a key alone never authorizes a send.
Every external observation requires an explicit field selection, a previewed
manifest, allow_external_jev=true, and the matching manifest SHA-256.
Run with the local key loaded when semantic sensing is needed:
uv run --env-file .env enzo-mcpExact approved Jev responses can optionally be remembered across Enzo processes.
Add a deployment-owned namespace and explicit model epoch to .env:
ENZO_JEV_CACHE_DIR=.enzo-cache
ENZO_JEV_CACHE_NAMESPACE=my-project-dev
ENZO_JEV_CACHE_MODEL_EPOCH=jev-1.13.0
ENZO_JEV_CACHE_TTL_SECONDS=86400
ENZO_JEV_CACHE_MAX_ENTRIES=10000The directory is ignored by Git. Use a new model epoch whenever the effective Jev
model changes; do not use the moving jev-latest alias as an epoch.
Add Enzo to Codex
Add this to ~/.codex/config.toml, replacing the path with your checkout:
[mcp_servers.enzo]
command = "uv"
args = ["run", "--env-file", ".env", "enzo-mcp"]
cwd = "/absolute/path/to/enzo-mcp"Restart Codex. Enzo will expose exactly three tools.
The three tools
Tool | Purpose |
| Admit one atomic claim, safely decompose it, or request refinement |
| Evaluate typed evidence and optionally invoke Jev with explicit consent |
| Return canonical investigation history, derived status, and the current frontier |
Atomicity has three outcomes:
ATOMIC— one operationalized predicate can be evaluated independently.DECOMPOSE— multiple safe, explicit child claims can vary independently.NEEDS_REFINEMENT— the claim appears composite or vague, but a mechanical split could change its meaning.
Observation has four honest outcomes:
VERIFIEDCONTRADICTEDUNKNOWNINSUFFICIENT_EVIDENCE
UNKNOWN is a useful result: it tells the LLM what must be learned next.
Example
Ask Enzo to atomize a compound security question:
{
"request": {
"root_goal": "Determine whether the production session cookie is hardened",
"question": "Does the cookie set Secure and HttpOnly?",
"subject": "the production session cookie",
"predicate": "sets Secure=true; sets HttpOnly=true",
"scope": "production session configuration",
"expected_value": true,
"evidence_requirements": [
{
"description": "Inspect the production cookie configuration",
"accepted_kinds": ["SCHEMA_VALIDATION"],
"deterministic_required": true
}
],
"verification_method": "SCHEMA_VALIDATION"
}
}Enzo returns DECOMPOSE and creates two independently falsifiable children:
Does the production session cookie set Secure=true?
Does the production session cookie set HttpOnly=true?Each child can now receive its own evidence, result, provenance, and parent impact.
Jev integration
PydanticJevSensor uses Pydantic AI's TypeSafe provider and maps Enzo's answer
contracts to Jev primitives:
Enzo answer type | Jev primitive |
|
|
|
|
|
|
Jev's native probability or confidence is preserved in sensor evidence. Enzo does not manufacture an aggregate confidence score. A deterministic instrument is used first whenever it can resolve the atom more reliably.
Jev receives only a caller-selected projection. The logical request contains the model, a minimal typed question contract, selected context keys, and selected evidence records. Expected answers, investigation IDs, evidence requirement IDs, unselected values, and prior sensor output remain local. Evidence payloads, assumptions, and provenance are excluded unless each category is explicitly enabled in the selection.
External dispatch is deliberately two-phase:
Call
enzo_observewithdispatch_selection. Enzo returns a canonical manifest and makes zero provider calls.Review that exact logical request in the host.
Repeat the observation with the same selection,
allow_external_jev=true, andapproved_dispatch_sha256set to the manifest digest.
Enzo rebuilds the logical provider inputs {model, state, questions} and sends only
when the digest still matches. Changing any selected content invalidates the prior
approval. A matching digest records what the caller approved; it is not proof that
a human saw or understood the request.
Selections are capped at eight context entries and eight evidence entries, with a 4 KiB limit per selected value and a 16 KiB canonical logical-request limit. Limit failures and hash mismatches make zero Jev calls.
Prior sensor output is never sent back into a later Jev request, preventing semantic
feedback loops. Exact observation retries are idempotent, while dependency changes
correctly invalidate replayed results. A manifest-bound provider failure is also
cached to prevent an ambiguous transport retry from making a duplicate external
call; an intentional retry must use a new approval_reference.
Jev response memory
Enzo includes an optional local response ledger inspired by JevCache. It follows the same useful core idea—reuse an answer for a stable request fingerprint—but keeps Enzo's stricter approval and privacy boundaries:
A cache lookup happens only after the normal two-phase manifest approval succeeds.
Identity binds the exact canonical request digest to a deployment namespace, an explicit model epoch, and Enzo's cache schema version.
The SQLite ledger stores only Jev's typed response and integrity metadata. It never stores selected context, evidence, atom IDs, investigation IDs, or approval references.
Entries expire, have a configurable upper bound, and are evicted oldest-accessed first.
Corrupt, expired, or contract-invalid entries are ignored. Cache read/write failure never changes the underlying provider result.
Existing cache directories, databases, and SQLite sidecars are hardened to private owner-only permissions; symbolic-link cache files are rejected.
An exact cache hit can be replayed without a provider key, but it still requires a fresh valid approval for that manifest.
Public sharing is deliberately not automatic. JevCache v0.1 currently publishes a prebuilt sidecar and global fingerprint index, but its public repository does not yet provide enough source and lifecycle detail for Enzo to independently audit deletion, expiry, tenant isolation, or redaction-equivalence behavior. The local ledger is the safe memory layer today; a remote JevCache adapter can be added behind the same narrow interface once that contract is reviewable. See the trust-boundary notes.
AG-UI boundary
AG-UI is a good future host adapter for displaying a manifest, pausing on a typed approval interrupt, and resuming with the reviewed digest. It is not a durable database or a crash-replay guarantee. Enzo therefore keeps the approval invariant in its core and leaves restart persistence to the integrating runtime. See the integration boundary.
Design principles
One atom tests exactly one independently falsifiable semantic claim.
Atomicity is semantic, not a measure of sentence length.
Deterministic evidence outranks semantic judgment.
Assumptions remain visibly distinct from facts.
Parent conclusions are derived from their dependency graph.
Contradictions remain explicit.
Unknowns expose gaps instead of becoming invented certainty.
The MCP surface stays small enough to understand.
Development
uv sync --group dev
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytest
uv buildThe current suite contains 85 tests covering contracts, atomicity, dependency derivation, revision history, manifest-bound dispatch, payload canaries and limits, generated replay and dependency invariants, offline evaluation fixtures, Jev answer validation, response-memory privacy and lifecycle controls, and the stdio MCP surface.
Privacy
Semantic observations are local-only by default. allow_external_jev=true alone
does nothing. An external call also needs an explicit dispatch_selection and a
matching approved_dispatch_sha256. Preview the manifest, select only necessary
fields, and keep payload, assumption, and provenance flags off unless Jev needs
them. Redact credentials and personal data before any external observation.
The optional response ledger does not retain request state, but Jev's typed answers can still be sensitive. Keep its directory private, use distinct namespaces per trust boundary, choose short retention where appropriate, and never commit the database.
Project status
Enzo is a focused v0.3 alpha implementation. Its three-tool surface is intentional; the contracts may evolve as real investigations expose better invariants.
Focused issues and pull requests are welcome. If the idea of turning large circles into testable small ones is useful to you, consider starring the repository.
License
MIT © 2026 Martin Harold Williams
Available Tools
3 toolsenzo_atomizeC
Validate one proposed claim and admit, decompose, or request refinement.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| atom | Yes | |
| reasons | Yes | |
| children | No | |
| decision | Yes | |
| investigation_id | Yes | |
| next_atom_needed | Yes | |
| validation_basis | No | |
| refinement_questions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It names the three possible actions but does not explain what 'admit', 'decompose', or 'request refinement' do in practice, whether the tool mutates state, or what prerequisites apply. This is insufficient for a validation/decomposition decision tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the verb and immediately names the three possible decisions. It contains no filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite an output schema being present, the tool has a highly complex input schema with nested types and enums, no annotations, and a one-line description. Critical context is missing: what constitutes a valid proposed claim, what triggers each outcome, how it relates to the sibling tools, and what state changes occur. The description is not self-sufficient for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, with one complex 'request' parameter buried in a large $defs structure. The description adds only the generic phrase 'one proposed claim' and does not explain how to fill the required AtomizeRequest fields or how they relate to the validation outcomes, so it fails to compensate for the absence of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('validate one proposed claim') and lists the three expected outcomes ('admit, decompose, or request refinement'), so an agent has a clear idea of what the tool accomplishes. It does not explicitly distinguish itself from sibling tools enzo_observe or enzo_state, so it loses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to choose this tool over enzo_observe or enzo_state, nor are the alternatives mentioned. The one sentence implies use when a proposed claim has been formed, but there are no exclusions or condition-based routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enzo_observeB
Evaluate typed evidence for one admitted atomic claim.
Deterministic evidence is authoritative. With TYPESAFE_API_KEY configured, semantic-only questions are evaluated by Pydantic AI's TypeSafe/Jev provider.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| source | Yes | |
| status | Yes | |
| atom_id | Yes | |
| evidence | Yes | |
| provenance | Yes | |
| parent_impact | Yes | |
| missing_information | No | |
| verification_method | Yes | |
| violated_constraints | No | |
| satisfied_constraints | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the behavioral disclosure burden. It adds meaningful context: deterministic evidence is authoritative, and semantic-only questions are routed to Pydantic AI's TypeSafe/Jev provider when the API key is present. However, it does not disclose whether the tool has side effects, requires specific authorization, or what happens when the API key is missing, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two compact sentences with no filler. Front-loaded with the core purpose, the second sentence adds a useful behavioral condition. It loses one point only because the jargon-laden first sentence is not expanded for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the ObserveRequest schema and the absence of annotations, this description is too thin. It does not explain key concepts like 'admitted atomic claim', the evidence kinds, the semantics of constraints fields, or whether the tool returns a judgment or mutates state. While an output schema exists, proper use still requires more contextual guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does very little. It hints at the 'deterministic' flag and maybe 'allow_external_jev', but does not explain how to construct the ObserveRequest, what 'evidence' versus 'missing_information' means, or how constraints are used. The rich schema remains largely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Evaluate') and resource ('typed evidence for one admitted atomic claim'), which clearly distinguishes it from the sibling tools enzo_atomize and enzo_state. However, 'admitted atomic claim' is domain jargon that may be unclear without additional context, so it misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when there is typed evidence for an atomic claim, and it notes a conditional (TYPESAFE_API_KEY configured) for semantic-only questions. It does not explicitly state when to use this tool versus enzo_atomize or enzo_state, nor does it mention exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enzo_stateB
Return canonical history and computed investigation status.
| Name | Required | Description | Default |
|---|---|---|---|
| investigation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| atoms | Yes | |
| root_goal | Yes | |
| branch_ids | Yes | |
| unknown_atoms | Yes | |
| contradictions | Yes | |
| verified_atoms | Yes | |
| current_frontier | Yes | |
| investigation_id | Yes | |
| unresolved_atoms | Yes | |
| contradicted_atoms | Yes | |
| insufficient_evidence_atoms | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must carry the behavioral burden. 'Return' and 'computed' imply a read-only snapshot that is derived rather than stored, but the description does not explicitly state whether it has side effects, authorization requirements, rate limits, or freshness guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with no filler; the main action and result type are front-loaded. Nothing wastes an agent's attention.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Low parameter complexity and an existing output schema reduce the need to spell out return values. Still, the missing usage guidance and lack of explicit read-only/behavioral confirmation leave the definition merely adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description partially compensates by tying the result to 'investigation status,' which makes investigation_id's role inferable. However, it adds no format, source, or validity details for the ID beyond what the property name already implies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: returning 'canonical history and computed investigation status.' It is clearly a state-retrieval tool, though it does not explicitly contrast itself with the siblings enzo_atomize and enzo_observe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus enzo_atomize or enzo_observe, no exclusions, and no alternative routing. The only usage signal is the tool name and the terse description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
enzo_atomize - First observed
enzo_observe - First observed
enzo_state
TDQS
Scored across 3 tools
Each tool maps to a distinct stage in the investigation workflow: atomize handles claim validation/admission, observe evaluates evidence against an admitted claim, and state returns the accumulated history and status. There is no meaningful overlap between tool purposes.
All tools share the consistent enzo_ prefix followed by a single lowercase word, creating a predictable enzo_<action> pattern. Even though 'state' is noun-like, the naming convention is uniform across the set.
Three tools is well-scoped for this narrow investigation workflow, and each tool is load-bearing: one for claims, one for evidence evaluation, and one for state. No tool feels redundant or missing for the stated purpose.
The server covers the full claim-investigation loop: claim validation/decomposition, evidence evaluation, and canonical history/status retrieval. An agent can iterate between atomize and observe and consult state at any point without hitting a dead end.
Related MCP Connectors
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
MCP-native web evidence and claim verification: cited, source-grounded evidence for AI agents.
Goal and task planning MCP for Codex and AI agents, with evidence-backed completion.
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceVerified-memory engine that decomposes AI agent memories into atomic claims with executable falsifiers and continuously re-verifies them against reality, returning facts with freshness verdicts via 30 MCP tools.1Apache 2.0
- FlicenseNot gradedqualityBmaintenanceEnables decomposition of complex research queries into multi-hop sub-queries and synthesis DAGs, with scientific consensus analysis, citation credibility verification, and deterministic JSON outputs for MCP-compliant clients.8-
- AlicenseNot gradedqualityCmaintenanceEnables structured reasoning with TypeSafe's Jev model through a single evaluate tool that returns answers, probabilities, and confidence for evidence-based questions.145 npm1MIT
- FlicenseNot gradedqualityBmaintenanceEnables autonomous agents and MCP clients to decompose text into atomic claims and verify them against reference knowledge, detecting ungrounded claims, numerical contradictions, and hallucinated entities.7-