DungeonQ
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DungeonQinvestigate the after-hours incident and report the route and digest"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DungeonQ — Deception Runtime
Divert a suspicious session into a persistent synthetic world. Observe what it does. Adapt within an approved policy.
DungeonQ is open-source infrastructure for security teams and developers handling human- or AI-driven suspicious sessions. At a designated entry point, its trusted gateway routes an admitted session into a separate synthetic world. The participant can read records, save changes and use a useful Wrong Ticket there. The operator retains the activity history and checks the protected origin through independent evidence.
The current self-hosted reference makes that sequence inspectable with real HTTP, MCP, SSH, a bounded PostgreSQL protocol and a private Unix workload broker. Trusted setup explicitly provisions which contexts are diverted; the protected origin and all records are artificial.
Follow the recorded diversion · Operate your own reference · Acceptance evidence
What happens after DIVERT
Step | Participant experience | Operator evidence and control |
Enter | A real client request reaches its designated synthetic world. | Authenticated routing is correlated with canonical and collector records. |
Work | Read and write persistent records; use a Wrong Ticket for a bounded world-only read. | A ticket never supplies origin authority; state and ticket consumption survive restart. |
Continue | Revisit the same world across supported adapters. | Observe stored requests, operations and world revisions across the session. |
Adapt | Approved follow-up records can appear after a first-time ticket use. | A finite operator grant controls the template, context, expiry and mutation budget. |
Verify | Synthetic success remains within that world. | Check the separate artificial-origin witness and route census; fence the context when required. |
A Wrong Ticket is deliberately useful, rather than an inert fake credential. Its signed scope permits work inside its issuing world. Approved adaptation extends that world with bounded records; participant instructions cannot authorize it. Transport failure never causes a diverted request to fall back to origin.
Related MCP server: Trust Gate MCP
Run the reference
Use Node.js 24.15.0+, npm and the prerequisites in the operation guide:
npm ci --ignore-scripts
npm run runtime -- --data-dir ../dungeonq-runtime-labOpen the printed Control room URL. Keep the private credential file outside Git and AI context. Give a participant only its actor token; keep operator authority separate. Follow the evaluator route to send a real request, consume a ticket, authorize one finite policy and inspect retained state after restart. Reuse the same private data directory to resume.
The client guide covers all five adapters, the CLI and Node/Python SDKs. The container profile adds measured network and file separation. The recorded website and older browser rehearsal are presentations; running the source provides the actual server roles.
Evidence you can inspect
The September 18 v0.11.0 Amazon distribution passed 460/460 tests. Its public CI run passed Ubuntu, macOS and source-bound runtime acceptance: 11/11 required rows and 16/16 container checks. The earlier reference also recorded restart continuity and an inconclusive result when infrastructure was stopped. These are dated observations; the validation record identifies each tested source.
Six recorded checkpoints · Evidence summary · Versioned runtime contract. The checkpoints group one recorded reference acceptance run by capability, with source pointers and timing; public CI is separate corroboration.
For fresh acceptance, follow the container guide and run the documented runtime:proof and runtime:gate commands against the exact clean candidate. Missing or stale evidence, unknown outcomes and failed checks prevent full admission. Public CI does not by itself establish protected-branch enforcement.
Scope and open questions
This is an owned reference using artificial resources, with explicitly provisioned contexts. It does not yet classify arbitrary attacks, transparently intercept host activity or certify production protection. SSH and PostgreSQL accept finite operations; the Unix broker mediates an explicitly connected workload. Local processes share an OS user, while the measured container profile still trusts Docker administration, the shared kernel and the gateway.
“Origin untouched” means the named artificial data/configuration and admission interval were checked; audit records change to record that measurement. Real deployment requires a named authorized connector, detection/identity integration, legitimate-traffic continuity, recovery and environment-specific acceptance.
Engineering success also does not establish deception efficacy. The complete research record retains 0/2 wrong-high-confidence study outcomes, 0/4 unsupported completion claims and the earlier 0/3 defense pilot. Later workspace participants accepted decoy values while explicitly qualifying their synthetic, common-source evidence. Study · Topology · Defense pilot · Workspace.
One core, reusable capabilities
The Proof Kernel and separate approval boundary support the runtime's decisions and evidence. The governed-assistant/Astra profiles exercise bounded clients; Defense and Orders Workspace retain incident response, notifications and account workflows; World, Study and Topology retain research instruments and their original results. These are supporting profiles of DungeonQ, not separate product definitions. Legacy account/email features do not automatically apply to the runtime's reference owner token.
Amazon contribution
The retained Alexa+ simulation lets an assistant investigate through real local MCP while a separate reviewer controls bounded effects. DungeonQ's defensive deception runtime is the product; the assistant workflow demonstrates how a client can participate without acquiring operator authority.
This entry retains Alexa+ primary + Open Source Mini. Contribution and reused foundations · Current submission copy · External MCP client. The assistant is a deterministic Alexa-style simulation, with no Alexa service, Echo integration or AWS deployment claim. The original v0.2 film demonstrates that earlier workflow, not Runtime v1.
Verify and contribute
npm run check
npm run verify:sourceOpen-source review route · Testing · Contributing · Security · Release history. Apache-2.0: preserve LICENSE, NOTICE and third-party notices. Source manifests check integrity, not independent certification. The earlier WebMCP repository, site, submission and evidence remain frozen.
The following dated sections preserve prior workflows, counts and launch statements. They are historical context, not fresh release or remote-submission confirmations.
v0.10.0: the alert belongs to the administrator
The current synthetic reference release adds verified administrator email binding, email-alias sign-in and a durable recipient-bound alert queue. Reauthenticate, verify your address, then inspect the message generated by a new honey contact. Actor and MCP tools cannot choose the recipient or approve a response.
Optional Google / GitHub sign-in adapters bind the verified provider email to an existing Owner, never create one from email alone. They require your registered application and credentials; the default demo has none. Apple sign-in is visibly unavailable in this local profile. Authentication and email delivery are separate services.
Run npm run defense:workspace for the full self-hosted lab. Its default simulated mailbox needs no account, cloud service or paid call. An optional explicitly configured, TLS-verified SMTP transport supports real mail; a simulated verification never becomes a real-mail verification. SMTP acceptance is not an inbox-delivery claim.
Email setup and judge route · Recorded synthetic proof. Reproduce with npm run email:proof -- --out ../dungeonq-email-proof. Prior world, governance, original video and all positive/negative model results below are retained. No enterprise integration or general deception efficacy is claimed.
v0.9.0: an inspectable decoy workspace
New Orders Workspace adds a coherent English snapshot → index → reconciliation trail, stable saved records and four real HTTP MCP tools. Evaluator-only labels are withheld from the participant, while the Owner retains origin truth, local alerts and separately approved rotation. Start a new installation with npm run defense:workspace; no cloud account, API key or paid call is required.
Two fresh Codex sessions used the workspace quantities 19/21 (fixed origin A=7) and recorded completion at 95% confidence. Both still qualified their claims as synthetic-interface, common-source evidence. This demonstrates decoy-data acceptance, not sustained origin misbelief or general attacker deception. The overly broad original metric and prior 0/3 negative pilot are disclosed, with complete raw records and projection/replay verification.
Installation, judge route and results · Recorded A · Recorded B. Run npm run workspace:verify -- evidence/workspace-pilot-v1. The retained video records v0.2, not this new workspace. This release remains synthetic-only and does not claim Alexa service certification.
Retained v0.8.0: one incident, a complete defensive response
Start with npm run defense:proof for the new same-incident reference profile: a configured honey contact opens a bounded synthetic world, creates a durable local alert and proposes a rotation for protected resource A. Separate Owner reauthentication approves the exact manifest. A real loopback TLS issuer rotates once, then checks the old key, new business request, old consumer and world-issued decoy token independently. The world remains usable afterward.
The recorded engineering proof passes 13 named checks, including HTTP/MCP parity, altered-evidence rejection and a complete restart. It uses SCRIPTED_FIXTURE and an automated Owner credential: not a human-presence proof, a live LLM evaluation or production defense. The actor receives no approval or origin-control tool. No API key or paid model call is needed.
Start the English Actor and Owner desks · Inspect the recorded proof · Three-minute OSS review.
npm ci --ignore-scripts
npm run defense:proof -- --out ../dungeonq-defense-proof
# First installation: a new, empty, private directory
mkdir -m 700 ../dungeonq-defense-lab
npm run defense -- --data-dir ../dungeonq-defense-lab --seed 42 --depth 4Use the URLs printed by the launcher. Keep the Owner password outside the Actor/MCP client. This profile supports one fixed synthetic A and one rotation per installation, a 30-day local worker credential and a 30-day self-signed certificate. Restart with the same directory, seed and depth. An expired unclaimed proposal can be refreshed; an uncertain execution uses readback reconciliation. Recovery and exact limits.
MCP is an interface, not transparent interception. A honey contact proves neither AI identity nor compromise of A. LOCAL_SINK_ONLY means local receipt, not email delivery; same-host services do not establish production isolation. An actor already holding a valid A credential needs origin enforcement and a credential/session fence in a real deployment. The reference does not claim to divert arbitrary traffic.
Retained v0.7.0: follow the packet, verify the destination
Run npm run topology for an English, persistent publishing workflow: saved note → edition → preview → queue → fixed filing worker. The actual visitor catalogue has a separate deposit/index path. HTTP and MCP share the same runtime; a separate Observer retains causal evidence. No key or paid model call is needed.
Two-minute workflow guide · All four model pilot records and limits.
Both procedural-memo participants followed the complete local branch (2/2), versus 0/2 early-explanation controls. False completion was not observed (0/4); all four ultimately verified the real goal. This is observed route following, not proof of a false causal belief or general deception efficacy. The older study below remains intact.
Retained v0.6.0: governed actions and a world you can question
The original governed MCP workflow stays intact. New npm run world and npm run study experiences add persistent abstract rooms and a finite two-feature causal experiment: predict → act → reflect, with consent, withdrawal, debrief and a separate-process Observer. No API key is needed.
Two-minute study route · World guide · Research results and raw evidence · Three-minute OSS reviewer guide.
The evidence includes 48 reference-learner conditions, not 48 subjects, and an N=2 Codex pilot with 0/2 wrong-high-confidence induction. Exact pilot model identity was not independently attested. No general human/LLM efficacy is claimed. This is finite feature learning, not exploitable vulnerabilities or attack chains.
The study UI remains Traditional Chinese; its English guide includes important labels. Overview and learner trace are actual recorded-evidence explorer captures, not a live hosted study. The independent Observer requires self-hosting. The original video below records the 0.2.0 governed workflow, not these new features.
Let it investigate. Decide before it acts.
DungeonQ is an Apache-2.0, locally runnable reference lab for testing the boundary between an assistant's request and a human-authorized effect. An assistant uses a real MCP server to investigate a synthetic incident and request containment. A separately authenticated reviewer approves an exact manifest. Only then can the assistant change one synthetic session and obtain a verifiable receipt.
For MCP developers, security engineers and reviewers who need more than an approval label: challenge the boundary, inspect what changed, tamper with the evidence, and rerun the checks.
Synthetic only. The transport, authentication, SQLite writes and signatures are real; the identities, incident and effects are artificial. The governed assistant is deterministic, not an LLM, the Alexa service, an Echo integration or production security software. No cloud account, API key, credit card, live attack or real enterprise connection is needed.
Watch the original 2:35 demo · Source releases · Review the evidence map · Public acceptance runs
Start in five minutes
Reference runtime: Node.js 24.15.0+, npm and OpenSSL with req -addext. Public CI targets Ubuntu 24.04 and macOS 14; the validation record identifies completed runs. Native Windows is not release-accepted. See installation and troubleshooting.
npm ci --ignore-scripts
npm run doctorRun these commands in the unpacked source directory containing package.json. The sections below retain the earlier governed-containment workflow; the v0.8.0 defense entry above is separate.
A. Reproduce the proof without a browser
npm run demo:proof -- --out ../dungeonq-proof-040Expect seven PASS checks, ending with Result: PASS. The command starts a fresh lab on random loopback ports, crosses actual HTTPS/CSRF and Streamable HTTP, verifies scope/signature/tamper/replay, stops the entire stack, and verifies persistence after restart. It exits nonzero on failure, refuses an existing output directory and removes only its own disposable private lab.
The output includes report.json, original and tampered evidence, and a public verification key. No passwords, worker tokens or private keys are exported. The automated driver controls both fixture roles; it does not prove human presence or give the runtime assistant an approval tool.
Independently check the signed receipt:
npm run verify:assistant-evidence -- ../dungeonq-proof-040/evidence.json ../dungeonq-proof-040/trusted-public-key.pem amazon-local-lab
npm run verify:assistant-evidence -- ../dungeonq-proof-040/tampered-evidence.json ../dungeonq-proof-040/trusted-public-key.pem amazon-local-labThe first returns receiptValid: true / exit 0. The second must return false / exit 1: that is the expected rejection, not a failed installation. The harness pins the key before exporting evidence. For artifacts supplied by someone else, obtain the trusted key independently; accepting an accompanying key proves no trusted origin.
B. Experience the human approval boundary
npm run amazonOpen https://127.0.0.1:4186/assistant. Sign in as owner-lab with the new disposable password printed in your terminal. Inspect and handle the local self-signed certificate warning yourself; do not install system trust or disable TLS validation.
Investigate this incident → see route
DENYand a decision digest; no asset changes.Request containment, then ask the assistant to Apply →
HUMAN_APPROVAL_REQUIRED; both assets stay active.In The approval boundary, review the exact session, digest, five-minute expiry and one-effect limit. Reauthenticate and approve.
Ask the assistant to Apply → target becomes
CONTAINED, version 1; the other session staysACTIVE, version 0.Verify receipt, Test tampering, Replay apply, Export evidence → valid original, rejected altered copy, same receipt and no second mutation.
Reviewer walkthrough covers expected outcomes, expiry and restart. The public video records the preceding 0.2.0 workflow; 0.4.0 adds scenario-aware proof, an independent MCP client and public release tooling without changing that UI or approval contract.
What you can actually verify
Question | Inspectable result |
Can the assistant approve its own request? | Six MCP tools, none for approval; an approval command is rejected. Human HTTPS reauthentication is separate. |
Does approval bind the effect's scope? | Exact digest, worker, observed asset/version, expiry and one-effect limit; unrelated asset remains unchanged. |
Can a retry execute twice? | Atomic claim and CAS; replay returns the original receipt without incrementing the asset again. |
Can an edited receipt pass? | Ed25519 verification against a previously pinned key rejects the modified copy. |
Does restart forget the decision? | Real full-stack restart preserves the account, request, effect and receipt in SQLite. |
Do combined failures grant more authority? | All 127 nonempty subsets of seven modeled failure flags remove capabilities; not 127 infrastructure attacks. |
Governance and security value maps claims to code, checks, limits and reuse. Architecture distinguishes modeled decisions, durable local effects and deferred production guarantees.
Bring your own synthetic scenario
Copy assistant/scenarios/after-hours.json and change its synthetic identifiers, seed, policy or signals. The Scenario Pack contract describes accepted fields and limits.
Run your file through the complete proof, not just an analysis preview:
npm run demo:proof -- --scenario ./my-synthetic-scenario.json --out ../my-scenario-proofEach run reports the scenario ID, input/decision/proposal digests, platform and named outcome. Equal admitted input produces equal decision digests; fresh keys, timestamps and receipts are intentionally not byte-identical. Invalid inputs or incorrect expected assertions exit nonzero. The proof cannot attach to or overwrite an existing lab.
Shipped proof case | Expected result | What it actually proves |
|
| Pre-approval denial, separate fixture approval, exact effect, signature, tamper, replay, restart |
|
| Modeled budget exhaustion prevents observation/request creation; assets unchanged |
|
| Rotation is analyzed but this adapter refuses execution; no receipt invented |
Use each path with --scenario and a new --out directory. The two rejection proofs export rejection.json, not a signed effect receipt. A PASS means the named expectation held; it does not mean every case executed an effect. A pack whose cost exceeds its declared budget is rejected even earlier at admission.
Upload in the UI to analyze through MCP. Uploading does not replace the installed lab or grant execution rights.
Install in a fresh lab to exercise the supported containment mapping:
npm run amazon -- --scenario ./my-synthetic-scenario.jsonExecution supports ISOLATE_SESSION, scope 1, AVAILABLE → containment only. Other modeled effects can be analyzed but cannot be executed by this adapter. URLs, real-looking credentials, extra fields, executable content and excess budget are rejected. Never use real incident data.
The reviewer guide names three fixed-seed cases and expected outcomes. MCP documentation includes the six tool authorities and an external local client.
For a separately running client and a Codex connection example, see External Agent walkthrough. Only the worker transport token belongs in that process; human approval stays in the HTTPS UI. CI tests the SDK client, not a live LLM or Alexa account.
Verify, contribute, release
npm run check
npm run verify:sourcecheck runs source tests, three fixed-seed goldens, release-pattern audit, type checking and build. It is not independent security certification. verify:source checks the distribution's file inventory and SHA-256 manifest; a hash manifest is not a signature or trusted timestamp.
Versioned validation record: tests, three goldens, full and rejection proofs; limits included.
Amazon-specific changes and provenance · Developer-tool feedback
Prepared for Build, Ship, Shape: Amazon Developer Hackathon — Alexa+ + Open Source Mini Challenge. The older WebMCP entry is separate and unchanged. Public availability and passing local tests do not imply contest acceptance, OpenAI endorsement or production readiness.
Apache-2.0: LICENSE, NOTICE, third-party notices, SBOM. The package is intentionally private to prevent accidental npm publishing; the GitHub source is open source.
This server cannot be deployed
Maintenance
Related MCP Connectors
Independent effect verification and signed receipts for consequential AI agent actions.
Deterministic authorization for one proposed AI agent action, returned with a signed receipt.
AI agent infrastructure for discovery, authorization, execution, identity, and signed receipts.
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
Related MCP Servers
- FlicenseCqualityBmaintenanceEnables AI assistants to autonomously perform site reliability engineering including monitoring, root-cause analysis, impact assessment, and remediation with cryptographic zero-trust enforcement.20-
- AlicenseAqualityBmaintenancePost-quantum, tamper-evident receipts for consequential agent actions. Provides tools for auditing, gating decisions, and egress classification with quantum-hardened security.7Apache 2.0
- FlicenseNot gradedqualityCmaintenanceEnables approval-gated incident response workflows that gather evidence through read-only MCP tools, perform idempotent writes, and preserve a durable audit trail.-
- AlicenseBqualityBmaintenanceEnables AI agents to execute actions under an accountability layer with identity passports, mandates, a permission gate, and a tamper-evident journal, while requiring human confirmation for irreversible operations.737 PyPI4AGPL 3.0