Crucible
Fetches live arXiv API data, normalizes it into feed envelopes, and renders it with live/fallback status for research paper feeds.
Provides live Hugging Face Hub facts about models, including whether weights are gated and whether a pinnable revision exists, supporting reproducibility checks for benchmark tasks.
Exports a runnable kaggle_benchmarks task, a self-contained Python grader, and CLI commands so the benchmark can be run and graded identically on Kaggle.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Crucibleforge a benchmark task from this failure: model ignored the regex and returned the wrong JSON path"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Crucible
Turn a real model failure into a deterministic, publishable benchmark task.
Live app · GitHub · Write-up · API · Agent · Issues
A benchmark that cannot be re-run is a rumour. Most evals are graded either by a language model asking whether an answer "looks right", or by a substring match so loose that a model can pass without doing the work.
Crucible grades a task with exact assertions — regex, JSON paths, numeric ranges — and then scores the task itself: is it deterministic, can it actually tell two models apart, are its inputs pinned, can a stranger reproduce it? Six weighted factors, every one produced from a measured quantity, each carrying the sentence that produced it.
Real output: a task forged through the UI, graded by crucible-grade-v1.0.0.
Captured from the deployed app by npm run browser.
✨ Features
Forge a task from a failure you actually saw. Name the failure, write the prompt that provokes it, write assertions that decide the answer. It is written to a real database immediately.
Grade with exact assertions only. Five deterministic kinds (regex, contains, not_contains, json_path_equals, number_between) and one explicit judge kind. Every outcome shows the exact span or value that decided it.
A grade for the task, not the model. Six factors, weights published in every API response, every export and the UI. A high score means the task is worth publishing.
The heat dial. Set how much of the grade a language-model judge decides. Watch the determinism factor, the score, the model ranking and the colour of the page respond — and the change is persisted and sealed.
Reproducibility, checked against reality. Live Hugging Face Hub facts tell you whether a model's weights are gated and whether a pinnable revision exists at all. Closed models publish no repository, so nobody outside the provider can reproduce a run against them — and the app says so.
A tamper-evident audit trail. Every create, update, grade, decision and delete appends to a per-task SHA-384 chain over canonical JSON. A replay endpoint recomputes it and names the first broken link.
A real agent interface. Eleven typed tools over JSON-RPC 2.0. Mutating tools call the same service layer the buttons call, and are idempotent on a key.
A Kaggle bundle that grades identically. Export a runnable
kaggle_benchmarkstask, a self-contained Python grader and the CLI commands. The test suite proves the Python port scores the same as the TypeScript engine.Dossiers you can keep. Markdown, JSON or CSV with the factor table, every transcript grade, provenance and the chain head.
Related MCP server: agentseed
🚀 Quickstart
Requires Node 22+. No API keys. No database to install.
git clone https://github.com/aniruddhaadak80/crucible.git
cd crucible
npm ci
npm run dev # http://localhost:3000Local development uses an embedded Postgres (PGlite — real Postgres compiled
to WebAssembly) stored under ./.crucible. It is created on first run and needs
no configuration.
Quality commands
Command | What it does |
|
|
| ESLint, no suppressions |
| 62 unit tests via |
| Production build |
| Exercises the real store: schema, seed, CRUD, audit, replay, tamper detection |
| Generates a Kaggle bundle and proves the Python grader matches the engine |
| Scans tracked files and the built client bundle for credentials; fails on a real one |
| Boots the app and walks the whole HTTP journey |
| Real Chromium pass on desktop and mobile |
| Full proof against a deployed URL |
| Everything above that does not need a browser |
Production environment variables
See .env.example. Only one is required:
DATABASE_URL=postgresql://user:password@host/db?sslmode=require
NEXT_PUBLIC_SITE_URL=https://your-alias.vercel.appCrucible refuses to start a production-mode process without a hosted
DATABASE_URL. It will not fall back to the embedded adapter, because a
serverless filesystem is not a durable store.
📐 Architecture
graph TB
subgraph client["Browser"]
UI["Pour column UI"]
DIAL["Heat dial"]
end
subgraph server["Next.js 16 · App Router"]
PROXY["proxy.ts · scope cookie"]
PAGES["Server components"]
API["REST route handlers"]
MCP["JSON-RPC 2.0 endpoint"]
SVC["Service layer"]
end
subgraph core["Deterministic core"]
GRADER["crucible-grader"]
ENGINE["crucible-grade"]
CHAIN["SHA-384 chain"]
end
UI --> PAGES
DIAL --> API
PAGES --> SVC
API --> SVC
MCP --> SVC
SVC --> GRADER
GRADER --> ENGINE
SVC --> CHAIN
SVC --> DB[("Postgres")]
SVC --> LIVE["Live feeds"]
classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
classDef agent fill:#34d399,color:#08080a,stroke:#34d399
classDef infra fill:#94a3b8,color:#08080a,stroke:#94a3b8
classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
class UI,DIAL,PAGES live
class GRADER,ENGINE,CHAIN engine
class MCP agent
class PROXY,API,SVC,DB infra
class LIVE externalData pipeline and honest fallback
graph LR
REQ["Route handler"] --> TIMEOUT["AbortSignal timeout"]
TIMEOUT --> RETRY["2 bounded retries"]
RETRY --> HUB["Hugging Face Hub API"]
RETRY --> AX["arXiv API"]
HUB --> NORM["Normalise to src/lib/types.ts"]
AX --> NORM
NORM --> ENV["FeedEnvelope · live | fallback"]
NORM -.-> SEALED["Sealed dated snapshot"]
ENV --> UI["Rendered with status tag"]
classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
class REQ,NORM,UI live
class HUB,AX external
class SEALED riskEvery feed returns status: "live" | "fallback", the time it was produced, the
upstream identity and — when degraded — the reason. A sealed snapshot is always
rendered with its capture date. It is never presented as current.
The deterministic engine
graph TB
TASK["Benchmark task"] --> GRADE["gradeTask"]
TX["Recorded transcripts"] --> GRADE
GRADE --> DET["determinism · 0.26"]
GRADE --> DISC["discrimination · 0.22"]
GRADE --> SEALF["fixture seal · 0.18"]
GRADE --> SPEC["assertion specificity · 0.16"]
GRADE --> REPRO["reproduction · 0.10"]
GRADE --> COST["cost fit · 0.08"]
DET --> SUM["Weighted sum · 0-100"]
DISC --> SUM
SEALF --> SUM
SPEC --> SUM
REPRO --> SUM
COST --> SUM
SUM --> VERDICT["Verdict + band + weakest factor"]
classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
class GRADE,DET,DISC,SEALF,SPEC,REPRO,COST,SUM engine
class TASK,TX live
class VERDICT riskgradeTask is the only scoring function. The task page, the REST endpoint, the
agent tools and the dossier export all call it. Weights sum to exactly 1.00 and
a test asserts it.
Discrimination is measured, not asserted. It is the population standard
deviation of recorded scores against a target of 0.22. A task every model
passes scores 0 on it, and the verdict says so.
Agent sequence
graph LR
CLIENT["MCP client"] --> INIT["initialize"]
INIT --> LIST["tools/list"]
LIST --> CALL["tools/call"]
CALL --> VALID["Validate arguments"]
VALID --> IDEM["Check idempotency key"]
IDEM -->|seen| RETURN["Return stored result"]
IDEM -->|new| SVC["Service layer"]
SVC --> MUTATE["Create / update / decide / retire"]
MUTATE --> SEAL["Append audit event"]
SEAL --> RESULT["Result + seal"]
classDef agent fill:#34d399,color:#08080a,stroke:#34d399
classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
classDef infra fill:#94a3b8,color:#08080a,stroke:#94a3b8
classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
class CLIENT,INIT,LIST,CALL,RESULT agent
class VALID,IDEM,SVC,MUTATE engine
class SEAL infra
class RETURN riskIntegrity and seal replay
graph TB
GEN["Genesis constant"] --> E1["seal 1 = SHA-384(prev || canonical(event))"]
E1 --> E2["seal 2"]
E2 --> E3["seal n"]
CANON["Canonical JSON<br/>sorted keys · stable arrays<br/>no silent NaN"] --> E1
TMB["Tombstone retained<br/>on delete"] --> E3
E3 --> REPLAY["Replay recomputes every seal"]
REPLAY --> OK["Intact"]
REPLAY --> BROKEN["First broken sequence named"]
classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
classDef agent fill:#34d399,color:#08080a,stroke:#34d399
classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
class GEN,CANON,E1,E2,E3 engine
class OK,TMB agent
class BROKEN riskTwo known digest vectors are pinned in the test suite, so a change to canonical form cannot pass silently. The smoke test also edits a stored event directly in the database and asserts that replay reports the correct sequence number.
Deployment
graph LR
PUSH["git push main"] --> CI["GitHub Actions<br/>typecheck · lint · test · build"]
CI --> VERCEL["Vercel production build"]
ENV["DATABASE_URL · SITE_URL"] --> VERCEL
VERCEL --> ALIAS["Production alias"]
ALIAS --> HEALTH["/api/health round trip"]
ALIAS --> LIVE["verify:live 12-point proof"]
LIVE --> HISTORY["Build history recorded"]
classDef infra fill:#94a3b8,color:#08080a,stroke:#94a3b8
classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
classDef agent fill:#34d399,color:#08080a,stroke:#34d399
classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
class PUSH,CI,VERCEL,ALIAS infra
class HEALTH,LIVE live
class HISTORY agent
class ENV external🔌 API
All endpoints are scoped to an anonymous HTTP-only session cookie. Cross-scope
reads return 404, indistinguishable from not-found.
Health
curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/health | jq{
"status": "ok",
"store": { "kind": "neon-postgres", "durable": true, "roundTrip": "SELECT 1 returned 1", "ok": true },
"seed": { "applied": true, "rows": 3, "error": null },
"engine": { "version": "crucible-grade-v1.0.0", "grader": "crucible-grader-v1.0.0" }
}This is a real round trip, not a static object. It returns 503 when the store
is unhealthy.
Create, read back, update, delete
# Create
curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks \
-H 'content-type: application/json' \
-d '{
"name": "Unit-of-measure drift",
"failureMode": "The model returns micrograms where the schema demands milligrams.",
"prompt": "Return ONLY JSON of the form {\"meta\":{\"units\":\"mg\"}}.",
"sealedFixtures": ["export.csv=sha256:6f1a2c9d4e8b73a5"],
"assertions": [
{"id":"u","kind":"json_path_equals","label":"units are mg","weight":3,"path":"meta.units","jsonExpected":"mg"},
{"id":"n","kind":"not_contains","label":"no micrograms","weight":2,"needle":"µg"}
],
"transcripts": [
{"modelId":"Qwen/Qwen2.5-72B-Instruct","completion":"{\"meta\":{\"units\":\"mg\"}}","latencyMs":1840,"tokensOut":96},
{"modelId":"google/gemini-2.5-flash","completion":"{\"meta\":{\"units\":\"µg\"}}","latencyMs":2100,"tokensOut":101}
],
"seed": 20260923, "targetTemp": 0, "tokenBudget": 400
}' | jq '.verdict.score, .verdict.factors[0]'
# Read back
curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id> | jq '.task.name, .grade.grades[0].outcomes'
# Update
curl -sX PATCH https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id> \
-H 'content-type: application/json' -d '{"name":"Unit-of-measure drift (v2)"}' | jq '.task.name'
# Delete — soft, and the tombstone keeps the chain replayable
curl -sX DELETE https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id> | jq '.retired, .integrity.ok'Run the engine
curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/grade | jq '{score:.verdict.score, band:.verdict.band.id, seal:.seal}'Verify integrity
curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/integrity | jq '.integrity'Export
curl -s "https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/dossier?format=markdown" -o crucible-task.md
curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/bundle | jq '.files[].path'Errors
Every failure returns the same envelope, with no stack traces or environment values:
{ "error": { "code": "validation_error", "message": "The task failed validation.",
"details": [{ "path": "assertions[0].pattern", "message": "pattern is not a valid regular expression" }] } }400 bad request · 404 not found · 409 conflict or broken chain · 422
validation · 429 rate limited · 503 store unavailable
🔌 Agent interface
JSON-RPC 2.0 over HTTP POST at https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp.
Manifest: /mcp.json.
curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp \
-H 'content-type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | jq '.result.tools[].name'Tool | Kind | Purpose |
| read | Every task in the session, with verdicts |
| read | Full record, per-assertion evidence, chain head |
| analysis | Re-grade and rank recorded models, with separation |
| analysis | Run the engine with live Hub facts, seal the verdict |
| read | Replay the chain, name the first broken link |
| read | Hub facts per model + newest arXiv papers |
| read | Generate the Kaggle Benchmarks bundle |
| write | Create a task (idempotent) |
| write | Patch a task (idempotent) |
| write | File adopt / iterate / discard (idempotent) |
| write | Soft delete, tombstone retained (idempotent) |
Retrying a mutating call with the same idempotencyKey returns the original
result and performs no second mutation:
curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":2,"method":"tools/call","params":{
"name":"forge_task",
"arguments":{
"name":"nested envelope collapse",
"failureMode":"The model wraps a tool payload in a second envelope and the caller parses data: null.",
"prompt":"Reply with ONLY the JSON the caller expects.",
"idempotencyKey":"readme-demo"
}}}'Point any MCP client at it:
{
"mcpServers": {
"crucible": { "type": "http", "url": "https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp" }
}
}📁 Project map
Route | User goal | Methods | Notes |
| Understand the product, see real grades |
| Live verdicts from the reference tasks |
| Browse everything you own |
| Filter state lives in the URL |
| Create a task |
| Dynamic assertion builder, live determinism |
| Inspect, decide, export, retire |
| Heat dial, assertion tape, chain replay |
| Rank recorded models |
| Next to live Hub reproducibility facts |
| Drive the agent interface |
| One-click calls, visible request/response |
| Export the suite |
| Markdown / JSON / CSV |
| Replay every chain |
| Reports the first broken link |
| Store health, weights, feeds |
| Includes a real store round trip |
API route | Purpose |
| Real store round trip, engine versions |
| List (paginated) and create |
| Read, revise, retire |
| File a decision |
| Run the engine, append a grade event |
| Replay the chain |
| Append-only history |
| Download Markdown / JSON / CSV |
| Kaggle bundle, or one file with |
| Normalised live data envelopes |
| Per-scope preferences |
| JSON-RPC 2.0 agent endpoint |
| Agent manifest |
Module map
Path | Responsibility |
| The grader and the six-factor engine. No dependencies. |
| Canonical JSON, SHA-384 sealing, chain verification |
| Normalised domain types, including every external source |
| Schema, adapter selection, typed repository, audit chain |
| Bounded fetching, Hub and arXiv normalisers, sealed fallbacks |
| The one layer REST, MCP and the UI all call |
| Tool catalogue and JSON-RPC dispatch |
| Kaggle bundle generator |
| Markdown / JSON / CSV export |
🎛 Every control, and what it actually does
The skill this project was built against forbids dead controls. Every one:
Control | Route | Request | Effect |
Forge a task |
|
| Row written, audit event appended |
Load an example |
| — | Fills the form from a worked example |
Add / remove assertion |
| — | Recomputes determinism live |
Heat dial |
|
| Re-grades, persists, seals |
Re-grade and seal |
|
| Appends a grade event, returns a new seal |
Transcript selector |
| — | Switches the assertion tape |
Record decision |
|
| Decision stored with the score at that moment |
Retire as tombstone |
|
| Soft delete, replay returned as proof |
Dossier downloads |
|
| Downloads a real file |
Kaggle bundle |
|
| Generates runnable Python |
Suite filters |
| — | Filter state in the URL, shareable |
Lineup task picker |
| — | Selection in the URL |
Tool buttons |
|
| Real JSON-RPC call, response shown |
|
|
| Real handshake |
Save settings |
|
| Persisted per scope |
🔐 Security model
No accounts. A visitor's work belongs to an unguessable 128-bit scope id in an HTTP-only cookie, set in
proxy.tsbefore the first render. Client script cannot read it.Cross-scope reads are
404, not403, so the API cannot be used to discover that a task exists.All input is validated before it reaches the database: string and collection limits, enum membership, regex compilation (an invalid pattern is rejected, never executed), JSON-path character classes and 500-character pattern cap.
All SQL is parameterised. Identifiers are never interpolated.
Deletes are soft. The tombstone is retained so the chain still replays.
Upstream hosts are allowlisted (
huggingface.co,export.arxiv.org) and model ids are constrained toowner/name, so a stored value cannot be turned into a request to another host.No secrets anywhere. The core product needs no API key.
DATABASE_URLis read server-side only and never reaches a client bundle.Rate limiting is best-effort. Anonymous write limits are per-container and reset on a cold start. That is a floor against a runaway client, not a boundary — see SECURITY.md.
📊 Data provenance
Source | Used for | Failure behaviour |
Gating, licensing, pinnable revisions, download counters | Sealed snapshot dated 2026-10-03 | |
Newest cs.CL capability-evaluation papers | Sealed snapshot dated 2026-10-03 |
Model metadata is the Hugging Face Hub's own record. Download and like counts are Hub counters, not usage telemetry. Preprint metadata is the author's own arXiv submission.
The three bundled reference tasks ship with seeded transcripts. They are labelled as bundled reference material everywhere they appear and are never presented as a live model run. User-created data is never replaced by fallback data.
⚠️ What this is not
Crucible does not score models and does not predict capability. A high grade means the task is worth publishing. It says nothing about how good any model is.
Specifically:
Where a grade would need a language-model judge, the weight is reported as undecided and the verdict is marked
degraded. It is never counted as a pass.Discrimination is a property of the recorded models. A task with two similar models recorded cannot be shown to separate them.
The Hub's absence of a repository is reported as unreproducible, not as a quality judgement about the model.
An audit chain detects rewriting history. It cannot stop somebody with write access from recomputing the whole log from genesis; that needs the head published somewhere append-only and independent.
🗺️ Roadmap
Now
Shipped and verified in this repository.
graph TB
A["Exact-assertion grader"] --> B["Six-factor engine"]
B --> C["Hash-chained audit"]
C --> D["MCP agent tools"]
D --> E["Kaggle bundle export"]
E --> F["Dossier exports"]
classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
classDef agent fill:#34d399,color:#08080a,stroke:#34d399
classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
class A,B engine
class C,D agent
class E,F liveNext
Published chain heads. Anchor each task's chain head to an append-only external log, so a whole-log rewrite becomes detectable rather than merely possible. Outcome: a reader can verify a task has not been rewritten since a date.
Suite-level runs. Record several tasks as one suite and report a suite verdict, so a whole benchmark can be graded rather than one task at a time. Outcome: a single export describing a benchmark, not a task.
Assertion diffing. When a task is revised, show which assertions changed and re-score every historical transcript against both versions. Outcome: you can see what a revision actually did to the numbers.
Partial-credit reporting. Export per-assertion pass rates as a matrix. Outcome: you can see which assertion every model fails, which is usually the interesting one.
graph LR
HEAD["Published heads"] --> SUITE["Suite runs"]
SUITE --> DIFF["Assertion diffing"]
DIFF --> MATRIX["Pass-rate matrix"]
classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
classDef agent fill:#34d399,color:#08080a,stroke:#34d399
class HEAD,SUITE live
class DIFF,MATRIX agentLater
Model-proxy passthrough. Run the exported task against a live proxy from the app and import the results, so the loop does not leave the browser. Outcome: a task can be measured against real models without hand-editing files.
Grader plugins. Let a project ship its own assertion kinds behind the same determinism accounting. Outcome: domain-specific assertions stay measurable without losing the guarantee.
Signed bundles. Sign a generated bundle so a reader can verify it came from a specific task record. Outcome: a shared task file is verifiably untampered.
graph TB
PROXY["Model-proxy passthrough"] --> PLUG["Grader plugins"]
PLUG --> SIGN["Signed bundles"]
classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
class PROXY external
class PLUG engine
class SIGN risk🤝 Contributing
Issues and pull requests are welcome, especially ones that find a case where a score is unearned.
npm ci
npm run gate # typecheck, lint, tests, build, store checks, bundle checksIf your change affects grading, add a test in src/lib/__tests__/ — the engine
is a pure module with no dependencies, so it is cheap to test. If it affects the
audit chain, the canonical form or the exported Python grader, the known-vector
and TypeScript-to-Python parity checks are the ones that matter.
See CONTRIBUTING.md.
📄 License
MIT © Aniruddha Adak
Built with Next.js 16, React 19, Tailwind CSS v4, PGlite, the Neon serverless driver, and Framer Motion.
This server cannot be deployed
Maintenance
Related MCP Connectors
Reproducible benchmarks and reliability evidence for agent tools.
Paid deterministic data-quality and execution-verification tools for AI agents.
26 pay-per-call AI agent tools for marketing, sales, content, research, verification, and integration testing. Agents pay per call via x402 micropayments (USDC on Base), raw JSON in and out. Includes delivery-acceptance receipts (ACCEPT/REJECT/REVIEW), a webhook capture/replay/conformance lab, and n8n workflow validation with bounded JSON-Patch repair.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides AI agents with honest benchmark rankings (Agentic Memory Index and Agentic Search Index) for AI tools, plus graded checks and telemetry for x402 endpoints.6 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables AI coding agents to enforce spec-driven development and verify code before it is marked done, using six tools that catch invented APIs, scan for hallucinated content, check plugin conformance, sandbox-run tests, validate schemas, and record audit evidence.771 npm8PolyForm Noncommercial 1.0.0
- AlicenseNot gradedqualityCmaintenanceCaptures AI agent runs and turns them into tamper-evident execution records showing tool use, timing, failures, recoveries, and human interventions. Records can be inspected, exported, and verified offline.MIT
- AlicenseNot gradedqualityBmaintenanceEnables governed tool-calling agents with policy decisions, optional human approval, hash-chained audit logging, and deterministic evaluation.MIT