Skip to main content
Glama

Crucible

Turn a real model failure into a deterministic, publishable benchmark task.

Live app Engine Agent tools License MIT Next.js 16 TypeScript strict Live feeds

Live app · GitHub · Write-up · API · Agent · Issues


A benchmark that cannot be re-run is a rumour. Most evals are graded either by a language model asking whether an answer "looks right", or by a substring match so loose that a model can pass without doing the work.

Crucible grades a task with exact assertions — regex, JSON paths, numeric ranges — and then scores the task itself: is it deterministic, can it actually tell two models apart, are its inputs pinned, can a stranger reproduce it? Six weighted factors, every one produced from a measured quantity, each carrying the sentence that produced it.

Real output: a task forged through the UI, graded by crucible-grade-v1.0.0. Captured from the deployed app by npm run browser.


✨ Features

  • Forge a task from a failure you actually saw. Name the failure, write the prompt that provokes it, write assertions that decide the answer. It is written to a real database immediately.

  • Grade with exact assertions only. Five deterministic kinds (regex, contains, not_contains, json_path_equals, number_between) and one explicit judge kind. Every outcome shows the exact span or value that decided it.

  • A grade for the task, not the model. Six factors, weights published in every API response, every export and the UI. A high score means the task is worth publishing.

  • The heat dial. Set how much of the grade a language-model judge decides. Watch the determinism factor, the score, the model ranking and the colour of the page respond — and the change is persisted and sealed.

  • Reproducibility, checked against reality. Live Hugging Face Hub facts tell you whether a model's weights are gated and whether a pinnable revision exists at all. Closed models publish no repository, so nobody outside the provider can reproduce a run against them — and the app says so.

  • A tamper-evident audit trail. Every create, update, grade, decision and delete appends to a per-task SHA-384 chain over canonical JSON. A replay endpoint recomputes it and names the first broken link.

  • A real agent interface. Eleven typed tools over JSON-RPC 2.0. Mutating tools call the same service layer the buttons call, and are idempotent on a key.

  • A Kaggle bundle that grades identically. Export a runnable kaggle_benchmarks task, a self-contained Python grader and the CLI commands. The test suite proves the Python port scores the same as the TypeScript engine.

  • Dossiers you can keep. Markdown, JSON or CSV with the factor table, every transcript grade, provenance and the chain head.


Related MCP server: agentseed

🚀 Quickstart

Requires Node 22+. No API keys. No database to install.

git clone https://github.com/aniruddhaadak80/crucible.git
cd crucible
npm ci
npm run dev          # http://localhost:3000

Local development uses an embedded Postgres (PGlite — real Postgres compiled to WebAssembly) stored under ./.crucible. It is created on first run and needs no configuration.

Quality commands

Command

What it does

npm run typecheck

tsc --noEmit, TypeScript strict

npm run lint

ESLint, no suppressions

npm run test

62 unit tests via node --test

npm run build

Production build

npm run verify:db

Exercises the real store: schema, seed, CRUD, audit, replay, tamper detection

npm run verify:bundle

Generates a Kaggle bundle and proves the Python grader matches the engine

npm run audit:secrets

Scans tracked files and the built client bundle for credentials; fails on a real one

npm run journey

Boots the app and walks the whole HTTP journey

npm run browser

Real Chromium pass on desktop and mobile

npm run verify:live

Full proof against a deployed URL

npm run gate

Everything above that does not need a browser

Production environment variables

See .env.example. Only one is required:

DATABASE_URL=postgresql://user:password@host/db?sslmode=require
NEXT_PUBLIC_SITE_URL=https://your-alias.vercel.app

Crucible refuses to start a production-mode process without a hosted DATABASE_URL. It will not fall back to the embedded adapter, because a serverless filesystem is not a durable store.


📐 Architecture

graph TB
  subgraph client["Browser"]
    UI["Pour column UI"]
    DIAL["Heat dial"]
  end
  subgraph server["Next.js 16 · App Router"]
    PROXY["proxy.ts · scope cookie"]
    PAGES["Server components"]
    API["REST route handlers"]
    MCP["JSON-RPC 2.0 endpoint"]
    SVC["Service layer"]
  end
  subgraph core["Deterministic core"]
    GRADER["crucible-grader"]
    ENGINE["crucible-grade"]
    CHAIN["SHA-384 chain"]
  end
  UI --> PAGES
  DIAL --> API
  PAGES --> SVC
  API --> SVC
  MCP --> SVC
  SVC --> GRADER
  GRADER --> ENGINE
  SVC --> CHAIN
  SVC --> DB[("Postgres")]
  SVC --> LIVE["Live feeds"]

  classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
  classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
  classDef agent fill:#34d399,color:#08080a,stroke:#34d399
  classDef infra fill:#94a3b8,color:#08080a,stroke:#94a3b8
  classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
  class UI,DIAL,PAGES live
  class GRADER,ENGINE,CHAIN engine
  class MCP agent
  class PROXY,API,SVC,DB infra
  class LIVE external

Data pipeline and honest fallback

graph LR
  REQ["Route handler"] --> TIMEOUT["AbortSignal timeout"]
  TIMEOUT --> RETRY["2 bounded retries"]
  RETRY --> HUB["Hugging Face Hub API"]
  RETRY --> AX["arXiv API"]
  HUB --> NORM["Normalise to src/lib/types.ts"]
  AX --> NORM
  NORM --> ENV["FeedEnvelope · live | fallback"]
  NORM -.-> SEALED["Sealed dated snapshot"]
  ENV --> UI["Rendered with status tag"]

  classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
  classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
  classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
  class REQ,NORM,UI live
  class HUB,AX external
  class SEALED risk

Every feed returns status: "live" | "fallback", the time it was produced, the upstream identity and — when degraded — the reason. A sealed snapshot is always rendered with its capture date. It is never presented as current.

The deterministic engine

graph TB
  TASK["Benchmark task"] --> GRADE["gradeTask"]
  TX["Recorded transcripts"] --> GRADE
  GRADE --> DET["determinism · 0.26"]
  GRADE --> DISC["discrimination · 0.22"]
  GRADE --> SEALF["fixture seal · 0.18"]
  GRADE --> SPEC["assertion specificity · 0.16"]
  GRADE --> REPRO["reproduction · 0.10"]
  GRADE --> COST["cost fit · 0.08"]
  DET --> SUM["Weighted sum · 0-100"]
  DISC --> SUM
  SEALF --> SUM
  SPEC --> SUM
  REPRO --> SUM
  COST --> SUM
  SUM --> VERDICT["Verdict + band + weakest factor"]

  classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
  classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
  classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
  class GRADE,DET,DISC,SEALF,SPEC,REPRO,COST,SUM engine
  class TASK,TX live
  class VERDICT risk

gradeTask is the only scoring function. The task page, the REST endpoint, the agent tools and the dossier export all call it. Weights sum to exactly 1.00 and a test asserts it.

Discrimination is measured, not asserted. It is the population standard deviation of recorded scores against a target of 0.22. A task every model passes scores 0 on it, and the verdict says so.

Agent sequence

graph LR
  CLIENT["MCP client"] --> INIT["initialize"]
  INIT --> LIST["tools/list"]
  LIST --> CALL["tools/call"]
  CALL --> VALID["Validate arguments"]
  VALID --> IDEM["Check idempotency key"]
  IDEM -->|seen| RETURN["Return stored result"]
  IDEM -->|new| SVC["Service layer"]
  SVC --> MUTATE["Create / update / decide / retire"]
  MUTATE --> SEAL["Append audit event"]
  SEAL --> RESULT["Result + seal"]

  classDef agent fill:#34d399,color:#08080a,stroke:#34d399
  classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
  classDef infra fill:#94a3b8,color:#08080a,stroke:#94a3b8
  classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
  class CLIENT,INIT,LIST,CALL,RESULT agent
  class VALID,IDEM,SVC,MUTATE engine
  class SEAL infra
  class RETURN risk

Integrity and seal replay

graph TB
  GEN["Genesis constant"] --> E1["seal 1 = SHA-384(prev || canonical(event))"]
  E1 --> E2["seal 2"]
  E2 --> E3["seal n"]
  CANON["Canonical JSON<br/>sorted keys · stable arrays<br/>no silent NaN"] --> E1
  TMB["Tombstone retained<br/>on delete"] --> E3
  E3 --> REPLAY["Replay recomputes every seal"]
  REPLAY --> OK["Intact"]
  REPLAY --> BROKEN["First broken sequence named"]

  classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
  classDef agent fill:#34d399,color:#08080a,stroke:#34d399
  classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
  class GEN,CANON,E1,E2,E3 engine
  class OK,TMB agent
  class BROKEN risk

Two known digest vectors are pinned in the test suite, so a change to canonical form cannot pass silently. The smoke test also edits a stored event directly in the database and asserts that replay reports the correct sequence number.

Deployment

graph LR
  PUSH["git push main"] --> CI["GitHub Actions<br/>typecheck · lint · test · build"]
  CI --> VERCEL["Vercel production build"]
  ENV["DATABASE_URL · SITE_URL"] --> VERCEL
  VERCEL --> ALIAS["Production alias"]
  ALIAS --> HEALTH["/api/health round trip"]
  ALIAS --> LIVE["verify:live 12-point proof"]
  LIVE --> HISTORY["Build history recorded"]

  classDef infra fill:#94a3b8,color:#08080a,stroke:#94a3b8
  classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
  classDef agent fill:#34d399,color:#08080a,stroke:#34d399
  classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
  class PUSH,CI,VERCEL,ALIAS infra
  class HEALTH,LIVE live
  class HISTORY agent
  class ENV external

🔌 API

All endpoints are scoped to an anonymous HTTP-only session cookie. Cross-scope reads return 404, indistinguishable from not-found.

Health

curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/health | jq
{
  "status": "ok",
  "store": { "kind": "neon-postgres", "durable": true, "roundTrip": "SELECT 1 returned 1", "ok": true },
  "seed": { "applied": true, "rows": 3, "error": null },
  "engine": { "version": "crucible-grade-v1.0.0", "grader": "crucible-grader-v1.0.0" }
}

This is a real round trip, not a static object. It returns 503 when the store is unhealthy.

Create, read back, update, delete

# Create
curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks \
  -H 'content-type: application/json' \
  -d '{
    "name": "Unit-of-measure drift",
    "failureMode": "The model returns micrograms where the schema demands milligrams.",
    "prompt": "Return ONLY JSON of the form {\"meta\":{\"units\":\"mg\"}}.",
    "sealedFixtures": ["export.csv=sha256:6f1a2c9d4e8b73a5"],
    "assertions": [
      {"id":"u","kind":"json_path_equals","label":"units are mg","weight":3,"path":"meta.units","jsonExpected":"mg"},
      {"id":"n","kind":"not_contains","label":"no micrograms","weight":2,"needle":"µg"}
    ],
    "transcripts": [
      {"modelId":"Qwen/Qwen2.5-72B-Instruct","completion":"{\"meta\":{\"units\":\"mg\"}}","latencyMs":1840,"tokensOut":96},
      {"modelId":"google/gemini-2.5-flash","completion":"{\"meta\":{\"units\":\"µg\"}}","latencyMs":2100,"tokensOut":101}
    ],
    "seed": 20260923, "targetTemp": 0, "tokenBudget": 400
  }' | jq '.verdict.score, .verdict.factors[0]'

# Read back
curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id> | jq '.task.name, .grade.grades[0].outcomes'

# Update
curl -sX PATCH https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id> \
  -H 'content-type: application/json' -d '{"name":"Unit-of-measure drift (v2)"}' | jq '.task.name'

# Delete — soft, and the tombstone keeps the chain replayable
curl -sX DELETE https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id> | jq '.retired, .integrity.ok'

Run the engine

curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/grade | jq '{score:.verdict.score, band:.verdict.band.id, seal:.seal}'

Verify integrity

curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/integrity | jq '.integrity'

Export

curl -s "https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/dossier?format=markdown" -o crucible-task.md
curl -s https://crucible-aniruddha-adaks-projects.vercel.app/api/tasks/<id>/bundle | jq '.files[].path'

Errors

Every failure returns the same envelope, with no stack traces or environment values:

{ "error": { "code": "validation_error", "message": "The task failed validation.",
  "details": [{ "path": "assertions[0].pattern", "message": "pattern is not a valid regular expression" }] } }

400 bad request · 404 not found · 409 conflict or broken chain · 422 validation · 429 rate limited · 503 store unavailable


🔌 Agent interface

JSON-RPC 2.0 over HTTP POST at https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp. Manifest: /mcp.json.

curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp \
  -H 'content-type: application/json' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | jq '.result.tools[].name'

Tool

Kind

Purpose

list_tasks

read

Every task in the session, with verdicts

get_task

read

Full record, per-assertion evidence, chain head

rank_lineup

analysis

Re-grade and rank recorded models, with separation

grade_task

analysis

Run the engine with live Hub facts, seal the verdict

verify_integrity

read

Replay the chain, name the first broken link

live_signals

read

Hub facts per model + newest arXiv papers

export_bundle

read

Generate the Kaggle Benchmarks bundle

forge_task

write

Create a task (idempotent)

revise_task

write

Patch a task (idempotent)

record_decision

write

File adopt / iterate / discard (idempotent)

retire_task

write

Soft delete, tombstone retained (idempotent)

Retrying a mutating call with the same idempotencyKey returns the original result and performs no second mutation:

curl -sX POST https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp -H 'content-type: application/json' -d '{
  "jsonrpc":"2.0","id":2,"method":"tools/call","params":{
    "name":"forge_task",
    "arguments":{
      "name":"nested envelope collapse",
      "failureMode":"The model wraps a tool payload in a second envelope and the caller parses data: null.",
      "prompt":"Reply with ONLY the JSON the caller expects.",
      "idempotencyKey":"readme-demo"
    }}}'

Point any MCP client at it:

{
  "mcpServers": {
    "crucible": { "type": "http", "url": "https://crucible-aniruddha-adaks-projects.vercel.app/api/mcp" }
  }
}

📁 Project map

Route

User goal

Methods

Notes

/

Understand the product, see real grades

GET

Live verdicts from the reference tasks

/suite

Browse everything you own

GET

Filter state lives in the URL

/forge

Create a task

POST

Dynamic assertion builder, live determinism

/task/[id]

Inspect, decide, export, retire

GET PATCH DELETE

Heat dial, assertion tape, chain replay

/lineup

Rank recorded models

GET

Next to live Hub reproducibility facts

/agent

Drive the agent interface

POST

One-click calls, visible request/response

/dossier

Export the suite

GET

Markdown / JSON / CSV

/chain

Replay every chain

GET

Reports the first broken link

/settings

Store health, weights, feeds

GET PUT

Includes a real store round trip

API route

Purpose

GET /api/health

Real store round trip, engine versions

GET/POST /api/tasks

List (paginated) and create

GET/PATCH/DELETE /api/tasks/[id]

Read, revise, retire

POST /api/tasks/[id]/decision

File a decision

POST /api/tasks/[id]/grade

Run the engine, append a grade event

GET /api/tasks/[id]/integrity

Replay the chain

GET /api/tasks/[id]/audit

Append-only history

GET /api/tasks/[id]/dossier

Download Markdown / JSON / CSV

GET /api/tasks/[id]/bundle

Kaggle bundle, or one file with ?file=

GET /api/feed

Normalised live data envelopes

GET/PUT /api/settings

Per-scope preferences

POST /api/mcp

JSON-RPC 2.0 agent endpoint

GET /mcp.json

Agent manifest

Module map

Path

Responsibility

src/lib/grade.ts

The grader and the six-factor engine. No dependencies.

src/lib/canonical.ts

Canonical JSON, SHA-384 sealing, chain verification

src/lib/types.ts

Normalised domain types, including every external source

src/lib/db/

Schema, adapter selection, typed repository, audit chain

src/lib/live/

Bounded fetching, Hub and arXiv normalisers, sealed fallbacks

src/lib/service.ts

The one layer REST, MCP and the UI all call

src/lib/mcp.ts

Tool catalogue and JSON-RPC dispatch

src/lib/kaggle.ts

Kaggle bundle generator

src/lib/dossier.ts

Markdown / JSON / CSV export


🎛 Every control, and what it actually does

The skill this project was built against forbids dead controls. Every one:

Control

Route

Request

Effect

Forge a task

/forge

POST /api/tasks

Row written, audit event appended

Load an example

/forge

—

Fills the form from a worked example

Add / remove assertion

/forge

—

Recomputes determinism live

Heat dial

/task/[id]

PATCH /api/tasks/[id]

Re-grades, persists, seals

Re-grade and seal

/task/[id]

POST …/grade

Appends a grade event, returns a new seal

Transcript selector

/task/[id]

—

Switches the assertion tape

Record decision

/task/[id]

POST …/decision

Decision stored with the score at that moment

Retire as tombstone

/task/[id]

DELETE /api/tasks/[id]

Soft delete, replay returned as proof

Dossier downloads

/task/[id], /dossier

GET …/dossier

Downloads a real file

Kaggle bundle

/task/[id]

GET …/bundle

Generates runnable Python

Suite filters

/suite

—

Filter state in the URL, shareable

Lineup task picker

/lineup

—

Selection in the URL

Tool buttons

/agent

POST /api/mcp

Real JSON-RPC call, response shown

initialize

/agent

POST /api/mcp

Real handshake

Save settings

/settings

PUT /api/settings

Persisted per scope


🔐 Security model

  • No accounts. A visitor's work belongs to an unguessable 128-bit scope id in an HTTP-only cookie, set in proxy.ts before the first render. Client script cannot read it.

  • Cross-scope reads are 404, not 403, so the API cannot be used to discover that a task exists.

  • All input is validated before it reaches the database: string and collection limits, enum membership, regex compilation (an invalid pattern is rejected, never executed), JSON-path character classes and 500-character pattern cap.

  • All SQL is parameterised. Identifiers are never interpolated.

  • Deletes are soft. The tombstone is retained so the chain still replays.

  • Upstream hosts are allowlisted (huggingface.co, export.arxiv.org) and model ids are constrained to owner/name, so a stored value cannot be turned into a request to another host.

  • No secrets anywhere. The core product needs no API key. DATABASE_URL is read server-side only and never reaches a client bundle.

  • Rate limiting is best-effort. Anonymous write limits are per-container and reset on a cold start. That is a floor against a runaway client, not a boundary — see SECURITY.md.


📊 Data provenance

Source

Used for

Failure behaviour

Hugging Face Hub API

Gating, licensing, pinnable revisions, download counters

Sealed snapshot dated 2026-10-03

arXiv API

Newest cs.CL capability-evaluation papers

Sealed snapshot dated 2026-10-03

Model metadata is the Hugging Face Hub's own record. Download and like counts are Hub counters, not usage telemetry. Preprint metadata is the author's own arXiv submission.

The three bundled reference tasks ship with seeded transcripts. They are labelled as bundled reference material everywhere they appear and are never presented as a live model run. User-created data is never replaced by fallback data.


⚠️ What this is not

Crucible does not score models and does not predict capability. A high grade means the task is worth publishing. It says nothing about how good any model is.

Specifically:

  • Where a grade would need a language-model judge, the weight is reported as undecided and the verdict is marked degraded. It is never counted as a pass.

  • Discrimination is a property of the recorded models. A task with two similar models recorded cannot be shown to separate them.

  • The Hub's absence of a repository is reported as unreproducible, not as a quality judgement about the model.

  • An audit chain detects rewriting history. It cannot stop somebody with write access from recomputing the whole log from genesis; that needs the head published somewhere append-only and independent.


🗺️ Roadmap

Now

Shipped and verified in this repository.

graph TB
  A["Exact-assertion grader"] --> B["Six-factor engine"]
  B --> C["Hash-chained audit"]
  C --> D["MCP agent tools"]
  D --> E["Kaggle bundle export"]
  E --> F["Dossier exports"]

  classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
  classDef agent fill:#34d399,color:#08080a,stroke:#34d399
  classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
  class A,B engine
  class C,D agent
  class E,F live

Next

  • Published chain heads. Anchor each task's chain head to an append-only external log, so a whole-log rewrite becomes detectable rather than merely possible. Outcome: a reader can verify a task has not been rewritten since a date.

  • Suite-level runs. Record several tasks as one suite and report a suite verdict, so a whole benchmark can be graded rather than one task at a time. Outcome: a single export describing a benchmark, not a task.

  • Assertion diffing. When a task is revised, show which assertions changed and re-score every historical transcript against both versions. Outcome: you can see what a revision actually did to the numbers.

  • Partial-credit reporting. Export per-assertion pass rates as a matrix. Outcome: you can see which assertion every model fails, which is usually the interesting one.

graph LR
  HEAD["Published heads"] --> SUITE["Suite runs"]
  SUITE --> DIFF["Assertion diffing"]
  DIFF --> MATRIX["Pass-rate matrix"]

  classDef live fill:#22d3ee,color:#08080a,stroke:#22d3ee
  classDef agent fill:#34d399,color:#08080a,stroke:#34d399
  class HEAD,SUITE live
  class DIFF,MATRIX agent

Later

  • Model-proxy passthrough. Run the exported task against a live proxy from the app and import the results, so the loop does not leave the browser. Outcome: a task can be measured against real models without hand-editing files.

  • Grader plugins. Let a project ship its own assertion kinds behind the same determinism accounting. Outcome: domain-specific assertions stay measurable without losing the guarantee.

  • Signed bundles. Sign a generated bundle so a reader can verify it came from a specific task record. Outcome: a shared task file is verifiably untampered.

graph TB
  PROXY["Model-proxy passthrough"] --> PLUG["Grader plugins"]
  PLUG --> SIGN["Signed bundles"]

  classDef external fill:#fbbf24,color:#08080a,stroke:#fbbf24
  classDef engine fill:#a78bfa,color:#08080a,stroke:#a78bfa
  classDef risk fill:#fb7185,color:#08080a,stroke:#fb7185
  class PROXY external
  class PLUG engine
  class SIGN risk

🤝 Contributing

Issues and pull requests are welcome, especially ones that find a case where a score is unearned.

npm ci
npm run gate     # typecheck, lint, tests, build, store checks, bundle checks

If your change affects grading, add a test in src/lib/__tests__/ — the engine is a pure module with no dependencies, so it is cheap to test. If it affects the audit chain, the canonical form or the exported Python grader, the known-vector and TypeScript-to-Python parity checks are the ones that matter.

See CONTRIBUTING.md.


📄 License

MIT © Aniruddha Adak

Built with Next.js 16, React 19, Tailwind CSS v4, PGlite, the Neon serverless driver, and Framer Motion.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI coding agents to enforce spec-driven development and verify code before it is marked done, using six tools that catch invented APIs, scan for hallucinated content, check plugin conformance, sandbox-run tests, validate schemas, and record audit evidence.
    771 npm
    8
    PolyForm Noncommercial 1.0.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Captures AI agent runs and turns them into tamper-evident execution records showing tool use, timing, failures, recoveries, and human interventions. Records can be inspected, exported, and verified offline.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables governed tool-calling agents with policy decisions, optional human approval, hash-chained audit logging, and deterministic evaluation.
    MIT