Skip to main content
Glama

Auto Interpretability Lab

Auto Interpretability Lab is an experimental toolkit for giving a research agent controlled access to a language model's internals. It combines model inspection and activation interventions with a persistent experiment ledger, bounded execution, and a separate human approval path for compute.

The intended research subjects are Qwen3-1.7B-Base and Qwen3-1.7B. The implementation on main has been tested locally with a small, randomly initialized Qwen3 model on CPU. Its cloud provider is a simulator. Live RunPod deployment is being developed separately in draft PR #1 on feat/live-deployment; that work is not included in main and is not a claim of completed GPU acceptance. Hidden evaluation and the independent scientific workflow also remain unfinished. No scientific discovery is claimed.

To try the local implementation, start with the development tests. Read the verification report for what has actually been tested and the service guides before configuring a host.

What main includes

  • Persistent experiments: a local SQLite ledger with job states, leases, cancellation, recovery, immutable accepted manifests, and a hash-chained audit log.

  • Model operations: activation capture, patching, ablation, steering, probe fitting, bounded generation, weight statistics, tensor slices, and module inspection.

  • Two numerical backends: Hugging Face/PyTorch as a reference and NNsight for interventions, with local parity tests.

  • Trusted execution: a separate worker process, authenticated loopback HTTP, a controller-side dispatcher, and hash-checked artifact transfer and retention.

  • Research tools over MCP: stdio access to jobs, history, hypotheses, artifacts, CPU experiments, and compute requests. Submitting a job does not approve a GPU start.

  • CPU Python sandbox: rootless Podman with restricted mounts, no network or GPU, and enforced time, memory, process, CPU, and output limits.

  • Compute-control foundation: one-time human approvals, price/runtime checks, and an independent watchdog, currently connected to a persistent provider simulator.

The repository is auto-interpretability-lab. The existing MCP command remains probe-mcp; the Python distribution and import package remain probe-core and probe_core. Service names, configuration paths and environment variables keep their existing probe identifiers so the documented commands continue to work.

Related MCP server: agent-orchestrator

Current limits

Area

Status

Cloud provider

Simulator on main; live RunPod work is separate in draft PR #1

GPU execution

CUDA paths exist; real GPU limits, SSH deployment, and canonical checkpoint parity still require validation

Flexible experiments

Arbitrary CPU Python plus fixed worker operations; no arbitrary agent-written GPU Python

Interventions

Apply to the prompt's prefill pass; generation is a separate, unmodified operation

Scientific evaluation

Confirmation/replication execution is refused until a private evaluator is implemented

Research workflow

Hypothesis storage and freezing exist; Explorer/Skeptic/Replicator orchestration and blind calibration remain pending

Advanced methods

SAE/Qwen-Scope, circuit tracing, and automated novelty adjudication are not implemented

Model weights, credentials, research datasets, and built container images are not included. Deployment files on main are examples and do not install or start services.

Architecture

Research agent
    │ stdio MCP
    ▼
Research facade ──────────► local ledger and retained artifacts
    │ compute request
    ▼
Human approval ───────────► trusted controller ──► provider simulator
                                  │                   ▲
                         approved job batch      watchdog stop
                                  ▼
                              dispatcher
                                  │ authenticated loopback HTTP
                                  │ (SSH tunnel for a remote deployment)
                                  ▼
                              model worker

Research facade ──────────► rootless CPU Python sandbox

The research interface cannot consume approvals, accept worker results, certify shutdown, or administer hidden evaluation. Production separation depends on distinct OS identities and correctly installed permissions; running every role under one account is a development setup. The live ledger belongs on local controller storage, with external backups, rather than a network filesystem.

Install and run the development tests

Requirements: Linux, Git, uv, and Python 3.13. The worker dependencies include PyTorch and require several gigabytes of disk. Use a maintained SQLite build; see the core guide.

git clone https://github.com/jaykobdetar/auto-interpretability-lab.git
cd auto-interpretability-lab
uv python install 3.13
uv sync --locked --all-extras
uv run --locked python -m pytest -q

These tests create their small model locally and need no cloud credentials or downloaded model checkpoint. Tests involving services bind local sockets. The nine real sandbox integration tests skip unless a suitable Podman environment and image are configured; a default test run does not validate containment.

The recorded local acceptance run on September 19, 2026 passed 429 tests with no skips, including the real sandbox gate. That is historical evidence for the source commit and environment recorded in verification details, not a new test result for every subsequent commit or a production-readiness claim. To reproduce a run, retain the checked-out commit, locked dependencies and any sandbox image/policy pins alongside its results.

To run the mandatory sandbox gate after preparing its pinned image:

PROBE_SANDBOX_REQUIRED=1 \
PROBE_SANDBOX_IMAGE="sha256:REPLACE_WITH_64_HEX_IMAGE_ID" \
uv run --locked python -m pytest tests/test_sandbox.py tests/test_sandbox_service.py -q

Replace the image placeholder before running. The sandbox guide covers the image build, subordinate user mappings, delegated cgroups, and service runtime requirements. With PROBE_SANDBOX_REQUIRED=1, missing capabilities fail the tests instead of skipping them.

Configure the services

Installing the package does not create a running lab. These guides describe the local implementation on main; follow them to prepare private configuration, service accounts, sockets, worker assets, and directories:

For the separate RunPod deployment effort, follow draft PR #1. Its installation procedures and acceptance records belong to that branch; do not infer their availability from these local examples.

After the research service is configured, an MCP client can launch probe-mcp with --socket pointing to its research socket and --service-uid set to the trusted service's actual UID. The controller approval interface belongs to the human account, not the research agent's MCP configuration.

Next milestones

  1. Complete and review the separate RunPod deployment work: trusted services, pinned model assets, storage, backups, and verified shutdown.

  2. Validate canonical BF16/CUDA execution, resource limits, and remote recovery.

  3. Implement scientific metrics, matched controls, private held-out evaluation, and trusted promotion rules.

  4. Add fresh Explorer, Skeptic, and Replicator sessions; pass blind calibration before beginning open-ended discovery.

  5. Run one narrow behavioral/mechanistic pilot, then expand into paired model and thinking-mode studies, advanced methods, and scientific-efficiency reporting.

Contributions should keep new claims tied to evidence, preserve the separation between research and compute authority, and include relevant regression tests. Larger changes to the execution model or scientific protocol benefit from an issue describing the proposed behavior and acceptance criteria first.

License

MIT. Model weights and third-party dependencies retain their own licenses.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server for mechanistic interpretability research, enabling agents to drive probe-causality and SAE-feature experiments via 8 typed tools on user's own compute (Colab).
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    MCP server for analyzing transformer language model internals, exposing tools to trace attention heads and neurons responsible for predictions, get activation statistics, ablate components, and sketch circuits.
    4
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to govern and access memory engines over MCP, with scoped read/write permissions, journaling, secret scrubbing, and bounded context delivery at session start.
    MIT