codecalc
# codecalc — universal code & logic calculator for AI models
<!-- mcp-name: io.github.The-40-Thieves/codecalc -->
**codecalc is an offline, self-hosted MCP server that gives an AI agent a
calculator, a code runner, and a logic checker — so it gets a *correct* answer
instead of a guessed one.** It runs code in **31 languages**, does exact
symbolic math, solves SMT/logic problems, and measures complexity, all exposed
as **49 MCP tools**.
**Fastest path:** `uvx 'codecalc[full]' setup --write` registers codecalc with your MCP client automatically. New to MCP, or want more detail first? See [QUICKSTART.md](QUICKSTART.md), or the Install section below.
Three things nobody else offers together cleanly:
- **Offline-core** — ships no model, no API key, no gateway, no telemetry.
The core opens no sockets; network access is opt-in and only where a
specific tool's job needs it (the Piston provider, `install_package`, the
runtime-update tools, executed code unless `no_net`, and a one-time
in-process grammar download on first `analyze_complexity` — full breakdown
in the network-boundary table below).
- **Safe execution of untrusted code** — an opt-in *strict* isolation
boundary (gVisor+Docker on Linux, AppContainer on Windows) layered above
the default rlimit sandbox, fail-closed and attested.
- **Verification tools** — `verify_translation` proves a port to another
language behaves identically, `verify_optimization` proves an optimization
preserved behavior, and `z3_check` proves or refutes logic with an SMT
solver.
## When to use codecalc
Use it when you want a free, local, private, hardened code-runner and
verifier that an MCP agent can call directly — no vendor account, no cloud
spend, nothing leaving the machine except where a tool's job explicitly
requires it.
Reach for something else when you want managed cloud scale instead of
self-hosting (a hosted sandbox like E2B or Modal), or when you're not
self-hosting at all and the model vendor's built-in code interpreter already
covers what you need.
## codecalc vs. the alternatives
codecalc is not a general cloud sandbox and not a vendor code interpreter. It
overlaps with several things and beats them in only one narrow place — forcing a
model to *measure* a claim instead of asserting it. Where that isn't what you
need, one of these is the better tool, and this table says so plainly.
| You want… | Better fit | Why |
|---|---|---|
| To just run some Python/JS quickly, zero setup | Your model vendor's built-in interpreter | Already there, already sandboxed, nothing to install. Anthropic's code-execution tool has internet access "completely disabled" and cannot install packages at runtime; OpenAI's hosted containers have no outbound network access by default, with an org-level `network_policy` allowlist as an opt-in. Both return output artifacts by reference (Anthropic a `file_id` via the Files API, OpenAI a `container_file_citation`) rather than inline (Anthropic code-execution tool docs, https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool; OpenAI shell/container tool guide, https://developers.openai.com/api/docs/guides/tools-shell; both retrieved 2026-09-07) |
| Heavy or multi-tenant workloads, managed scale | A cloud sandbox (E2B, Modal, Daytona) | Per-tenant Firecracker/gVisor isolation codecalc does not claim by default |
| Pure arithmetic or symbolic math, nothing else | A small calculator or SymPy MCP | Lower token cost; none of the 31-language runtime machinery |
| **A model to stop _guessing_ numbers, equivalence, and speedups — locally, privately, with graded evidence** | **codecalc** | Exact rationals, `verify_translation`/`verify_optimization`, and `unenforced`/grade honesty — offline, no account |
**Do not reach for codecalc if** you need multi-tenant or network-exposed
isolation (its threat model is explicitly single-operator, local, stdio), if
zero-setup convenience matters more than measurement, or if a hosted interpreter
already covers your case. It earns its keep only when the *correctness of the
claim* — not just "it ran" — is the point.
## Install
### Quickstart with `codecalc setup`
The fastest path to a working MCP connection, without reading the rest of
this section:
```bash
uvx 'codecalc[full]' setup # prints what it would do — nothing on disk changes
uvx 'codecalc[full]' setup --write # applies it: merges your client's config, copies the skill
```
It detects which MCP client is installed (Claude Desktop, Claude Code,
Cursor, VS Code, Zed — pass `--client=NAME` if none or several are found),
reuses `codecalc doctor`'s own backend/extras/grammar-cache checks, prints the
exact config block in that client's own JSON shape with absolute paths
already filled in, runs two real canaries (`execute_code`, `evaluate_expression`)
to prove the connection would work, and ends in one verdict: `ready` /
`degraded` / `not-ready`. `--write` is the only mode that changes anything —
it MERGES the `codecalc` entry into your existing client config (every other
server stays exactly as it was) and backs up the original to
`<path>.codecalc-bak` first. `codecalc --help` lists every subcommand.
> [!NOTE]
> Published as **`codecalc` 0.12.0** on PyPI (`pip install codecalc`) and the
> **`codecalc-exec` 0.12.0** executor on crates.io
> ([#91](https://github.com/The-40-Thieves/codecalc/issues/91)). Every release
> artifact carries a keyless sigstore **build-provenance attestation** — verify
> one with `gh attestation verify <file> --repo The-40-Thieves/codecalc`; PyPI
> wheels additionally carry PEP 740 attestations.
### Where to find codecalc
| Where | What you get | Link |
|---|---|---|
| PyPI | `pip install codecalc` / `uvx codecalc` | [pypi.org/project/codecalc](https://pypi.org/project/codecalc/) |
| crates.io | the `codecalc-exec` Rust executor crate | [crates.io/crates/codecalc-exec](https://crates.io/crates/codecalc-exec) |
| GitHub Releases | wheels for every platform, the executor binaries, the `.mcpb` bundle, an SBOM, and `SHA256SUMS` | [github.com/The-40-Thieves/codecalc/releases](https://github.com/The-40-Thieves/codecalc/releases) |
| MCP registry (official) | the `io.github.The-40-Thieves/codecalc` server entry `server.json` publishes to | [registry.modelcontextprotocol.io](https://registry.modelcontextprotocol.io/v0/servers?search=codecalc) |
| Smithery | hosted listing and one-click client config | [smithery.ai/servers/@The-40-Thieves/codecalc](https://smithery.ai/servers/@The-40-Thieves/codecalc) |
| Glama | hosted listing and the score badge above | [glama.ai/mcp/servers/The-40-Thieves/codecalc](https://glama.ai/mcp/servers/The-40-Thieves/codecalc) |
| MCPB (Claude Desktop) | the drag-and-drop bundle, attached to every GitHub Release | see GitHub Releases, above |
| Docker MCP Catalog | the `mcp/codecalc` image Docker builds from this repo's `docker/mcp-server.Dockerfile`, for Docker Desktop's MCP Toolkit (`docker mcp server enable codecalc`) | submitted as [docker/mcp-registry#5025](https://github.com/docker/mcp-registry/pull/5025); listed at [hub.docker.com/mcp/server/codecalc](https://hub.docker.com/mcp/server/codecalc) once merged |
**Not yet listed:** PulseMCP and mcp.so do not carry a codecalc entry yet; the Docker MCP Catalog entry is pending review.
PulseMCP's own submission page (checked 2026-09-08) says it is not accepting
new submissions and that publishing to the official MCP registry — already
done, row above — is what it indexes from once submissions reopen, so there
is nothing to submit there today. mcp.so takes a submission through its own
form. See [docs/distribution.md](docs/distribution.md) for the exact steps,
kept there rather than here because submitting is an action for whoever runs
it, not a fact about the current release.
**The published install** (simplest — no build step, and what most people want):
```bash
uvx 'codecalc[full]' # run it directly, no environment to manage
# or
pip install 'codecalc[full]' # into your own virtualenv
```
**From source**, if you would rather build the executor yourself:
```bash
git clone https://github.com/The-40-Thieves/codecalc
cd codecalc
uv sync --all-extras # or: pip install -e '.[full]'
cargo build --release --manifest-path executor/Cargo.toml
mkdir -p bin # bin/ is gitignored, so a fresh clone has none
cp executor/target/release/codecalc-exec bin/
uv run codecalc doctor # verify: backend should read `rust`
```
Without the `cargo build`, everything still runs on the pure-Python fallback —
`doctor` will say so, and the network table below says what that costs.
**Why `[full]`.** The base install is the MCP surface and the sandbox executor:
31 language runtimes, sessions, packages, ~32 MB. The symbolic half — sympy and
z3 — is 88.6 MB measured, and a caller who only runs code should not download an
SMT solver to do it. So it is an extra:
| install | size | what you get |
|---|---|---|
| `codecalc` | ~32 MB | execute_code, sessions, packages, complexity-free tools |
| `codecalc[symbolic]` | +83 MB | evaluate_expression, solve, limits, truth tables, z3, units |
| `codecalc[parsing]` | +5 MB installed, **+89 MB fetched on first use** | analyze_complexity via tree-sitter |
| `codecalc[full]` | ~120 MB | everything |
Nothing fails silently: a tool whose extra is missing returns
`{"ok": false, "error": "sympy is not installed. It ships in the 'symbolic'
extra: pip install 'codecalc[symbolic]' ..."}`, and `codecalc doctor` lists
which extras are present before you make a call.
### Editions
Four names for the capability sets above, plus the two that live outside
`pyproject.toml` entirely — a Docker image and an opt-in isolation boundary.
**The invariant that makes "edition" a meaningful word here:** in the edition
that lists a tool, that tool is functional — never listed-but-missing-its-extra.
A tool an edition doesn't have returns the `dependency_missing` contract error
naming the extra that provides it (see above), not a silent failure or a
tool that appears to exist and doesn't work.
| Edition | Install | What you get |
|---|---|---|
| **Full** | `uvx 'codecalc[full]'` / `pip install 'codecalc[full]'` | The recommended local product: the native (Rust) executor, symbolic tools (`evaluate_expression`, `symbolic`, `z3_check`, …), and parsing (`analyze_complexity`). Everything this README documents actually runs. |
| **Core** | `uvx codecalc` / `pip install codecalc` | Execution + non-symbolic tools only — the base install in the table above. Every symbolic/parsing tool is still *listed* by `tools/list` (MCP doesn't support per-install schemas), but calling one returns `dependency_missing` naming the extra, before any other work happens. |
| **Docker** | `docker build -f docker/mcp-server.Dockerfile .` | The MCP server itself, packaged to run as an ordinary container. Core-shaped **by default**: ships `python3`/`node`/`ruby`/`php`/`perl`/`gawk`/`lua`/`c`/`cpp`/`jq`/`sqlite3` and the default rlimit sandbox — symbolic/parsing are absent by design (no `[full]` in the base image; see the Dockerfile's own comment for why, including an arm64 z3-solver wheel gap). `--build-arg CODECALC_EXTRA=full` adds them. This image cannot nest the Strict Host boundary below inside itself (no privileged docker-in-docker), and `codecalc doctor` inside it says so rather than claiming a boundary it doesn't have. |
| **Strict Host** | opt-in; `CODECALC_STRICT_URL` (client) or the gVisor+Docker host itself (server) — see [`docs/deployment/README.md`](docs/deployment/README.md) | Not an install, a *boundary*: the gVisor `runsc` sandbox on Linux, or AppContainer hardening on Windows, layered above whichever install above is already running. Fails closed — no digest pinned, no fallback to unenforced local execution. |
`codecalc doctor` reports which of these you're actually running (backend,
extras present, `strict_runtime` prerequisites) — read it before assuming a
capability rather than after a tool call surprises you.
`.github/workflows/release.yml` publishes a platform-tagged wheel per target
(Linux x86_64/aarch64 musl, macOS x86_64/aarch64, Windows x86_64), each
carrying the matching `codecalc-exec` binary and — where the platform has
one — its `--no-net` shim, so `executor.backend() == "rust"` on install
without a manual build step. No wheel for your platform, or installed from
source instead? Everything still runs; see the network table above for what
falls back and to `unenforced` in that case.
Point an MCP client at the installed command. **The key differs by client** —
`mcpServers` for most, `servers` for VS Code, `context_servers` for Zed — so
these are given separately rather than as one snippet to adapt:
**Claude Desktop** — `~/Library/Application Support/Claude/claude_desktop_config.json`
(macOS), `%APPDATA%\Claude\claude_desktop_config.json` (Windows) · **Cursor**
(`.cursor/mcp.json`) and **Claude Code** (`.mcp.json`) use the same shape:
```json
{ "mcpServers": { "codecalc": { "command": "uvx", "args": ["codecalc[full]"] } } }
```
**VS Code** — `.vscode/mcp.json`, top-level key is `servers`:
```json
{ "servers": { "codecalc": { "command": "uvx", "args": ["codecalc[full]"] } } }
```
**Zed** — `~/.config/zed/settings.json`, key is `context_servers`:
```json
{ "context_servers": { "codecalc": { "command": "uvx", "args": ["codecalc[full]"], "env": {} } } }
```
**Windows paths need doubled backslashes** in JSON. If you installed into a venv
rather than using `uvx`, point at the interpreter directly:
```json
{ "mcpServers": { "codecalc": {
"command": "C:\\path\\to\\venv\\Scripts\\python.exe",
"args": ["-m", "codecalc"] } } }
```
Run `codecalc doctor` to print a config block with the absolute paths of *your*
install already filled in.
**Install the skill too.** The tools cannot help a model that never reaches for
them — a model confident about `0.1 + 0.2` does not feel uncertain, it feels
finished. `codecalc/SKILL.md` ships inside the package and says when calling is
mandatory (any non-integer, any comparison you will state, anything past 2^53,
any number stated as a claim), when it is noise (`2 + 3 + 4` needs no tool), and
how results must be reported — `passed: true` means "equivalent on N inputs",
never "verified". `codecalc doctor` prints its path; copy it into your client's
skills directory. `check_claims.py` gates it, so it cannot name a tool that does
not exist or a field no tool returns.
Not sure what your install actually resolved? Ask it, rather than finding out
from a tool call later:
```bash
codecalc doctor # or: python -m codecalc doctor
```
**This is the install verification step.** It exits `0` when the install can
execute — a writable workspace and a resolved backend — and `1` when it cannot,
so it works unchanged in a Dockerfile, a provisioning script or a CI job. A
missing optional extra or an uninstalled Haskell does **not** fail it: those are
facts about the host, not a broken install, and a check that goes red for them
is one people learn to ignore.
It prints the execution backend and the binary behind it, whether installs are
confined, the status of every one of the 31 runtimes, whether the workspace is
writable, and a client config block with absolute paths filled in. All of that
is otherwise discoverable only by making a tool call and reading `backend`,
`unenforced`, or a failure.
```bash
codecalc doctor --json # the same report, for scripts
codecalc doctor --deep # actually RUN each runtime, and read its version
```
`--json` emits the report and nothing else, against a published schema
([`docs/contract/doctor-v1.schema.json`](docs/contract/doctor-v1.schema.json))
carrying the same `contract_version` and the same policy as a tool result.
Each runtime reports one of four states, and the difference between two of them
is which measurement was actually taken:
| state | means |
|---|---|
| `supported` | codecalc knows the language; nothing for it resolves here |
| `installed` | its command resolves and is executable — **not run** |
| `unhealthy` | resolves but cannot run, or was run and failed |
| `available` | actually executed here and answered — `--deep` only |
`status_basis` says which pass produced them. Without `--deep` nothing is ever
reported `available`, because nothing was executed, and claiming otherwise for a
binary that was merely found on `PATH` would be a stronger measurement than was
taken.
Under `--deep`, a runtime whose version probe never gets an answer (a spawn
failure or a timeout) is `unhealthy` too; a nonzero exit alone only counts
when the flag used is one confirmed correct for that command (`go version`,
`lua -v`, `zig version` — none of them speak GNU `--version`, so a bare
nonzero exit there is reported as merely unmeasured, not broken). A captured
failure lands in `probe_error`, never in `version`, which holds a version
string or nothing. A compile-then-run language whose `run` step needs a
SECOND, different tool (kotlin: `kotlinc` compiles, but `run` launches
`java` directly) reports `installed` only when BOTH resolve; `detail` names
whichever half is missing.
Building the Rust core yourself, or running from a checkout? See "Build the
Rust core" and "Run the server" below.
### Use it from an MCP client
One-click install: both buttons register `uvx codecalc[full]` (the
recommended Full edition) and require `uv` to be installed.
[](cursor://anysphere.cursor-deeplink/mcp/install?name=codecalc&config=eyJjb21tYW5kIjoidXZ4IiwiYXJncyI6WyJjb2RlY2FsY1tmdWxsXSJdfQ==)
[](vscode:mcp/install?%7B%22name%22%3A%22codecalc%22%2C%22type%22%3A%22stdio%22%2C%22command%22%3A%22uvx%22%2C%22args%22%3A%5B%22codecalc%5Bfull%5D%22%5D%7D)
[](https://glama.ai/mcp/servers/The-40-Thieves/codecalc)
The shortest version of the config above — this registers codecalc as a
stdio MCP server. The console entry point is `codecalc`, so `uvx codecalc`
launches it directly:
```json
{
"mcpServers": {
"codecalc": { "command": "uvx", "args": ["codecalc"] }
}
}
```
Installed with `pip install codecalc` instead? Point at the resolved
command with no args:
```json
{
"mcpServers": {
"codecalc": { "command": "codecalc" }
}
}
```
## Network boundary
**CodeCalc's core opens no sockets.** No model gateway or telemetry is built
in. `tests/test_offline.py` asserts this for the top-level core modules. The
opt-in Piston provider is the deliberate exception: its wire client lives under
`codecalc/provider_adapters/` and is registered only when
`CODECALC_PISTON_URL` is configured.
That is a claim about the **package**, not about every tool call, and the
difference is worth stating rather than leaving a reader to discover:
| layer | reaches the network? |
|---|---|
| CodeCalc core | **No HTTP client, model gateway, or telemetry.** One dependency exception: `analyze_complexity` may download a grammar on first use (see below) |
| configured Piston provider | **Yes, explicitly.** Calls only the operator-supplied `CODECALC_PISTON_URL`; credentials stay in its authorization header and are redacted from results |
| `install_package` | **Yes, by design.** It runs uv / npm / gem / cargo, which fetch from their registries. Installer hooks also run *outside* the sandbox — see [SECURITY.md](SECURITY.md) |
| `runtimes_status`, `update_runtimes` | **Yes.** They shell out to mise / rustup / swiftly / npm, which check remote versions |
| code you execute | **Yes, unless `no_net=True`** — and that guarantee needs the native executor (seccomp-bpf where the Linux kernel supports it, a symbol shim otherwise; see the guarantee table below), so the pure-Python fallback reports it in `unenforced` instead of applying it. Set `CODECALC_REQUIRE_NATIVE=1` to turn "fallback in use" into a startup failure instead of a result you have to notice by reading `unenforced` |
| `execute_code` / `session_run` / `execute_code_stream` / `run_submit` with declared `dependencies` | **Yes, before the sandboxed step, through the confined `install_package` path.** A PEP 723 block (python3) or the `dependencies` argument is installed BEFORE the code runs — never inside the sandbox — and refused (`capability_not_requested`, no fetch attempted) when `no_net=True` was requested or the capability policy denies or strictly limits network. `run_submit`'s install runs on its own background worker, same as the code that follows it — the call itself still returns a run_id immediately |
These distinctions are stated precisely on purpose: a guarantee described more
broadly than it is enforced is exactly the failure mode this project works to
avoid, so "offline-core" is scoped to what the structural test can actually
support rather than claimed as a blanket "no network calls".
**A PEP 723 block alone, with no `dependencies` argument, can trigger the
install above.** `execute_code`/`session_run` read the block out of the source
text itself — a caller who passes no `dependencies` argument at all still gets
a confined `uv`/`npm` subprocess and real egress if the code they submit
happens to carry a `# /// script` block, whenever the refusal rule above does
not apply. This is logged distinctly (`dependency_install_implicit` in the
audit trail, alongside `install_denied`) so an operator can tell "source text
alone triggered this" from an explicit `install_package`/`dependencies=` call.
To disable it: `no_net=True` on the call, or a `deny-network`/`strict`
`CODECALC_CAPABILITY_POLICY` — either one refuses before any fetch, block or
no block. `execute_code_stream` and `run_submit` read the block the same way
`execute_code` does. `compare_execution` is the one holdout: it fans out
across several languages with no per-language install plumbing behind it, so
it REJECTS an explicit `dependencies` argument with a `validation` error
rather than approximating one, and DISCLOSES rather than silently drops an
inline PEP 723 block it finds in a snippet — that row's result carries
`dependencies: {"status": "unsupported", "reason": ...}` instead of
installing from it.
**Two ceilings govern a dependency-bearing run, not one.** The run's own
`timeout` bounds the sandboxed step; it says nothing about installing
dependencies FIRST, outside the sandbox. A separate, fixed budget
(`codecalc.dependencies.DEFAULT_DEPENDENCY_INSTALL_BUDGET_SECONDS`, 120s,
aggregate across every dependency of one run) bounds that step instead —
exceeding it refuses the run with a stamped `timeout` naming the budget,
before the run's own `timeout` clock even starts. A sessionless run's
dependency workdir is also held to a disk quota — reusing
`CODECALC_SESSION_DISK_QUOTA_MB` (below), the same cap a session workspace
already has — and a run that grows past it after a successful install is
refused with a stamped `resource_exhausted` naming the measured size and the
cap.
**The grammar download, stated plainly, because it is the one that is easy to
miss.** The other three paths above go through a CHILD PROCESS, which is what
`tests/test_offline.py` says it cannot see. This one does not:
`tree-sitter-language-pack` ships a ~5 MB extension and fetches each grammar on
first use, **in-process**, into a local cache — 28 grammars, 89 MB, about 15
seconds on a cold cache. So the first `analyze_complexity` call for a given
language opens a socket from inside the server.
It is verified (the pack checks a signature and raises on a checksum mismatch),
it is cached, and it never happens again for that language. So the offline-core
claim is scoped to steady state: this first-use grammar fetch is the one
in-process exception, which is why it is called out here rather than glossed
over.
**For an offline or egress-restricted install**, warm the cache first — it is one
command, and afterwards nothing here reaches the network. If you installed
codecalc (`pip install`/`uvx`, not a source checkout), `scripts/` did not come
with it, so use the shipped console script instead:
```bash
codecalc-prefetch-grammars # installed: fetch all 28 grammars
codecalc-prefetch-grammars --print-cache-dir # installed: the directory to copy
```
Building from source? The script still works and calls the same code:
```bash
python scripts/prefetch_grammars.py # fetch all 28 grammars
python scripts/prefetch_grammars.py --print-cache-dir # the directory to copy
```
`codecalc doctor` reports whether that cache is populated, so this is
discoverable before it matters rather than after a tool call degrades.
## Architecture (language-per-strength)
| Layer | Language | Why |
|---|---|---|
| Executor core (`executor/`) | **Rust** | Sandbox + rlimits + process-group kill + JSON CLI. No `eval()` anywhere near user input; memory-safe host; single static binary |
| Logic layer (`codecalc/logic.py`) | **Python** | sympy (symbolic math, equation solving) and z3 (SMT) have no Rust equivalents |
| MCP server (`codecalc/server.py`) | **Python** | the official `mcp` SDK (2.0) generates tool schemas from type hints; protocol **2026-07-28** |
Python orchestrates; Rust executes; sympy/z3 reason. Each layer does what it's
best at. The Rust binary is preferred automatically; a pure-Python executor is
the fallback if the binary is missing.
## Older-computer support
- **No modern instruction-set requirements** — rustc targets a generic CPU by
default and nothing overrides it. (`executor/.cargo/config.toml` explains why
`-C target-cpu=generic` is deliberately NOT written there: it would be a
no-op that reads like a guarantee.)
- **Static musl builds** run on any Linux regardless of glibc version:
`bin/codecalc-exec-x86_64-musl`, `bin/codecalc-exec-aarch64-musl` (~430K each;
the exact size moves with every toolchain bump, so it is not pinned here)
- Size-optimized profile (`opt-level="z"`, LTO, panic=abort, stripped) —
**measured, not assumed**: against an otherwise identical `opt-level=3` build,
`z` came out 1.02 ± 0.26 times faster on the executor's own path (i.e. no
detectable difference) while being 16% smaller. The executor spends its time
in syscalls, not arithmetic, so there was nothing for a higher optimisation
level to speed up.
- **Lazy sympy/z3 imports.** Both are imported on first use, so a session that
only executes code never pays for them. This claimed "~40ms, not ~600ms" for a
long time while being wrong in both directions: the server took **1.9s** to
start, and sympy was not actually lazy — `units.py` imported it at module
scope and `server.py` imports `units`, so every start paid 437ms for it.
Deferring that took spawn-to-first-response from **1888ms to 1243ms**
(measured, median of 7). The remaining ~870ms is the `mcp` SDK's own import,
which is not ours to remove.
- **The fork-bomb measurement is taken once, and only when it is needed.**
Sizing `RLIMIT_NPROC` means reading `/proc/<pid>/status` for every process on
the machine. That walk used to run during argument parsing and again for every
step: a C compile-and-run opened 1767 status files on a 590-process box to
answer one question three times, and `--lang notalanguage` paid the full cost
to produce a one-line error. Measured lazily and cached, an error costs 1.1ms
instead of 13.3ms and a compiled run 78ms instead of 104ms.
- `list_languages` probes runtime availability and reports which languages
actually work on the machine (graceful degradation on minimal installs)
## Build the Rust core
```bash
cd executor
cargo build --release # native
cargo zigbuild --release --target x86_64-unknown-linux-musl # static x86_64 (uses zig)
cargo zigbuild --release --target aarch64-unknown-linux-musl # static arm64
# Copy the executable AND its --no-net shim together. build.rs rebuilds the
# shim whenever blocknet.c changes, but the executor looks for it beside the
# BINARY, so installing only the binary leaves the previous shim in place — and
# a stale shim silently enforces the old policy while every "is it there?"
# check still passes. Copy both or neither.
mkdir -p ../bin # bin/ is gitignored, so a fresh clone has none
cp target/release/codecalc-exec target/release/blocknet.so ../bin/
```
Requires: Rust 1.97+, a C compiler for the `--no-net` shim (the build warns and
carries on without one; on macOS, or a Linux kernel without seccomp support,
`--no-net` then reports itself in `unenforced` rather than pretending — a
Linux kernel with seccomp support enforces it in-kernel either way), and
[cargo-zigbuild](https://github.com/rust-cross/cargo-zigbuild)
for the static cross-builds (zig is used as the linker; no x86_64 GCC needed).
## MCP tools (49) + MCP resources
Every session file is also exposed as an MCP resource:
`codecalc://session/<session_id>/files/<path>` — images render inline for the
model, text returns as text, other files download.
**Graphical results in MCP Apps hosts**: `verify_translation` and
`verify_optimization` also carry an [MCP Apps](https://modelcontextprotocol.io)
`ui://` view (`ui://verify-translation/view.html`,
`ui://verify-optimization/view.html`) — a per-case diff table and a per-size
timing chart, respectively, rendered inline by a host that supports the
extension. Both are self-contained (inline CSS/JS, no network, no external
assets); a host without MCP Apps support sees exactly today's text/structured
result, unchanged.
**Exact arithmetic & programmer-mode**: exact rationals, threshold checks, bit
analysis, binary64 introspection.
| Tool | Description |
|---|---|
| `calc_exact` | EXACT arithmetic: `0.1+0.2 == 0.3` is True; arbitrary-precision ints, bitwise ops inline, whitelisted math funcs, pi/e/tau |
| `compare_threshold` | Exact threshold verdict with shortfall: `('1/25', '>', '0.05')` → False, shortfall 1/100 |
| `percentage` | Exact share and percentage of PART/TOTAL (rationals accepted) |
| `calc_stats` | mean, median, sample stdev, **CV** (CV > 0.2 = noise swamps the effect) |
| `percentiles` | p50/p90/p95/p99 by nearest-rank AND interpolation; warns n<100 |
| `collision_probability` | Birthday-bound hash collision: 1e5 items/32 bits ≈ 0.69, 1e6/64 ≈ 2.7e-8 |
| `data_sizes` | Byte sizes both ways: KiB/MiB (binary) AND KB/MB (decimal) |
| `human_duration` | Humanised duration + per-day/per-30d rates |
| `epoch_time` | Epoch s/ms/µs/ns → ISO 8601 UTC, implausible readings suppressed |
| `radix_convert` | Any base 2..36, fractions included, non-termination flagged (`0.1` base 2) |
| `float_repr` | What binary64 actually stores: exact value, raw bits, ULP, neighbours, representable-or-not |
| `bits` | Programmer-mode integer facts and operations, selected by `mode`: `analysis` (was `bit_analysis`), `op` (was `bitop`), `widths` (was `int_widths`), `repr` (was `base_repr`) — those four former standalone tools were retired in 0.12.0 (CHANGELOG.md) |
| `algebraic_equiv` | Are `(a*b)/c` and `a*(b/c)` identical? refactor verification (with float/truncation caveat) |
| `symbolic` | Symbolic algebra, selected by `op`: `solve` (was `solve_expression`), `solve_linear` (was `solve_linear`, name unchanged), `simplify` (was `simplify_expression`), `limit` (was `limit_expression`) — those four former standalone tools were retired in 0.12.0 (CHANGELOG.md) |
**Core tools**
| Tool | Description |
|---|---|
| `list_languages` | 31 languages with extension, compile flag, runtime availability |
| `list_execution_providers` | Execution-provider identity, interface version, host class, and machine-readable capabilities |
| `execute_code` | Run code in any language → stdout/stderr/exit_code/**verdict** (OK/TLE/MLE/OLE/RTE)/cpu_ms/peak_memory_kb; per-call limits (`max_memory_mb`, `max_output_kb`, `max_cpu`), `no_net`, `compact`. With a session and no explicit `max_output_kb`, oversized output **spills** into the session workspace (`stdout_spill`/`stderr_spill`) instead of just truncating |
| `execute_code_stream` | Provider-selected execution using the same canonical limits as `execute_code`, with progress + partial output when the provider supports streaming |
| `trace_execution` | **python3 only.** Runs the same sandboxed executor `execute_code` uses, plus a per-line event trace (`events`: line/call/return/exception, changed locals per step) and a static branch/line-coverage report (`branches`, `lines_executed`, `lines_never_executed`) from an AST parse — answers "which lines ran, in what order, and why" rather than just "what did it print" |
| `branch_reachability` | **python3 only.** Decides, with z3, which if/elif/else arms and while/for(range, static bounds) loops in ONE function can ever be taken for ANY input — `reachable`/`dead`/`unknown` per branch, a `witness` when reachable, and `boundary_inputs` (min/max/equality-edge, via z3 Optimize) shaped for `compare_edge_cases`'s `test_inputs`. Refuses up front, naming the construct and line, for anything outside `+ - * // %`/`and or not`/`== != < <= > >=`/`abs min max len` on int/bool/str |
| `run_submit` | Submit code for **background execution**; returns a `run_id` immediately instead of holding the call open |
| `run_inspect` | Poll a background run: status while running, the full `execute_code` result shape once terminal |
| `run_cancel` | Cancel a background run; idempotent on an already-terminal run, honest about providers that cannot cancel mid-flight |
| `session_start` | Persistent session; python3/node get a stateful REPL worker (variables/imports persist across calls), other languages a workspace dir |
| `session_stop` / `session_list` | Session lifecycle |
| `session_files` / `session_read_file` / `session_write_file` | Workspace file tools, jailed to the session dir; listings support `page_size`/`cursor`, and reads return images inline (`as_image`) |
| `session_run` | **Multi-file programs**: execute an entry file that imports other session files (helper.py, data/...) in the workspace |
| `session_artifacts` | List files created by executed code (results, images, CSVs) |
| `session_snapshot` | Archive a session's workspace to a snapshot stored OUTSIDE the jailed workspace (`action="save"`), or restore one into a new session or, with `replace=True`, back into the same session (`action="restore"`); `action="list"`/`"delete"` manage them. Files only — never a stateful session's REPL variables. Snapshots die with `session_stop` unless `keep_snapshots=True` |
| `install_package` | Install packages (uv pip/npm/gem/go/cargo...) into a session or shared cache |
| `verify_translation` | **Prove a port is equivalent**: you write the translation, the executor runs both versions on the same inputs and reports match / diverged / inconclusive per input. A pass is graded `cross_checked` (see [Grade vocabulary](#grade-vocabulary)) |
| `verify_optimization` | **Prove an optimisation**: you write the candidate, the executor confirms it still agrees with the original AND times both — accepted only if equivalent and measurably faster. Accepted is graded `cross_checked` |
| `extract_function` | Pull a named function + its dependency closure (imports, referenced helpers) into a standalone program and run it (ast-exact for python3, best-effort elsewhere) |
| `compare_edge_cases` | Run the same logic in N languages on edge-case inputs (empty, zero, negative, float precision) and flag behavioral divergence |
| `convert_units` | Dimensional unit conversion via sympy: length, mass, time, speed, energy, power, force, pressure, temperature (°C/°F/K), volume, area, data, frequency |
| `physical_constants` | 22 physical constants with values (c, h, N_A, k_B, G, g, m_e, R, ...) |
| `list_units` | All 140+ unit aliases for convert_units |
| `evaluate_expression` | Symbolic math: `integrate(x**2, x)`, `sqrt(144) + 2**10` |
| `truth_table` | Boolean algebra: `a and b or not c`, `p xor q`, `a implies b` |
| `z3_check` | SMT-LIB2 satisfiability + model. An `unsat` verdict is graded `solver_proven`; `sat` is graded `ungraded` (decided, but not proof-shaped — see [Grade vocabulary](#grade-vocabulary)) |
| `matrix` | Structured matrix ops: det/inverse/eigenvalues/transpose/rank/trace on a `rows` array — never a caller string through sympify, so `evaluate_expression`'s `[`/`]` RCE screen never applies. Each entry screened individually |
| `analyze_complexity` | Static Big-O estimate from code structure, parsed with **tree-sitter** (every supported language). Reports `analysis: tree-sitter\|regex-fallback` so you can tell a parse from a guess |
| `benchmark` | Empirical Big-O: runs code at increasing N, fits growth curve |
| `compare_execution` | Same code across N languages side-by-side |
| `runtimes_status` | **Non-mutating** update check: current vs latest for every language runtime, which package manager owns it, and the command that would run |
| `update_runtimes` | Update runtimes. **Dry-run by default** (`apply=False` returns the commands); `apply=True` executes them |
### Retired tool aliases
The 2026-09-08 merge (0.11.0) folded two lexically-overlapping clusters —
flagged by Glama's public review as indistinguishable from a description
alone — into one enum-selected tool each, keeping every mode's own
parameters, annotations and result shape. The eight old names stayed
registered as thin aliases for one minor release so nothing broke
mid-upgrade, then were removed in 0.12.0 (CHANGELOG.md). Calling one of
them now gets the MCP SDK's own unknown-tool error, not a result:
| Retired name | Replacement |
|---|---|
| `bit_analysis` | `bits(mode="analysis")` |
| `bitop` | `bits(mode="op")` |
| `int_widths` | `bits(mode="widths")` |
| `base_repr` | `bits(mode="repr")` |
| `solve_expression` | `symbolic(op="solve")` |
| `solve_linear` | `symbolic(op="solve_linear")` |
| `simplify_expression` | `symbolic(op="simplify")` |
| `limit_expression` | `symbolic(op="limit")` |
## Grade vocabulary
`verify_translation`, `verify_optimization` and `z3_check` return `grade` +
`grade_basis` (+ `grade_rules_version`) on top of their own result. The grade
names how strong the evidence for a success actually is; it is derived from
evidence those tools already emit, in `codecalc/grades.py` — the verifiers
never assign their own grade.
| Grade | Means | Emitted by |
|---|---|---|
| `cross_checked` | Two independently authored programs were both actually run and their outputs agreed. `grade_basis` names the runtime(s) that did the checking. | `verify_translation` (source vs. port), `verify_optimization` (original vs. candidate) |
| `solver_proven` | Z3 returned `unsat` within its timeout — a machine-checked refutation, not a heuristic. `grade_basis` names the engine version and the timeout bound. **Not** `sat`: see below. | `z3_check` |
| `executed` | Reserved: the claimed computation ran and produced the reported result, with no independent second opinion. Not currently emitted by any tool above — every one of them also clears the `cross_checked`/`solver_proven` bar. | — |
| `ungraded` | Explicit non-grade for a mismatch, an inconclusive comparison, a rejected optimisation candidate, a measurement failure, a Z3 `unknown` verdict, and — deliberately — a Z3 `sat` verdict. A real value on `grade`, never an absent key. **Never** a softened stand-in for one of the three grades above. | any of the above, on a non-success |
`z3_check`'s `sat` verdicts are graded `ungraded`, not `solver_proven`, even
though `sat` is just as decisive a verdict as `unsat`. The ticket's motivating
pattern is proving a property P by asserting not-P and checking `unsat`; a
caller running that pattern who gets `sat` back has learned P is FALSE, and
`solver_proven` on that result would let a reader who skims `grade` without
`result` mistake a counterexample for a proof. `sat`'s `grade_basis` says so
explicitly: satisfiability was decided, but `solver_proven` is reserved for
`unsat` so a counterexample can never wear a proof grade. Widening `sat` back
into `solver_proven` later is additive; narrowing it after callers depend on
the wider behaviour would not be, so this ships narrow now. Full reasoning:
`codecalc/grades.py`'s module docstring.
`algebraic_equiv` is deliberately NOT graded: it compares two expressions via
`sympy.simplify(a - b) == 0`, a CAS transformation rather than a decision
procedure with a checkable certificate, and it is one simplifier's opinion
rather than two independent implementations agreeing. None of the three
grades describes that evidence honestly.
## Runtime self-update
Every language is mapped to its package manager, and codecalc can update its own
runtimes:
| Manager | Languages | Update command |
|---|---|---|
| mise | python3, node, bun, deno, ruby, go, erlang, elixir, gleam, zig, java, kotlin, sqlite, duckdb, gradle | `mise up` |
| rustup | rust (stable/nightly toolchains) | `rustup update` |
| swiftly | swift | `swiftly update` |
| apt | c, c++, fortran, csharp, php, perl, lua, tcl, r, jq, bash, zsh | `apt-get install --only-upgrade` (language packages only) |
| npm | typescript/tsc | `npm update -g` |
| uv | mojo | `uv tool upgrade mojo` |
| nix | haskell (on-demand) | nothing persistent |
`runtimes_status` is always safe. `update_runtimes` refuses to mutate unless
`apply=True` is passed explicitly — and it only touches the package manager
that owns each language (never the Rust sandbox, which has no update powers).
One of those managers is elevated: apt updates system packages, so its command
starts with `sudo`. `apply=True` is an argument a connected model controls, so
that branch takes a second key the model does not have — the host must set
`CODECALC_ALLOW_RUNTIME_APPLY=1`. Without it the apt command is reported as
skipped with `ok: false` and the variable named, while the unprivileged managers
still run. `sudo -n` already fails closed where a password is required; this
covers the passwordless-sudo rule common on developer machines and CI images,
which is exactly where `-n` does not stop it.
## Run the server
```bash
cd /path/to/codecalc && .venv/bin/python -m codecalc.server
# stdio transport — register with any MCP client
# The identical tool/resource registry over stateless Streamable HTTP:
.venv/bin/python -m codecalc.server serve-http --host 127.0.0.1 --port 8000
```
Streamable HTTP binds to loopback by default. Bearer-token auth
(`CODECALC_HTTP_TOKEN`) is **required** for any non-loopback bind — `serve-http`
refuses to start on a routable address if neither it nor `--oauth-issuer` (below)
is set, and the static-token comparison is constant-time — and **optional** on
loopback, where an MCP client spawning the process is already inside the trust
boundary. Setting a token does not change the single-operator threat model: put
an authenticating reverse proxy and the stronger process/container isolation
described in `SECURITY.md` in front of it before exposing it beyond one
operator's own machine.
For hosted use only, `serve-http` also accepts `--oauth-issuer URL` (or
`CODECALC_OAUTH_ISSUER`) as an alternative to the static token — **off by
default**; the static-token path above is unchanged when it is unset, and
setting the variable costs nothing outside `serve-http` itself: `doctor`,
`--help`, `serve-strict`, and the bare stdio server never touch the network
over it, only `serve-http`'s own startup does. The issuer and JWKS URLs must
be `https://` unless the host is loopback (for local testing); an issuer that
is plain `http://` on a real host, or cannot be reached at all, fails
`serve-http`'s startup outright with a message on stderr rather than starting
a server no token could ever pass. Given a reachable issuer, `serve-http`
validates each bearer token as a JWT against that issuer's own JWKS
(RS256/ES256; the JWKS URL is discovered once from
`<issuer>/.well-known/openid-configuration`, or pinned with
`--oauth-jwks-url`) and checks its issuer, audience, expiry, and
not-before. It also serves RFC 9728 Protected Resource Metadata at
`/.well-known/oauth-protected-resource/mcp`, and a request with a missing or
invalid token gets `WWW-Authenticate: Bearer resource_metadata="..."` pointing
at it, per the MCP authorization spec (2025-06-18 and later). `--oauth-audience`
defaults to this server's own resource URL; `--oauth-scopes "s1 s2"` requires
every named scope on the token, checked by the SDK's own auth middleware. If
both a static token and an issuer end up configured at once, **the issuer
wins** — a request bearing the static token's exact value is rejected like any
other invalid bearer value, and a warning naming both settings is printed to
stderr at startup. codecalc runs no `/authorize` or `/token` endpoint of its
own (it is a resource server only, never an authorization server), so there is
no dynamic client registration surface here either; a 2026-07-28-era client
that needs one uses a Client ID Metadata Document against its OWN
authorization server, not against codecalc.
Point an MCP client at it:
```json
{ "mcpServers": { "codecalc": { "command": "/path/to/codecalc/.venv/bin/python",
"args": ["-m", "codecalc.server"],
"env": {
"PYTHONPATH": "/path/to/codecalc",
"CODECALC_RUNTIME_PATH": "/path/to/mise/shims:/usr/local/bin:/usr/bin:/bin"
} } } }
```
## MCP protocol
Protocol revision **2026-07-28**, on the official `mcp` SDK 2.0. Not fastmcp:
fastmcp 3.x pins `mcp>=1.24,<2.0` and so cannot reach this revision at all.
Verifying that is less obvious than it looks. `mcp.types.LATEST_PROTOCOL_VERSION`
reads `2026-07-28` regardless of what a given connection negotiated, and the
*same server* answers on either protocol depending only on how you connect:
| client | negotiated | cache hints |
|---|---|---|
| `ClientSession.initialize()` | `2025-11-25` | dropped |
| `Client(..., mode="auto")` | **`2026-07-28`** | applied |
So `tests/test_mcp_protocol.py` asserts the negotiated value from a real
connection. The legacy path still works — backward compatibility is a feature —
it just must not be mistaken for the new protocol.
Worth noting for anyone reading the spec's headline change: 2026-07-28 removes
protocol-level sessions, and directs servers needing cross-call state to use
"explicit, server-minted handles passed as ordinary tool arguments". That is
exactly what codecalc's `session_id` already is.
## The result contract
Every result carries `contract_version`, currently **1.6.0**. The published
schema is [`docs/contract/result-v1.schema.json`](docs/contract/result-v1.schema.json)
and the policy behind it — what MAJOR/MINOR/PATCH may change, the twelve-month
deprecation window, worked success/failure/timeout examples, and the migration
path from unversioned servers — is in
[`docs/contract/README.md`](docs/contract/README.md).
For in-process Python use, the supported protocol-neutral service boundary—and
the session/storage internals that are deliberately not public—is documented in
[`docs/embedding.md`](docs/embedding.md).
Two things a caller should know before reading anything else:
- **`ok` means "ran and exited 0".** A program that behaves exactly as intended
and exits 3 comes back `ok: false`, `exit_code: 3`, `verdict: "RTE"`. To tell
a failed *program* from a failed *request*, read `verdict` — a request that
never reached a runtime has no `verdict` at all, and has a `code` instead.
- **`code` is the branch target, not `error`.** Eight stable values; the prose
in `error` is free to improve and is not a contract. An unrecognised `code`
must be treated as `internal` — that is what lets a `1.x` client survive a
`2.0.0` server, though adding a code is still a MAJOR change, because the
published enum is closed and a strict validator rejects the result first.
- **Truncation reports a size, not just a flag.** `output_truncated` says output
was cut; `stdout_bytes` / `stderr_bytes` say by how much — the bytes the
program actually produced, before the cap. A 200 000-character `print` under
`max_output_kb=1` returns 1 039 bytes of `stdout` and `stdout_bytes: 200001`,
so a caller can size a retry instead of guessing. `null` there means not
measured (nothing ran); a program that printed nothing reports `0`.
The schema is JSON Schema 2020-12 — the dialect MCP 2026-07-28 defaults tool
`outputSchema` to — so a client can validate our results with it directly.
`scripts/check_contract.py` regenerates it from `codecalc/contract.py` and fails
on a diff, and separately re-derives both backends' verdict vocabularies from
`main.rs` and `executor.py`: `check_parity.py` compares the two backends' *key
sets* and is structurally blind to a new verdict *value*, which would leave the
published enum short and make a strictly validating client reject a good result.
## Configuration
All optional. codecalc runs with none of these set.
| Variable | Default | What it does |
|---|---|---|
| `CODECALC_HTTP_TOKEN` | *(unset)* | Bearer token for the Streamable HTTP transport (`serve-http`). Unset, the transport is loopback-only — binding a non-loopback address without this set is refused outright. Set, the token gates every request via a constant-time comparison; stdio ignores this entirely. |
| `CODECALC_HTTP_URL` | `http://127.0.0.1:8000` | What the HTTP transport's auth metadata advertises as its own URL. Only consulted when `CODECALC_HTTP_TOKEN` is set; the loopback default matches the offline-by-default posture rather than guessing a public one. |
| `CODECALC_OAUTH_ISSUER` | *(unset)* | Same as `--oauth-issuer`: validate `serve-http` bearer tokens as JWTs against this issuer instead of the static `CODECALC_HTTP_TOKEN`. Off by default. If both end up set, the issuer wins and the static token is rejected — see "Run the server" above. |
| `CODECALC_OAUTH_AUDIENCE` | this server's own resource URL (`CODECALC_HTTP_URL` + `/mcp`) | Same as `--oauth-audience`: the expected JWT `aud` claim, and the RFC 8707 resource this server advertises at its own `/.well-known/oauth-protected-resource`. Only consulted when `CODECALC_OAUTH_ISSUER` is set. |
| `CODECALC_OAUTH_JWKS_URL` | *(unset)* — discovered from `<issuer>/.well-known/openid-configuration` | Same as `--oauth-jwks-url`: pin the JWKS endpoint instead of discovering it. Only consulted when `CODECALC_OAUTH_ISSUER` is set. |
| `CODECALC_OAUTH_SCOPES` | *(unset)* | Same as `--oauth-scopes`: space-separated scopes a token must carry. Unset, any token that otherwise verifies is accepted regardless of scope. Only consulted when `CODECALC_OAUTH_ISSUER` is set. |
| `CODECALC_RUNTIME_PATH` | the server's own `PATH`, else `/usr/local/bin:/usr/bin:/bin` | The `PATH` executed code resolves runtimes on. **Set this when an MCP client spawns the server**: clients often launch with a stripped environment, so an inherited `PATH` can miss a toolchain manager's shims entirely and most languages silently become unavailable. `list_languages` reports what actually resolved. |
| `CODECALC_EXEC_BIN` | `bin/codecalc-exec` (arch-matched) | Override the sandbox binary. Without one, codecalc falls back to a pure-Python executor — `list_languages` and `execute_code` still work, but the Rust path is the production one. |
| `CODECALC_REQUIRE_NATIVE` | *(unset)* | Fail-closed: refuse to start if no usable `codecalc-exec` binary was found (checked at import, so this is also a server-start check), instead of silently answering every call on the weaker Python fallback. Raises naming `CODECALC_REQUIRE_NATIVE` and the paths that were checked. |
| `CODECALC_EXECUTION_PROVIDER` | `local` | Default execution-provider ID. Explicit `execute_code(provider=...)` selection still wins. Setting this to an unregistered provider fails explicitly; it never falls back. |
| `CODECALC_PISTON_URL` | *(unset)* | Register the non-local open-source Piston v2 provider at this absolute HTTP(S) base URL. No public service is contacted by default. |
| `CODECALC_PISTON_AUTHORIZATION` | *(unset)* | Exact value for Piston's `Authorization` header. It is scoped to the Piston transport and redacted from normalized results, descriptors, health, and receipts. |
| `CODECALC_STRICT_URL` | *(unset)* | Activate the current OS's `<host>-strict` provider as an authenticated client of the Linux strict execution service. Without it, strict selection fails closed. The adapter verifies the remote enforcement handshake before sending source. |
| `CODECALC_STRICT_AUTHORIZATION` | *(unset)* | Exact value for the strict service's `Authorization` header. It is never published in descriptors, doctor output, errors, or receipts. |
| `CODECALC_RUN_STATE_DIR` | `~/.codecalc/runs` | Durable metadata-only journal backing `run_submit`/`run_inspect`/`run_cancel`, for every provider (not only managed strict runs). Source, stdin, output, and credentials are never written there. On restart, recorded orphan runs are cancelled and cleaned through their owning provider where it supports that; where it does not (the built-in `local` provider), there is nothing to signal and the record is simply marked recovered. |
| `CODECALC_MAX_ACTIVE_RUNS` | `64` | Admission cap for `run_submit`: how many runs may be running/cancelling at once before further submissions are refused with a `resource_exhausted` error. Bounds the in-memory run table and its thread pool against an unbounded burst or a caller that never inspects/cancels what it starts. An empty, non-numeric or non-positive value falls back to `64` with a message on stderr — a set-but-empty variable is a shell and compose-file commonplace, and it used to abort the server's import. |
| `CODECALC_ALLOW_RUNTIME_APPLY` | *(unset)* | Permit `update_runtimes(apply=True)` to run the **elevated** update commands (apt, via `sudo`). Unset, they are skipped with `ok: false` naming this variable, and the unprivileged managers still run. Deliberately an environment variable rather than a tool argument: `apply` is something a connected model can flip, and this is not. Accepts `1`/`true`/`yes`/`on`; an empty value is not consent. |
| `CODECALC_SESSION_ROOT` | `~/.codecalc/sessions` | Where session workspaces live. **Keep this codecalc-private.** `codecalc cleanup --write --include-unmarked` removes plain, session-shaped subdirectories under it on a heuristic (name shape + age) that is a loose filter, not a strong one — never point it at a directory anything else writes into. |
| `CODECALC_CLEANUP_ABANDONED_AGE_HOURS` | `24` | How old (and untouched) a marker-less, session-shaped directory must be before `codecalc cleanup --include-unmarked` will consider it abandoned. Only consulted with `--include-unmarked`; the default `cleanup` invocation never reads it. |
| `CODECALC_PACKAGE_ALLOWLIST` | *(unset)* | Deny-by-default allowlist for `install_package`. Unset, any syntactically valid package name may be installed (today's behaviour). Set, only listed packages install — anything else is refused before any subprocess or network work, with the stable `permission_denied` code. Comma-separated; each entry is `<language>:<name>` (scoped to one ecosystem) or a bare `<name>` (every ecosystem). Matches the bare name, ignoring `[extras]` and `==version` pins. |
| `CODECALC_SESSION_IDLE_TTL_SECONDS` | *(unset)* | Idle-expiry for stateful (python3/node) session workers: a session untouched for longer than this is reaped — worker killed via the same teardown `session_stop` uses — on its next access. Unset, a session worker lives until `session_stop` or server exit, same as before this existed. A subsequent call on an expired session gets `ok: false` with the stable `worker_failure` code, never a silent respawn. |
| `CODECALC_SESSION_DISK_QUOTA_MB` | `512` | Per-session ceiling on total workspace disk. `session_write_file` and oversized-output spilling refuse BEFORE writing (`resource_exhausted`, no partial file); code run via `execute_code(session_id=...)`/`session_run` is checked before it starts and, since its own writes cannot be pre-checked, again after — an over-quota run still returns its result, now with `disk_quota_exceeded` plus usage/limit, and the session's next write/run is refused until usage (re-measured fresh each time) drops back under the line. Also the cap a SESSIONLESS run's per-run dependency workdir is held to (`codecalc/dependencies.py`, checked after each successful install) — reused rather than a second, independently-tunable constant, since it is the same kind of workspace in every way that matters here. |
| `CODECALC_TOTAL_DISK_QUOTA_MB` | `8192` | Global ceiling on disk summed across every session workspace on this host — closes the gap where staying under the per-session quota by opening many sessions would otherwise be unbounded. Same enforcement points and `resource_exhausted` contract as `CODECALC_SESSION_DISK_QUOTA_MB`. |
| `CODECALC_MAX_ARTIFACT_BYTES` | `16777216` (16 MiB) | Per-write size ceiling for anything a session write path creates — independent of the total quotas above, so one runaway file cannot hide under a generous session/global total. A WRITE-time cap; distinct from `RESOURCE_MAX_BYTES` (4 MiB), which caps what a *read* may serve back. |
| `CODECALC_MAX_ARTIFACT_COUNT` | `500` | Per-session ceiling on the number of artifact files — catches a session writing one byte at a time into thousands of tiny files, a shape no byte-sized cap alone bounds. Only a write that creates a NEW file is checked; overwriting an existing one always succeeds regardless of the count. |
| `CODECALC_MIN_HOST_FREE_MB` | `256` | Refuse a session write when the HOST's free disk space drops below this — protects the host even when every quota above is generous, since a shared host can be driven low by something that is not a codecalc session at all. Measured with `shutil.disk_usage`, which works identically on Windows, unlike `statvfs`. |
| `CODECALC_MAX_SNAPSHOT_BYTES` | `268435456` (256 MiB) | `session_snapshot(action="save")` refuses to archive a workspace whose files sum to more than this — independent of the SESSION disk quotas above, since a snapshot is written OUTSIDE any session's own workspace and quota. |
| `CODECALC_MAX_SNAPSHOTS_PER_SESSION` | `10` | Per-session ceiling on the number of snapshots kept at once — catches many small snapshots the byte cap alone would not, the same "count cap alongside the byte cap" shape `CODECALC_MAX_ARTIFACT_COUNT` already applies to workspace files. |
| `CODECALC_CAPABILITY_POLICY` | *(unset)* | Capability broker. Unset, no brokering — a job's capabilities run as requested (today's behaviour); the execution receipt still discloses them under `provider.capabilities` with `brokered: false`. Set, comma-separated directives narrow them: `deny-network` forces `no_net` on a job that did not request network (enforced where the provider can, disclosed as `effective` where it cannot); `allow-network` explicitly grants network to a job that requested it; `strict` rejects a job whose denial the provider cannot enforce. The broker never approves a capability the request did not ask for — an escalation is refused with `permission_denied` / `capability_not_requested`, before any side effect. |
| `CODECALC_AUDIT_LOG` | `~/.codecalc/audit/audit.log` | Append-only JSON-lines audit stream for broker decisions and security-relevant side effects (denied capability, refused install, cleanup). Each event carries a source-safe timestamp, the run/session id, the decision and reason, and never the executed source or a credential. Set to a path to relocate it; set empty to disable. Best effort — a write failure never fails a run. |
| `CODECALC_PROCESS_HEADROOM` | `512` | Fork-bomb guard. `RLIMIT_NPROC` is a **uid-wide task budget**, not a per-sandbox one — the kernel compares it against every thread your user owns, machine-wide. So codecalc measures the ambient count per execution and sets the limit to *ambient + headroom*: a bomb can add at most this many tasks, while a runtime wanting a few threads always has room however busy the box is. |
| `CODECALC_MAX_PROCESSES` | *(unset)* | Escape hatch: pin `RLIMIT_NPROC` to an absolute value and skip the measurement. |
The strict service runs on Linux x86_64 or ARM64 with Docker Engine, cgroup v2,
and an explicitly registered gVisor `runsc` runtime. Its executor image must be
pinned by `@sha256:` digest on the execution path. That image is published to
GHCR (`ghcr.io/the-40-thieves/codecalc-exec`, multi-arch amd64+arm64) by the
`publish-executor-image` workflow, which an operator dispatches
(`workflow_dispatch`); the workflow commits the immutable digest into
`docker/executor-image.lock`, and `published_strict_image()` resolves it as the
production default. Until that first dispatch no digest is pinned and the
execution path **fails closed** — it never falls back to the mutable local
diagnostic tag (`codecalc-exec:strict`), which `doctor` and the conformance
suite keep using. The default `systrap` platform works without KVM, so the same
authenticated service can be used from Linux, macOS, and Windows; strict clients
never fall back to native local execution.
Provisioning and running any of the three strict backends in production —
the gVisor+Docker host, Windows AppContainer hardening, and the macOS/Windows
remote-client configuration — is covered in
[`docs/deployment/README.md`](docs/deployment/README.md), separate from the
provider interface itself in
[`docs/contract/provider-v1.md`](docs/contract/provider-v1.md).
Both backends resolve `CODECALC_RUNTIME_PATH` identically, and
`scripts/check_parity.py` fails CI if the Rust and Python copies of that
contract ever drift — including if a machine-specific home directory finds its
way back into the default.
## Tool-definition token cost
codecalc's `tools/list` returns 49 definitions. Measured with `o200k_base` as a
proxy on the served JSON, that is 78,586 bytes / 20,630 tokens of descriptions
and input schemas (up from 64,643 bytes / 17,277 tokens at the same 49 tools),
and every client pays it before the first user message. The number
has grown with the descriptions, not the count: the disambiguation sentences
and the per-mode text on `bits`/`symbolic` are what a selection-accuracy-first
server spends its tokens on. The latest jump (+13,943 bytes, +3,353 tokens) is
every one of the 152 tool PARAMETERS gaining its own `description` in the
input schema (`Annotated[<type>, Field(description=...)]`) — the docstrings
above did not change, so `scripts/tool_select_eval.py`'s selection-accuracy
numbers (it scores only name + docstring, never the input schema) are
unaffected by this change.
A follow-up pass then edited the DOCSTRINGS themselves, now that every
parameter's own syntax/default/range lives in its schema description and no
longer needs restating in prose: measured with the same `o200k_base` proxy
(a `mcp.Client.list_tools()` dump, `by_alias=True, exclude_none=True`, one
tool per JSON object), **77,237 bytes / 20,439 tokens** — down from the
152-parameter figure above (**-1,576 bytes, -412 tokens**), short of the
~19,000-token target this pass aimed for. What moved which way: the
execute_code/execute_code_stream/run_submit/session_run cluster and six
other tools (`truth_table`, `session_files`, `algebraic_equiv`,
`compare_threshold`, `percentage`, `percentiles`) each gained one
disambiguation/usage sentence naming a sibling tool, `verify_optimization`'s
docstring was cut to about 60% of its length, and pure schema-restating
sentences ("`languages` = comma-separated subset", explicit default/range
call-outs, repeated operation lists) were removed across the file — but
several of THOSE removals had to be partly reverted once
`scripts/tool_select_eval.py` showed they deleted discriminating
vocabulary BM25 actually leans on (see `scripts/data/tool_select_baseline.json`,
regenerated by this pass: `full` 151→150/234, `dev` 128→130/184, `core`
79→78/116 top-1 hits, every surface within the eval's own
`DEFAULT_EPSILON_HITS`), which is most of why the net reduction is smaller
than the additions alone would suggest.
A per-tool `icons` field (2025-11-25+) was tried and measured, not assumed:
one tiny inline `data:image/svg+xml;base64,...` glyph per tool GROUP, under
300 bytes even for the largest of six — small per icon, but `Tool.icons` is a
per-TOOL field, so each of the tools repeats its group's full base64
payload on the wire, and base64 tokenizes far worse than prose under a BPE
encoder. Measured on the full served `tools/list` payload: **+11,540 bytes,
+6,665 tokens** (`o200k_base`) — real and non-trivial on a server whose whole
pitch (see "Reducing the tool surface" below and `docs/design/
2026-08-10-tool-facade.md`) is that tool SELECTION accuracy matters more than
saving a few tokens elsewhere. Removed. codecalc's `MCPServer` still carries
one SERVER-level icon plus a `website_url` — both ride on `initialize`, once
per **connection**, not once per **tool**, so they do not touch `tools/list`
at all: measured before/after, the served `tools/list` payload is
byte-identical (59,902 bytes / 15,952 tokens either way, as measured at the
time on the then 52-tool surface) — **+0** on the
number this section exists to track.
codecalc does not hide its tools behind a discovery facade, and that is
deliberate: the tool surface is where per-operation approval prompts, audit
names and typed schemas live, and collapsing 49 tools into one dispatcher makes
`install_package` and `percentage` look like the same permission to a client
that approves by tool name. The cost is real, but the client is the better place
to solve it, because the client can defer definitions **without** giving up the
schemas or the per-tool boundary.
If you are paying too much for codecalc's definitions:
- **Claude Code** defers every MCP tool by default — tool search is on by
default, with no token floor codecalc needs to clear. `auto` loads a
server's tools upfront only while their definitions total under 10% of the
context window and defers all of them once that 10% is reached; `false`
loads everything upfront regardless of size (Claude Code MCP docs,
https://code.claude.com/docs/en/mcp, retrieved 2026-09-07). `calc_exact`,
`execute_code`, `verify_translation`, `verify_optimization`, and
`list_languages` carry `_meta["anthropic/alwaysLoad"]` (per that same doc,
"your 3-5 most frequently used tools") so they stay loaded even when a
client defers everything else; `install_package` and `update_runtimes`
carry `_meta["anthropic/requiresUserInteraction"]`, which forces a
permission prompt on every call regardless of the session's permission
mode — both change the host and both fetch from a registry.
`execute_code`, `execute_code_stream`, `session_run`, `compare_execution`,
and `run_inspect` carry `_meta["anthropic/maxResultSizeChars"] = 499520`
(`2 * 240 KiB + 8_000`), the truncation hint for the one tool family whose
output can legitimately approach it — `run_inspect` carries it because
its terminal reply, once a `run_submit`-started run finishes, is the same
envelope `execute_code` returns. This bounds the serialized TEXT `content`
block only — the JSON result as the string a client renders as the tool's
reply — not the whole MCP response: every one of these five tools except
`session_run` also declares `outputSchema`, so the SDK additionally
attaches `structuredContent` with the same JSON ("MCP server developers
can configure custom output limits for individual tools by specifying
`_meta['anthropic/maxResultSizeChars']` in the tool's listing, up to a
hard maximum of 500,000 characters", same doc as above, describes the
text result specifically) — so a typed tool's total wire payload
approaches twice this hint. `session_run`'s inlined artifact content
blocks (image/text/link, up to 8 within a 4 MiB encoded budget — see its
own docstring) are likewise separate blocks outside this bound. 240 KiB
per stream is a hard CEILING `max_output_kb` is clamped to on every tool
that accepts it (`execute_code`, `execute_code_stream`, `run_submit`),
separate from the 64 KiB DEFAULT `0` selects — chosen as the largest
round-KiB ceiling that keeps the TEXT-block hint under that 500,000-char
limit; there is no other ceiling on that parameter today, and raising it
further would push the hint over that limit. A run whose real output
needs more than 240 KiB belongs in a session instead: leaving
`max_output_kb` at its default with `session_id` set spills oversized
output to a full-fidelity file, readable in full via `session_read_file`,
rather than truncating it. `compare_execution` takes no `max_output_kb`
of its own, but accepts an unbounded number of `snippets`, so a
many-language comparison can still legitimately exceed the hint.
- **Claude API, via the MCP connector**, takes `defer_loading` once on the
toolset's `default_config`, or per tool in `configs`. Deferred definitions stay
out of the system-prompt prefix, prompt caching is preserved, and a matching
tool is expanded into its full definition when the model searches for it.
- **OpenAI's Responses API** has the same knob under a different name:
`defer_loading: true` on an MCP server tool definition (OpenAI Responses MCP
tool guide, https://developers.openai.com/api/docs/guides/tools-connectors-mcp,
retrieved 2026-09-07).
- **VS Code** caps a single chat request at 128 enabled tools and groups
excess tools behind "virtual tools" above a configurable threshold (VS Code
agent tools docs, dated 2026-09-02,
https://code.visualstudio.com/docs/copilot/agents/agent-tools). **Windsurf /
Cascade** caps at 100 total tools (Cascade MCP docs,
https://docs.devin.ai/desktop/cascade/mcp, retrieved 2026-09-07).
- **The MCP specification itself has no deferral mechanism** — no tool search,
grouping, tags, or toolsets; a server can only publish `ttlMs`/`cacheScope`
hints and paginate `tools/list` (MCP spec 2026-07-28,
https://modelcontextprotocol.io/specification/2026-07-28/server/tools,
retrieved 2026-09-07). A client without one of the mechanisms above pays the
full cost regardless of what codecalc does.
- **Any client** can filter which of the 49 tools it exposes to the model.
Nothing here requires codecalc to change.
A server-side facade remains under consideration for clients with no such
mechanism (`docs/design/2026-08-10-tool-facade.md`), and is not implemented.
Trimming a description to cut this cost is exactly the change
`scripts/tool_select_eval.py` exists to gate: an offline, labeled eval of
whether a deterministic lexical (BM25) selector still picks the right tool
for a plain-language ask, scored against the live `tools/list` text.
Measured v1 baseline (196 hand-labeled prompts, none containing their own
target tool's name — see the script's own docstring): **60.71%** top-1 /
75.51% top-3 accuracy on the `full` surface (62.75% / 63.0% top-1 on `dev` /
`core` respectively). It is a lexical proxy, not a model — see the script's
module docstring for exactly what a green run does and does not prove.
The checked-in baseline PINS the exact labeled corpus by content hash
(`prompt_set_sha256`); a `--baseline` compare against a corpus that no
longer hashes to it fails with a distinct "corpus changed" error rather than
silently scoring a smaller, easier prompt set against the old numbers. And
because a tool can be top-1-wrong against `full`'s 51 distractors (zero
headroom to lose) while still having real headroom against `core`'s much
smaller distractor set, both the regression compare and the ablation
self-check (replacing real descriptions with a generic stub, one tool at a
time, across every candidate tool — no sampling) run separately against all
three of `full`/`dev`/`core`, wired into CI via
`tests/test_tool_select_eval.py` so the gate is proven live, on every
surface, on every run — not just at the PR that added it.
BM25 is a lexical proxy, not a model — `scripts/tool_select_llm_eval.py` is
the model-driven half, calling a real chat model over a live gateway with
the identical tool catalog and labeled corpus; it is opt-in (`workflow_dispatch`,
advisory rather than a hard gate) rather than wired into every PR, and its
measured numbers live in `docs/tool-selection-eval.md` next to BM25's own.
## Reducing the tool surface
For an operator who would rather not configure every client, codecalc also has
a first-party knob: `CODECALC_TOOLS` registers only a chosen slice of the
56-tool surface, so a client that never enables tool search still pays for a
smaller `tools/list`.
On a client with no deferral mechanism of its own, the client's own allow-list
does the same job from the other end — OpenAI's `allowed_tools`, Gemini CLI's
`includeTools`/`excludeTools`, or Codex CLI's `enabled_tools`/`disabled_tools`
all narrow what a given session sees without touching the server.
Every tool also now carries a `ToolAnnotations` hint (`readOnlyHint`,
`destructiveHint`, `idempotentHint`, `openWorldHint` — see
`codecalc/server.py`'s `GROUP_ANNOTATIONS`/`TOOL_ANNOTATION_OVERRIDES` tables
for the value on each of the 49). Codex CLI's `writes` approval mode
(v0.144.0+) reads `readOnlyHint` directly: a tool marked `readOnlyHint: true`
skips the approval prompt, everything else still asks. That covers the whole
`calculator` group (19/19 pure) plus the read-only members of the mixed
groups — `list_languages`/`list_execution_providers`/`runtimes_status`/
`branch_reachability` in `execution`, `z3_check`/`algebraic_equiv` in
`verification`,
`session_list`/`session_files`/`session_read_file`/`session_artifacts`/
`run_inspect` in `sessions`, and `analyze_complexity` in `analysis` — without
codecalc doing anything client-specific; the annotation is the same hint
every MCP client reads, `writes` just happens to be the mode that consumes it.
**This is not the facade** the section above declines to build. Every tool a
group activates keeps its own name, its own typed input schema and its own
per-tool approval prompt — a group that is not active simply never registers
its tools with the MCP SDK at all, so they are absent from `tools/list` *and*
rejected by `tools/call`, not merely hidden behind a dispatcher a client could
still invoke by guessing the name.
Every tool belongs to exactly one group:
| Group | Tools |
|---|---|
| `calculator` (19) | `calc_exact`, `compare_threshold`, `percentage`, `calc_stats`, `percentiles`, `collision_probability`, `data_sizes`, `human_duration`, `epoch_time`, `bits`, `radix_convert`, `float_repr`, `symbolic`, `convert_units`, `physical_constants`, `list_units`, `evaluate_expression`, `truth_table`, `matrix` |
| `verification` (5) | `verify_translation`, `verify_optimization`, `algebraic_equiv`, `compare_edge_cases`, `z3_check` |
| `execution` (8) | `list_languages`, `list_execution_providers`, `execute_code`, `execute_code_stream`, `trace_execution`, `branch_reachability`, `compare_execution`, `runtimes_status` |
| `sessions` (12) | `session_start`, `session_stop`, `session_list`, `session_files`, `session_write_file`, `session_read_file`, `session_run`, `session_artifacts`, `session_snapshot`, `run_submit`, `run_inspect`, `run_cancel` |
| `analysis` (3) | `analyze_complexity`, `benchmark`, `extract_function` |
| `admin` (2) | `install_package`, `update_runtimes` |
`CODECALC_TOOLS` takes a comma-separated list of group names, preset names, or
both:
| Preset | Expands to |
|---|---|
| `core` | `calculator` |
| `dev` | `calculator`, `execution`, `verification`, `analysis` |
| `full` | every group (the default) |
```bash
CODECALC_TOOLS=calculator # just the calculator (19 tools)
CODECALC_TOOLS=core # same thing, by preset name
CODECALC_TOOLS=calculator,execution # two groups, unioned
CODECALC_TOOLS=dev # a coding-assistant slice (35 tools)
```
Unset or empty registers every group — 49 tools, same as today —
so nothing changes for an operator who does not set this. An unknown group or
preset name is a loud startup failure naming the bad value and every known
group/preset, never a silent fallback to "everything" or "nothing": either
direction would turn a typo into a footgun nobody notices until it matters.
`codecalc doctor` prints the active groups, the full group→tools mapping, and
how many tools this process actually registered, whatever `CODECALC_TOOLS` is
set to.
Client-side deferred loading (the section above) and this env var compose
cleanly: point a client with no deferred-loading mechanism at a
`CODECALC_TOOLS`-restricted process, or use both — a smaller declared surface
still benefits from being deferred.
## Test
Each file is a standalone script that prints one `PASS`/`FAIL` line per
assertion and exits non-zero if any failed — no test runner, no plugins.
```bash
cd /path/to/codecalc
# everything. `|| break` used to be `|| break` alone, which stopped at the
# first failure AND left the loop exiting 0 — a red suite reported success to
# anything wrapping this command. This form runs them all and carries the
# failure out.
fail=0
for f in tests/test_*.py; do PYTHONPATH=. .venv/bin/python "$f" || { echo "FAILED: $f"; fail=1; }; done
for f in scripts/*.py; do PYTHONPATH=. .venv/bin/python "$f" || { echo "FAILED: $f"; fail=1; }; done
[ "$fail" -eq 0 ] # the exit status of the whole run
# or individually
PYTHONPATH=. .venv/bin/python tests/test_smoke.py # every language, via the Rust executor
PYTHONPATH=. .venv/bin/python tests/test_mcp_all.py # every tool over MCP stdio, answers checked
PYTHONPATH=. .venv/bin/python tests/test_executor_sweep.py # sandbox regressions
```
71 test files and 19 CI-invoked scripts, **2184 assertions**. "CI-invoked"
means referenced by path (`scripts/<name>.py`) from a job in
`.github/workflows/*.yml` — `scripts/check_claims.py` derives the count that
way and gates it, so a script wired into a workflow without this sentence
changing, or this sentence bumped without a workflow change, fails the build.
Nothing in the suite
needs the internet, so none of it is ever skipped for lack of a network.
It **can** skip for lack of a *capability*, and that is correct rather than a
regression: a machine without a symlink privilege, without a given language
runtime, or without a built native executor cannot exercise the cases that
need them. The suite reports three distinct outcomes — the property holds, the
property is broken, and this machine cannot exercise it — and every skip names
its real cause. A nonzero skip count on Windows or in fallback mode is the
healthy result; what would be wrong is a skip reading as a pass.
This paragraph previously claimed **zero skips** unconditionally. That became
false the moment the suite learned to distinguish the third outcome, and
nothing gated it: `check_claims.py` gates the counts below, not the prose
around them. The counts are gated by
`scripts/check_claims.py`: they were written by hand once and were stale within
three pull requests, which is exactly the failure the rest of that script
exists to prevent. Four of the files are regression suites named after the
sweep that produced them — `test_bug_sweep`, `test_executor_sweep`,
`test_python_sweep`, `test_network_modules` — and each one's docstring states
the defect it locks out and how it was reproduced, because a regression test
whose reason has been forgotten is the first one deleted.
Two rules the suite holds itself to, learned from breaking both:
- **Assert the value, not the shape.** Three of these files once had no
assertions at all: they called tools, printed the output and exited 0. They
caught a crash and never a wrong answer — a `runtimes_status` total replaced
with `-999` passed, printing `total = -999`.
- **Don't pin what varies.** `benchmark` and `compare_execution` rank by
measured time, so their winner moves under load; their structure is asserted
and their timing is not. `runtimes_status` is checked against itself — the
summary must agree with the data it summarises — so it holds on any machine
rather than describing this one.
## Platform support
Linux, macOS and Windows. The three do not offer the same primitives, and the
executor reports which ones it could **not** apply in an `unenforced` array on
every result rather than letting a caller assume they all held.
The native table below describes the `local` provider and is **not a hostile-code
security boundary**. On macOS, `<host>-strict` instead uses the explicitly
configured Linux strict service: the macOS binary performs provider selection,
attestation, supervision, and result validation, while untrusted code executes
inside the remote cgroup/namespace/seccomp/Landlock boundary. A missing or
incomplete service fails before source leaves the Mac and never falls back to
native execution.
Symbolic evaluation carries the same idea. Every symbolic tool runs SymPy in
a forked child under CPU and memory ceilings with a wall clock the parent
enforces, so an expression nobody anticipated is still bounded — SymPy's own
maintainers abandoned their attempt at a `safe=` flag as "security theater", so
the screen in `safe_expr.py` buys time and the child buys the bound. Where
there is no `fork`, the result reports `expression_bound_not_enforced_without_fork`
rather than implying a guarantee.
A second field, `output_error`, covers the other way a result can be wrong:
absent means `stdout`/`stderr` are what the program produced, present means at
least one of them is **not**, and names which stream and the OS error. That
distinction did not exist until [#80](https://github.com/The-40-Thieves/codecalc/issues/80)
— an output file that could not be read came back as a program that printed
nothing, on a run reported as successful. `ok` now accounts for it on both
backends.
| Guarantee | Linux | macOS | Windows |
|---|---|---|---|
| Wall-clock timeout | yes | yes | yes |
| Kill the whole process tree | `killpg` + `PDEATHSIG` | `killpg` | `TerminateJobObject` |
| Fork-bomb guard | `RLIMIT_NPROC` (uid-wide) | `RLIMIT_NPROC` (uid-wide) | Job `ActiveProcessLimit`, **reported unverified**⁵ |
| Memory ceiling | `RLIMIT_AS` | reported unenforced¹ | Job `ProcessMemoryLimit` |
| CPU-time ceiling | `RLIMIT_CPU` | `RLIMIT_CPU` | Job `PerProcessUserTimeLimit`⁴ |
| Open-file ceiling | `RLIMIT_NOFILE` | `RLIMIT_NOFILE` | reported unenforced |
| Output cap | yes | yes | yes (on read) |
| `no_net` | seccomp-bpf filter⁶ (falls back to `LD_PRELOAD` shim²) | `DYLD_INSERT_LIBRARIES`²˒³ | reported unenforced |
| Stateful sessions | yes | yes | yes |
¹ Darwin accepts `setrlimit(RLIMIT_AS)` but does not enforce address space the
way Linux does, so setting it would buy an illusion.
² Dynamically-linked programs only — a statically linked binary (Go, by default)
ignores it.
⁴ Applied via `JOB_OBJECT_LIMIT_PROCESS_TIME`, which Windows has supported
since XP — this was reported as `cpu_limit_unavailable_on_windows` until
2026-08-08, and the table said the same, so code and docs agreed with each other
and disagreed with Windows. It is not identical to `RLIMIT_CPU` and the
difference is reported rather than glossed: it counts **user-mode time only**, so
a process burning kernel time is not capped by it, and the system checks
periodically rather than immediately. Runs on Windows carry
`cpu_limit_counts_user_time_only_on_windows` in `unenforced` to say so.
³ Weaker still on macOS, in two ways. SIP and the hardened runtime strip
`DYLD_INSERT_LIBRARIES` for protected and hardened-signed binaries (most signed
interpreters), and dyld interposing does not reach calls made *inside* the
shared cache where libSystem lives — a program's own `connect()` is intercepted,
a system framework opening a connection internally is not. Treat macOS `no_net`
as a speed bump, never as isolation.
⁶ Linux only. The executor installs a seccomp-bpf filter in the sandboxed
child that refuses the `socket(AF_INET/AF_INET6)` SYSCALL in-kernel — not a
libc symbol, so `ctypes`/`dlsym` and raw `syscall()` calls cannot route around
it the way they can around the `LD_PRELOAD` shim. `AF_UNIX` still works. Falls
back to the shim (with its symbol-level bypass, disclosed in `unenforced` as
`no_net_best_effort_shim`) when the kernel refuses the filter.
Both are exercised by the suite on every platform. The fork-bomb probe measures
the EAGAIN boundary precisely but needs `os.fork`, so it is POSIX-only; a second
probe SPAWNS processes instead, which is the portable operation, and pins the
ceiling low through `CODECALC_MAX_PROCESSES` so it costs two dozen short-lived
processes rather than walking up to the fallback. Verified to track the limit
rather than something incidental: a headroom of 24 bounds it at 22 children and
a headroom of 300 bounds it at 298.
⁵ Windows' `ActiveProcessLimit` is scoped to the **job** rather than to the uid,
so it avoids the failure mode that broke 14 of 31 runtimes on Linux. CodeCalc
now supplies that job at process creation, makes it non-nestable with the
minimal `JOB_OBJECT_UILIMIT_EXITWINDOWS` restriction, and allowlists only the
three standard I/O handles inherited by the child.
Measured on Windows 11 Pro: **400 of 400 spawns succeeded against a ceiling of
24**, reproduced from two unrelated launchers including Task Scheduler. This is
not a failed API call — `SetInformationJobObject` and `AssignProcessToJobObject`
both return success and the correct limit reaches the job. It is topology.
`ActiveProcessLimit` is **not** one of the limits combined across a nested job
chain; those take the most restrictive value, while this one comes from the
process's **immediate** job. A post-creation `AssignProcessToJobObject` places
the child somewhere in that chain rather than at its end: measured, the child's
immediate job reported `0x3000 / APL 0` while codecalc's reported
`0x230A / APL 24`, so codecalc's ceiling was never consulted.
No parent-side Win32 call returns another process's immediate job or its
effective `ActiveProcessLimit`, so this cannot be closed by inspection. Every
compatibility run that assigns the child after creation therefore carries
`process_limit_enforcement_unverified_on_windows` in `unenforced`. Four further
strings can each positively *prove* a failure; none can prove success, so their
silence does not imply enforcement.
Creation-time assignment is the default. It was verified on Windows 11 Pro with
a direct Python runtime: 23 children succeeded against a total limit of 24 and
the next spawn failed with WinError 1816. Runtime launchers that require an
inner job now fail rather than silently escaping the limit; configure a direct
runtime executable. `CODECALC_WIN_JOB_AT_CREATION=0` retains the old path only
as an explicitly unverified compatibility escape hatch.
**AppContainer security isolation is a DIFFERENT guarantee from the Job Object's
resource limits.** The Job Object above caps *resources* — memory, process count,
user-mode CPU — and each run names in `unenforced` which of those did not bind.
The optional AppContainer backend adds a *security* boundary layered on the same
creation-time topology: a least-privilege AppContainer profile
(`CreateAppContainerProfile`, no capability SIDs, so no network), launched with
`SECURITY_CAPABILITIES` in the same `STARTUPINFOEX` attribute list as the job
assignment. Access is granted two ways, deliberately split. The sandbox
**workdir** is granted to the run's *own* AppContainer SID — per-run, so
concurrent runs cannot reach each other's workdirs, and it vanishes with the
ephemeral directory. The **interpreter directory** is granted read+execute to the
fixed *ALL APPLICATION PACKAGES* SID (`S-1-15-2-1`) as an explicit,
non-inheritable ACE applied per file across the tree — because a real
interpreter's pre-existing files are inheritance-protected and no inheritable
grant reaches them. That interpreter grant is **persistent and cached** (a marker
in codecalc's own state dir; the several-thousand-file walk runs once per
interpreter): a deliberate trade-off that leaves a read-only ACE, readable by any
AppContainer on the machine, on a public interpreter — rather than re-walking
every run. The intended property is that a payload cannot read the user profile,
write outside its workdir, or reach the network. It is **OFF by default** (opt in
with `CODECALC_WIN_APPCONTAINER=1`) and **fails closed** — if profile creation,
SID derivation or an ACL grant fails, the launch is refused rather than dropped
to an unconfined process. The isolation has been verified on a Windows 11 box
(AppContainer SID present, user-profile secrets unreadable, writes confined to
the workdir, network denied, ambient privileges reduced to the two benign ones
Windows keeps), yet every run that takes this path still emits
`appcontainer_isolation_unverified_on_windows`: a Server-SKU CI runner cannot
exhibit AppContainer behaviour, and the guarantee ultimately depends on the
deployment's OS and configuration, so the shipped default stays conservatively
disclosed rather than claiming a universal proof.
Two things degrade rather than fail on a given platform: languages whose runtime
is absent (`list_languages` reports `available: false`), and the shell-wrapped
plans — `gleam` and `haskell` — which need a POSIX shell to scaffold a project
and report `available: false` on Windows outright rather than resolving through
a bash that cannot run them. `csharp` left that set: .NET 10 runs a single `.cs`
file directly, so it is shell-free on every platform.
### Reliability tiers
`available`/`status` above is a claim about **resolution**: did this machine
find the command on PATH. It is not a claim about **reliability**: has
codecalc's own CI ever actually run this language and checked the output. The
two are orthogonal, and they disagree in practice — a review's own smoke test
found the rust and csharp *host toolchains* failing on a machine where both
`rustc` and `dotnet` resolved cleanly. `list_languages`, `runtimes_status`, and
`codecalc doctor` all report a `tier` alongside resolution to make that gap
visible instead of silent:
| Tier | Meaning |
|---|---|
| `tested` | A CI job genuinely **executes** this language and asserts on its real output, on every PR. Currently `python3`, `node`, `rust`, and `go` — kept deliberately conservative, and gated by `scripts/check_runtime_tiers.py` so a language cannot claim it without a CI check backing it, or silently drop out of CI while still claiming it. `python3`/`node` earn it from the stateful-worker sweep (all three OS legs); `rust`/`go` from `tests/test_tier_evidence.py`, which compiles and runs a real program in each and asserts a per-run computed stdout — on the **Linux** leg, where skips are promoted to failures. The tier claims "CI executes this on every PR", not per-platform coverage. |
| `best_effort` | Declared, with a local smoke fixture (`tests/test_smoke.py`), and plausibly works on a normal install with the right toolchain — but no CI job runs it, so nothing would notice it silently breaking. Every other language, including `csharp`, `java`, and the rest. |
| `plan_only` | A registry entry never validated on any runner, anywhere, not even locally. None today. |
`codecalc doctor`'s text output prints both axes side by side rather than
folding tier into the resolution summary, so a `best_effort` runtime that
happens to be `installed` on your machine reads as exactly what it is:
resolved, unverified by codecalc, may be broken.
## Sandbox guarantees
- Fresh temp dir per run, deleted on exit (source + binaries + outputs). The
deletion is **identity-checked**: the directory's device and inode are
recorded at creation and re-checked before removal, because executed code runs
with that directory as its cwd and can rename another one into its place. A
caller-supplied `--workdir` is a session workspace and is never deleted.
If the filesystem supplies no file index to identify the directory by, the
deletion is **refused** rather than performed unverified, so temp directories
accumulate there instead of the wrong one being removed. That trade is stated
because it is the one this guarantee actually makes: it was previously
implemented in the Rust executor only, and the Python fallback deleted
unconditionally, which CI caught on Windows.
- rlimits: CPU (timeout+8s), address space 2TiB (V8/JVM need huge VA),
file size 256MiB, 256 FDs, core dumps off
- The **timeout is a total budget**: compile and run share it, so `--timeout 10`
cannot take twenty seconds. `duration_ms` is the run alone; `compile_ms` and
`total_ms` are reported separately.
- Wall-clock timeout kills the whole process **group** (SIGKILL). So does
SIGTERM to the executor — `PR_SET_PDEATHSIG` reaches only the direct child, so
a group kill is what covers its descendants, and the executor is the only
participant that knows the group id.
- Output capped at 64KiB per stream, on every path including stateful sessions.
Exceeding it is reported as `OLE`, and the file-size rlimit is kept strictly
above the cap so that overflow stays *detectable* — tying the two together
turned a truncated 4MB output into a silent `verdict: OK`.
- Fork-bomb guard via `RLIMIT_NPROC`, sized from the **measured** ambient task
count plus headroom rather than a fixed number. This is a mitigation, not
isolation: the budget is shared with every other process your user owns, so
concurrent executions draw on the same pool. cgroup v2 `pids.max` is the real
per-sandbox answer and needs delegated cgroup access a stdio MCP server cannot
assume — reach for it when this moves behind a container.
- `no_net` blocks the **network**, not every socket: it refuses `AF_INET` and
`AF_INET6` and forwards everything else, so `AF_UNIX` local IPC keeps working.
- No network namespace isolation (single-host tool; containerize for untrusted code).
**Note, 2026-09-07:** this bullet describes the pre-#242 state. Since
[#242](https://github.com/The-40-Thieves/codecalc/pull/242) (2026-08-21),
Linux additionally enforces `no_net` in-kernel via a seccomp-bpf filter —
see the `no_net` row in `SECURITY.md`'s "Known limitations" table for the
current per-platform breakdown. A full network *namespace* is still only the
strict (gVisor) backend's job; this note does not change that.
- Every result carries a `backend` field (`"rust"` or `"python"`) so a caller
never has to infer which sandbox actually ran from an absent key — that was
possible to confuse with an older build that never reported it at all. The
pure-Python fallback cannot provide everything above: it has no `no_net`
shim (reported in `unenforced`, not silently dropped), and `peak_memory_kb`
comes back `None` rather than a number, because `ru_maxrss` is a
process-lifetime high-water mark this path has no way to attribute to one
run. `CODECALC_REQUIRE_NATIVE=1` turns "running on the fallback" into a
startup failure instead of a guarantee you have to notice was quietly
weaker.
### Sessions
A session is a persistent workspace; `python3` and `node` additionally get a
long-lived REPL worker so variables and imports survive between calls. What that
does and does not buy you:
| | workspace session | stateful worker |
|---|---|---|
| Fresh sandboxed process per call | yes | no — one worker serves every call |
| `max_memory_mb` / `max_cpu` / `no_net` | applied | **reported in `unenforced`** |
| `RLIMIT_AS` / `NPROC` / `FSIZE` / `NOFILE` | per call | applied once, at worker start |
| Output cap + `OLE` | yes | yes |
| Per-call wall clock | yes | yes — a worker that blows it is killed |
A worker cannot take a per-call rlimit after the fact, and `--no-net` is decided
at exec time. Rather than accept those arguments and drop them, the result lists
them in `unenforced` — the same field the executor already uses to say "asked
for, not applied". Omit `session_id`, or use a workspace session, when a ceiling
has to be real.
The worker protocol does not share a file descriptor with executed code, and
every response carries the id of the request it answers. Both matter: `sys.stdout`
is a Python-level rebind that a subprocess writes straight past, and a corrupted
stream that is not resynchronised returns every later call the *previous* call's
result — a well-formed answer to a different question.
The channel differs by platform and the guarantee does not. POSIX hands the
worker an out-of-band pipe; Windows has neither `pass_fds` nor `preexec_fn`, so
the worker appends responses to a file whose path arrives in the environment.
Either way a child spawned with inherited stdio writes to fd 1 and cannot reach
the protocol. Tests force the file-backed channel on every platform, because an
unexercised fallback is one that works until it is needed.
### Local operations: status & cleanup
Two CLI-only commands for an operator running a long-lived server, not MCP
tools — they don't count toward the tool surface above:
```bash
codecalc status # read-only snapshot: sessions, disk usage, quotas, audit log
codecalc status --json # the same report, for scripts
codecalc cleanup # DRY RUN (the default) — marker-based only
codecalc cleanup --write # actually removes marker-based candidates
codecalc cleanup --write --include-unmarked # ALSO sweep old, unmarked, session-shaped dirs
```
`status` reports `SESSION_ROOT`, how many sessions exist and which of them
are idle-expired (the on-disk `.codecalc-session-expired` marker the
idle-TTL reaping leaves behind), per-session and global workspace disk
usage, the configured disk quotas and current headroom, the audit log's
path and size, and a one-line runtime reliability-tier summary. It changes
nothing — no session is started, stopped, reaped, or written to.
`cleanup` reclaims disk from session directories under `SESSION_ROOT`.
`--dry-run` is the default — nothing is removed until `--write` is passed.
Because `cleanup` runs as a SEPARATE process from any server that may be
using `SESSION_ROOT` right now, it has none of that server's in-memory
bookkeeping to consult — only what is on disk.
**Directory mtime is deliberately NOT trusted as a liveness signal.** An
earlier version of this feature did trust it, and an adversarial review
proved that wrong live: a REPL worker doing purely in-memory work touches
no file at all, and even an in-place file overwrite bumps only that file's
own mtime, never its parent directory's — so a genuinely-active worker
session can look, by directory mtime alone, identical to an abandoned one.
The real signal is a per-worker-session **liveness lockfile**: the server
writes its own pid into the session directory the moment a stateful
(python3/node) worker starts, and removes it the moment that worker is
actually gone (reaped or `session_stop`). `cleanup` checks this for every
candidate and refuses outright — regardless of marker, age, or the mtime
floor below — whenever the lockfile names a pid that is still alive. That
is what makes **a session any running codecalc server is using is never
deleted** true for worker sessions.
By default, `cleanup` considers ONLY directories carrying the idle-expiry
marker — the risk-free path, since a marker only ever exists after
sessions.py's own idle-TTL reaper has already closed that specific worker
for good (session ids are never reused). `--include-unmarked` additionally
sweeps old (`CODECALC_CLEANUP_ABANDONED_AGE_HOURS`, default 24h),
session-shaped, marker-less directories — the one path with real residual
risk, because a **workspace-only** session (no worker) never gets a
lockfile to check against, so this path falls back to age + a hard
recency floor (nothing modified in the last few minutes is ever touched)
as a heuristic, not a proof. Turn it on deliberately, and never point
`CODECALC_SESSION_ROOT` at anything but a codecalc-private directory —
the "looks like a session dir" name filter is loose, not strict.
Other safety properties, unconditional on every path: only a direct child
of `SESSION_ROOT` is ever a candidate (never `SESSION_ROOT` itself); a
symlink there is refused, never followed; and removal itself is
identity-checked (device/inode, re-verified immediately before the
delete) the same way `session_stop`'s own workspace teardown is, so a
directory swapped out from under a stale scan is refused rather than
deleted.
## Language list
python3, node, bun, deno, typescript, ruby, php, perl, lua, tcl, r, elixir,
erlang, bash, zsh, mojo, swift, c, cpp/c++, rust, go, fortran, zig, java,
kotlin, csharp, gleam, haskell, sqlite, jq, awk — 31 runtimes.
codecalc does not install any of them. It runs whatever is already on
`CODECALC_RUNTIME_PATH`, and `list_languages` probes each one and reports which
actually resolved, so a minimal machine degrades to the subset it has rather
than failing opaquely.
## Notes
- Java uses single-file source launch (JEP 330). Kotlin compiles to a jar.
- gleam/haskell scaffold a temp project (gleam new / nix-shell); csharp runs the file directly (.NET 10 file-based apps).
- `benchmark` uses the stdin-N contract: code reads N from stdin, work sized by N.
## CI
Five workflows, each documented inline with what it gates and — where a tool was
considered and rejected — why it is not there.
| Workflow | Gates |
|---|---|
| `ci-rust` | clippy `-D warnings`; the executor's **JSON contract**, asserted by running the built binary (OK/TLE/OLE/unknown-language) and confirming a canary secret in the executor's own env does not reach executed code; both **static musl** cross-builds, checked with `file` for static linkage; `blocknet.so` built `-Werror`, symbol-checked, and confirmed to actually block an outbound connection |
| `ci-python` | `ruff` at a genuine zero residual (ruleset and every exception in `pyproject.toml`, each with a reason); calc parity on 3.11 and 3.14; the security suite against the **Rust** backend, with an assertion that the Rust backend is the one under test; MCP stdio round-trip |
| `ci-security` | `scripts/check_no_eval.py` (the CRITICAL-01 invariant), `scripts/check_parity.py` (the three security constants duplicated in Rust and Python must match), `scripts/check_claims.py` (README counts and licence), `actionlint`, `gitleaks`, `trufflehog`, `osv-scanner`, `cargo-deny`, `cargo-audit`, and `opengrep` on a schedule |
| `ci-quality` | `typos`. **Not** `shellcheck` — the repo's last shell script was removed with `executor/zig-cc.sh`, so the gate would have matched zero files and reported success for scanning nothing; `actionlint` in `ci-security` shellchecks every embedded `run:` block instead. The workflow says so inline. |
| `dco` | `Signed-off-by` on every non-merge commit |
Two conventions run through all of them, both borrowed from harder-won experience:
- **Actions are pinned by commit SHA and downloaded tools by SHA-256.** A tag is
mutable; a digest is not.
- **Every scan asserts it scanned something.** A linter pointed at a renamed
directory, a dependency scanner with no lockfile to read, and a clean repo all
produce the same output — exit 0. Each gate counts its inputs first and fails
if the count is implausible.
## Licence
Apache-2.0. See [LICENSE](LICENSE).
Contributions require a DCO sign-off (`git commit -s`); `dco.yml` enforces it.
TDQS
Scored across 49 tools
Many tools are clearly distinct (e.g., session_*, run_*, symbolic vs calc_exact), but there is notable overlap among execution tools: execute_code, execute_code_stream, run_submit, session_run, and compare_execution all run code with different modes, and several math tools (calc_exact, evaluate_expression, symbolic, calc_stats, percentiles) have adjacent purposes. The descriptions are detailed enough to disambiguate with careful reading, but the sheer number of similar execution/math tools creates real selection risk.
Most tools follow a clear verb_noun pattern (execute_code, session_start, run_inspect, convert_units, verify_optimization). There are minor deviations: 'bits' and 'symbolic' are noun/adjective-only names, and 'matrix' is a single noun, but these are documented as mode-based consolidated tools. Overall the pattern is consistent and predictable.
49 tools is a very large surface for a code calculation/execution server. While the server covers many domains (execution, sessions, math, units, verification, runtimes), the count is heavy and includes several consolidated mode-based tools that could reduce the count further. It exceeds the typical well-scoped range and will burden agent tool selection.
The server covers its apparent domains thoroughly: code execution (sync, stream, background, session), file/session management, math/units/constants, verification (optimization, translation, edge cases), and runtime management. Minor gaps exist (e.g., no explicit session_run cancellation, no direct file deletion tool), but the core workflows are well covered and there are no obvious dead ends.