Carbon Tracking MCP
# Carbon Tracking MCP
MCP servers that estimate the energy used by your Claude Code and Codex sessions, so you can ask for it directly from within a chat: "how much energy has this chat used?" or "how much has this whole project cost?"
They wrap a small pricing-ratio energy model (`src/mcps/energy_estimate.py`) around each tool's local session logs — no telemetry, no network calls. The Claude Code server reads `~/.claude/projects/**/*.jsonl`; the Codex server reads Codex's local thread index (`~/.codex/state_5.sqlite`) and rollout logs under `~/.codex/sessions/`. Both run over the stdio MCP transport (`src/mcps/stdio/`), spawned fresh per chat by Claude Code/Codex.
## What it does
Each server exposes the same two MCP tools, scoped to its own provider:
| Tool | Answers |
|---|---|
| `current_session_energy` | How much energy has *this* chat used so far? |
| `collate_project_sessions_energy` | How much energy has *every* chat in this project used, in total? |
Example output:
```json
{
"session_id": "afc721a3-0d77-4c5f-b1f5-24074d03fa7d",
"file": "/Users/you/.claude/projects/-Users-you-my-project/afc721a3-....jsonl",
"request_count": 60,
"estimated_kwh": 1.018,
"estimated_kg_co2": 0.144,
"comparisons": [
{"id": "washing_machine_cycle", "label": "washing machine cycle", "count": 0.21},
{"id": "ev_car_km", "label": "km driven in a 2020 Tesla Model 3", "count": 1.78}
]
}
```
```json
{
"project_dir": "/Users/you/.claude/projects/-Users-you-my-project",
"session_count": 3,
"total_estimated_kwh": 1.545,
"sessions": [
{"session_id": "afc721a3-...", "request_count": 60, "estimated_kwh": 1.018},
{"session_id": "8ee8cb6f-...", "request_count": 5, "estimated_kwh": 0.073},
{"session_id": "2aabc643-...", "request_count": 38, "estimated_kwh": 0.86}
],
"estimated_kg_co2": 0.218,
"comparisons": [
{"id": "washing_machine_cycle", "label": "washing machine cycle", "count": 0.31}
]
}
```
`estimated_kg_co2` and `comparisons` (both tools' full comparison list is longer than shown above — see `carbon_equivalents.json`) convert the estimated energy into CO2eq using a rough UK grid carbon intensity figure, then express it against everyday activities (washing machine cycles, EV charges, flights, ...). See "CO2eq comparisons" below.
Both tools return a typed, field-described [pydantic](https://docs.pydantic.dev/) model (`src/mcps/schema.py`), so MCP clients get a real JSON schema for the response shape rather than an untyped object. If a session/project can't be found, the tool raises a proper MCP tool error instead of returning a disguised "successful" result.
On the Claude Code server, `current_session_energy` identifies "this chat" via the `CLAUDE_CODE_SESSION_ID` environment variable that Claude Code sets on every process it launches (including the server). Codex doesn't set an equivalent env var for MCP subprocesses it spawns, so the Codex server checks `CODEX_THREAD_ID` opportunistically and otherwise falls back to the most-recently-modified indexed thread whose recorded `cwd` matches the server's — a best-effort heuristic that's right unless you have multiple Codex chats open in the same project directory at once. Either tool's `collate_project_sessions_energy` then sums every other session belonging to the current project — every sibling `.jsonl` for Claude Code, every indexed thread with a matching `cwd` for Codex.
## Install
Requires [uv](https://docs.astral.sh/uv/).
```bash
git clone https://github.com/<you>/carbon-tracking-mcp.git
cd carbon-tracking-mcp
uv sync
```
This installs console-script entry points (`carbon-tracking-stdio-claude`, `carbon-tracking-stdio-codex`) backed by the `mcps` package under `src/`.
Register whichever server(s) you use **globally**, so they're available in every project rather than just this one.
Use `uv run --project` (not `--directory`) to launch them: `--directory` changes the subprocess's working directory to this repo before running, which breaks `collate_project_sessions_energy()`'s cwd-based "which project is this?" scoping for every *other* project you use these servers from. `--project` only points uv at this repo for dependency resolution and leaves the subprocess's cwd alone.
### Claude Code
```bash
claude mcp add --scope user carbon-impact-claude -- uv run --project /path/to/carbon-tracking-mcp carbon-tracking-stdio-claude
```
Restart or start a new session and the tools become available. Verify with:
```bash
claude mcp get carbon-impact-claude
```
### Codex
```bash
codex mcp add carbon-impact-codex -- uv run --project /path/to/carbon-tracking-mcp carbon-tracking-stdio-codex
```
`codex mcp add` writes to `~/.codex/config.toml`, which Codex CLI, the IDE extension, and the desktop app all share — there's no per-project scope to choose, so this is global by default. Restart or start a new session and check `/mcp` inside Codex to verify the server is connected. Alternatively, add the entry by hand:
```toml
[mcp_servers.carbon-impact-codex]
command = "uv"
args = ["run", "--project", "/path/to/carbon-tracking-mcp", "carbon-tracking-stdio-codex"]
```
Codex won't reliably call the tools on its own unless you also tell it to: append this repo's [AGENTS.md](AGENTS.md) to your global `~/.codex/AGENTS.md` (create the file if it doesn't exist yet). That's a separate, global instructions file Codex reads in every project — without it, Codex only calls these tools when you explicitly say "use the MCP".
## Standalone CLI
`src/mcps/energy_estimate.py` also works as a plain script, independent of MCP:
```bash
uv run carbon-tracking-energy-estimate ~/.claude/projects/<project>/<session-id>.jsonl
```
```
60 deduplicated requests
Estimated session energy: 1.0181 kWh
```
## The energy model
Follows [Simon P. Couch's methodology](https://simonpcouch.com/blog/2026-01-20-cc-impact/): Wh-per-million-tokens rates are estimated from Epoch AI's ChatGPT-4o energy figures, using the *price ratio* between input/output/cache tokens as a proxy for their *energy ratio* (Anthropic doesn't publish energy numbers directly).
Couch's article was published when models still used 200k token context windows. Models have a context window of 1M tokens now. We assume one is smart and keeps their context window around half this. So, we added 500K-token context as a fourth anchor point, based on code from EpochAI: https://colab.research.google.com/drive/1dnhL0lkjsk-isAH-j02pFUbv1g12hEm1#scrollTo=dN0Ezyr4qXAQ, which is used to compute energy.
**Read this before trusting the numbers:**
- "Energy scales with price" is an assumption, not a measurement.
- Cache-read/cache-write rates are a flat napkin-math ratio applied to the input rate, not real per-model pricing.
- A single fixed rate is used for every request regardless of its actual context length, so short requests are overestimated and very long ones (near 1M tokens) are underestimated relative to a context-scaled model.
- EpochAI note that cost/energy scales quadratically with input length as expected due to the attention mechanism. However, there are certainly innovations that improve on quadratic scaling. So this is a pessimistic estimation.
Treat every number here as **order-of-magnitude and directional** — useful for comparing sessions against each other, not as an audited carbon/energy figure.
## CO2eq comparisons
`carbon_equivalents.py` converts an energy estimate (Wh) into kgCO2eq using a grid carbon intensity figure (default: the 2025 UK average of 218 gCO2e/kWh, from [Ember's yearly electricity data](https://files.ember-energy.org/public-downloads/generation/outputs/release_generation_yearly_global.csv)) and expresses that total against everyday activities defined in `carbon_equivalents.json` — washing machine cycles, EV charges, flights, a serving of beef, and so on. Entries in that file are direct `kg_co2` figure. See the JSON for sources.
Same caveat as above: this is a rough, directional comparison, not an audited figure — grid intensity varies by country, time of day, and year.
## Dashboard
A small web dashboard lets anyone (for example a sustainability team) enter input, output and cached token counts and see the energy, CO2e and everyday equivalents, with a grid-intensity selector. It will be published at <https://nathanwbailey.github.io/carbon_tracking_mcp/>.
The dashboard reads its rates from `web/src/data/model.json`, generated from the Python model, so it cannot drift from the MCP tools: `uv run pytest` fails if the file is stale. Regenerate it with `uv run python scripts/export_dashboard_data.py`.
```bash
cd web
npm install
npm run dev # local dev server
npm test # parity tests against the Python model
npm run build # production build into web/dist
```
Pushes to `main` that touch the dashboard or the model deploy it via `.github/workflows/pages.yml` (repo Settings -> Pages -> Source: GitHub Actions).
## Project layout
```
src/mcps/
schema.py # pydantic models (with field descriptions) for every tool input/output
energy_estimate.py # the energy model + Claude Code/Codex session-log parsers (also runnable as a CLI)
carbon_equivalents.py # Wh -> kgCO2eq conversion + everyday-activity comparisons
carbon_equivalents.json # the comparison database (grid intensity + activity list)
mcp_results.py # shared MCP tool-result shaping used by every server below
claude_sessions.py # Claude Code session discovery (glob under ~/.claude/projects)
codex_sessions.py # Codex session discovery (sqlite thread index + ~/.codex/sessions)
stdio/
server_claude.py # FastMCP server for Claude Code sessions
server_codex.py # FastMCP server for Codex sessions
scripts/
export_dashboard_data.py # writes web/src/data/model.json from the Python model (--check in CI)
web/ # Vite + React dashboard, deployed to GitHub Pages
```
## License
[MIT](LICENSE)
TDQS
Scored across 2 tools
The two tools have clearly distinct scopes: one reports energy/CO2eq for the current Claude Code session, while the other aggregates all sessions in the current project. There is no meaningful overlap in purpose, making selection straightforward.
Both tools use snake_case, which is consistent. However, one is a noun phrase (current_session_energy) and the other uses a verb+noun pattern (collate_project_sessions_energy), a minor grammatical inconsistency in an otherwise predictable scheme.
The server is narrowly scoped, but two tools feels thin for a carbon-tracking surface. It is borderline: enough for a minimal monitor, but likely too few to support broader tracking workflows.
The tools cover current-session and current-project energy/CO2eq, but notable gaps remain. There is no way to query historical trends, other projects, or account-wide totals, which limits the surface for a carbon-tracking server.