Skip to main content
Glama
nathanwbailey

Carbon Tracking MCP

Carbon Tracking MCP

MCP servers that estimate the energy used by your Claude Code and Codex sessions, so you can ask for it directly from within a chat: "how much energy has this chat used?" or "how much has this whole project cost?"

They wrap a small pricing-ratio energy model (src/mcps/energy_estimate.py) around each tool's local session logs — no telemetry, no network calls. The Claude Code server reads ~/.claude/projects/**/*.jsonl; the Codex server reads Codex's local thread index (~/.codex/state_5.sqlite) and rollout logs under ~/.codex/sessions/. Both run over the stdio MCP transport (src/mcps/stdio/), spawned fresh per chat by Claude Code/Codex.

What it does

Each server exposes the same two MCP tools, scoped to its own provider:

Tool

Answers

current_session_energy

How much energy has this chat used so far?

collate_project_sessions_energy

How much energy has every chat in this project used, in total?

Example output:

{
  "session_id": "afc721a3-0d77-4c5f-b1f5-24074d03fa7d",
  "file": "/Users/you/.claude/projects/-Users-you-my-project/afc721a3-....jsonl",
  "request_count": 60,
  "estimated_kwh": 1.018,
  "estimated_kg_co2": 0.144,
  "comparisons": [
    {"id": "washing_machine_cycle", "label": "washing machine cycle", "count": 0.21},
    {"id": "ev_car_km", "label": "km driven in a 2020 Tesla Model 3", "count": 1.78}
  ]
}
{
  "project_dir": "/Users/you/.claude/projects/-Users-you-my-project",
  "session_count": 3,
  "total_estimated_kwh": 1.545,
  "sessions": [
    {"session_id": "afc721a3-...", "request_count": 60, "estimated_kwh": 1.018},
    {"session_id": "8ee8cb6f-...", "request_count": 5, "estimated_kwh": 0.073},
    {"session_id": "2aabc643-...", "request_count": 38, "estimated_kwh": 0.86}
  ],
  "estimated_kg_co2": 0.218,
  "comparisons": [
    {"id": "washing_machine_cycle", "label": "washing machine cycle", "count": 0.31}
  ]
}

estimated_kg_co2 and comparisons (both tools' full comparison list is longer than shown above — see carbon_equivalents.json) convert the estimated energy into CO2eq using a rough UK grid carbon intensity figure, then express it against everyday activities (washing machine cycles, EV charges, flights, ...). See "CO2eq comparisons" below.

Both tools return a typed, field-described pydantic model (src/mcps/schema.py), so MCP clients get a real JSON schema for the response shape rather than an untyped object. If a session/project can't be found, the tool raises a proper MCP tool error instead of returning a disguised "successful" result.

On the Claude Code server, current_session_energy identifies "this chat" via the CLAUDE_CODE_SESSION_ID environment variable that Claude Code sets on every process it launches (including the server). Codex doesn't set an equivalent env var for MCP subprocesses it spawns, so the Codex server checks CODEX_THREAD_ID opportunistically and otherwise falls back to the most-recently-modified indexed thread whose recorded cwd matches the server's — a best-effort heuristic that's right unless you have multiple Codex chats open in the same project directory at once. Either tool's collate_project_sessions_energy then sums every other session belonging to the current project — every sibling .jsonl for Claude Code, every indexed thread with a matching cwd for Codex.

Related MCP server: Regen Compute

Install

Requires uv.

git clone https://github.com/<you>/carbon-tracking-mcp.git
cd carbon-tracking-mcp
uv sync

This installs console-script entry points (carbon-tracking-stdio-claude, carbon-tracking-stdio-codex) backed by the mcps package under src/.

Register whichever server(s) you use globally, so they're available in every project rather than just this one.

Use uv run --project (not --directory) to launch them: --directory changes the subprocess's working directory to this repo before running, which breaks collate_project_sessions_energy()'s cwd-based "which project is this?" scoping for every other project you use these servers from. --project only points uv at this repo for dependency resolution and leaves the subprocess's cwd alone.

Claude Code

claude mcp add --scope user carbon-impact-claude -- uv run --project /path/to/carbon-tracking-mcp carbon-tracking-stdio-claude

Restart or start a new session and the tools become available. Verify with:

claude mcp get carbon-impact-claude

Codex

codex mcp add carbon-impact-codex -- uv run --project /path/to/carbon-tracking-mcp carbon-tracking-stdio-codex

codex mcp add writes to ~/.codex/config.toml, which Codex CLI, the IDE extension, and the desktop app all share — there's no per-project scope to choose, so this is global by default. Restart or start a new session and check /mcp inside Codex to verify the server is connected. Alternatively, add the entry by hand:

[mcp_servers.carbon-impact-codex]
command = "uv"
args = ["run", "--project", "/path/to/carbon-tracking-mcp", "carbon-tracking-stdio-codex"]

Codex won't reliably call the tools on its own unless you also tell it to: append this repo's AGENTS.md to your global ~/.codex/AGENTS.md (create the file if it doesn't exist yet). That's a separate, global instructions file Codex reads in every project — without it, Codex only calls these tools when you explicitly say "use the MCP".

Standalone CLI

src/mcps/energy_estimate.py also works as a plain script, independent of MCP:

uv run carbon-tracking-energy-estimate ~/.claude/projects/<project>/<session-id>.jsonl
60 deduplicated requests
Estimated session energy: 1.0181 kWh

The energy model

Follows Simon P. Couch's methodology: Wh-per-million-tokens rates are estimated from Epoch AI's ChatGPT-4o energy figures, using the price ratio between input/output/cache tokens as a proxy for their energy ratio (Anthropic doesn't publish energy numbers directly).

Couch's article was published when models still used 200k token context windows. Models have a context window of 1M tokens now. We assume one is smart and keeps their context window around half this. So, we added 500K-token context as a fourth anchor point, based on code from EpochAI: https://colab.research.google.com/drive/1dnhL0lkjsk-isAH-j02pFUbv1g12hEm1#scrollTo=dN0Ezyr4qXAQ, which is used to compute energy.

Read this before trusting the numbers:

  • "Energy scales with price" is an assumption, not a measurement.

  • Cache-read/cache-write rates are a flat napkin-math ratio applied to the input rate, not real per-model pricing.

  • A single fixed rate is used for every request regardless of its actual context length, so short requests are overestimated and very long ones (near 1M tokens) are underestimated relative to a context-scaled model.

  • EpochAI note that cost/energy scales quadratically with input length as expected due to the attention mechanism. However, there are certainly innovations that improve on quadratic scaling. So this is a pessimistic estimation.

Treat every number here as order-of-magnitude and directional — useful for comparing sessions against each other, not as an audited carbon/energy figure.

CO2eq comparisons

carbon_equivalents.py converts an energy estimate (Wh) into kgCO2eq using a grid carbon intensity figure (default: the 2025 UK average of 218 gCO2e/kWh, from Ember's yearly electricity data) and expresses that total against everyday activities defined in carbon_equivalents.json — washing machine cycles, EV charges, flights, a serving of beef, and so on. Entries in that file are direct kg_co2 figure. See the JSON for sources.

Same caveat as above: this is a rough, directional comparison, not an audited figure — grid intensity varies by country, time of day, and year.

Dashboard

A small web dashboard lets anyone (for example a sustainability team) enter input, output and cached token counts and see the energy, CO2e and everyday equivalents, with a grid-intensity selector. It will be published at https://nathanwbailey.github.io/carbon_tracking_mcp/.

The dashboard reads its rates from web/src/data/model.json, generated from the Python model, so it cannot drift from the MCP tools: uv run pytest fails if the file is stale. Regenerate it with uv run python scripts/export_dashboard_data.py.

cd web
npm install
npm run dev     # local dev server
npm test        # parity tests against the Python model
npm run build   # production build into web/dist

Pushes to main that touch the dashboard or the model deploy it via .github/workflows/pages.yml (repo Settings -> Pages -> Source: GitHub Actions).

Project layout

src/mcps/
  schema.py                # pydantic models (with field descriptions) for every tool input/output
  energy_estimate.py       # the energy model + Claude Code/Codex session-log parsers (also runnable as a CLI)
  carbon_equivalents.py    # Wh -> kgCO2eq conversion + everyday-activity comparisons
  carbon_equivalents.json  # the comparison database (grid intensity + activity list)
  mcp_results.py           # shared MCP tool-result shaping used by every server below
  claude_sessions.py       # Claude Code session discovery (glob under ~/.claude/projects)
  codex_sessions.py        # Codex session discovery (sqlite thread index + ~/.codex/sessions)
  stdio/
    server_claude.py       # FastMCP server for Claude Code sessions
    server_codex.py        # FastMCP server for Codex sessions
scripts/
  export_dashboard_data.py # writes web/src/data/model.json from the Python model (--check in CI)
web/                       # Vite + React dashboard, deployed to GitHub Pages

License

MIT

Available Tools

2 tools
collate_project_sessions_energyCollate Project Sessions EnergyB

Total + per-session energy and CO2eq for every Claude Code chat in the current project.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
providerYesCoding-agent provider these sessions belong to ('claude' or 'codex').
sessionsYesPer-session breakdown of request count and estimated energy.
comparisonsYesEveryday-activity equivalents for the total estimated CO2eq.
project_dirYesAbsolute path to the project's session-log directory that was scanned.
session_countYesNumber of sessions included in this total.
estimated_kg_co2YesEstimated CO2-equivalent emissions for the total, in kilograms.
total_estimated_kwhYesTotal estimated energy across all included sessions, in kilowatt-hours.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says nothing about read-only safety, computation cost over many sessions, or any limits, leaving real gaps for a tool that aggregates across an entire project.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the measured quantities and scope front-loaded and no filler. The '+' shorthand for total-plus-per-session is slightly cryptic but compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description states the metrics and scope needed to invoke a zero-param tool. Only the sibling relationship and any cost caveat are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics for the description to add. Baseline 4 applies per the scoring rule for 0-param tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (collate) and resource (project sessions energy/CO2eq) and clearly scopes it to 'every Claude Code chat in the current project'. The project-wide scope implicitly distinguishes it from the current_session_energy sibling, but the sibling is never named, so an agent must infer the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The project-wide scope hints at when this is appropriate versus a single-session tool, but there is no explicit when-to-use, when-not-to-use, or named alternative. Usage is left to inference rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

current_session_energyCurrent Session EnergyB

Estimated energy (kWh) and CO2eq for the calling Claude Code session.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYesAbsolute path to the session's local log file that was parsed.
providerYesCoding-agent provider this session belongs to ('claude' or 'codex').
session_idYesClaude Code session ID or Codex thread ID being reported on.
comparisonsYesEveryday-activity equivalents for this session's estimated CO2eq.
estimated_kwhYesEstimated energy used by this session, in kilowatt-hours.
request_countYesNumber of deduplicated API requests/turns in this session.
estimated_kg_co2YesEstimated CO2-equivalent emissions for this session, in kilograms.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It never states that this is a side-effect-free read, nor does it disclose the reliability of the estimate, latency, or behavior when no session data exists; 'estimated' hints at approximation but that is all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the two returned quantities and the scope. Zero filler and no redundancy with the schema or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and a zero-parameter tool is simple. However, the near-total absence of usage and behavioral guidance against zero annotations leaves the agent to infer whether and when this should be called over its sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to explain. Per the baseline for parameter-less tools, a 4 is appropriate; the description adds nothing here, but nothing is missing either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource ('calling Claude Code session') and the quantified outputs ('energy (kWh) and CO2eq'), so an agent knows what is measured. It implies a read of the current session, distinguishing it from the project-wide sibling, but never names collate_project_sessions_energy explicitly, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, and no reference to the alternative tool for project-level aggregation. The scope word 'calling session' implicitly limits it, but the agent receives no explicit routing instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedcollate_project_sessions_energy
    • First observedcurrent_session_energy

TDQS

B3.4/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct scopes: one reports energy/CO2eq for the current Claude Code session, while the other aggregates all sessions in the current project. There is no meaningful overlap in purpose, making selection straightforward.

Naming Consistency4/5

Both tools use snake_case, which is consistent. However, one is a noun phrase (current_session_energy) and the other uses a verb+noun pattern (collate_project_sessions_energy), a minor grammatical inconsistency in an otherwise predictable scheme.

Tool Count3/5

The server is narrowly scoped, but two tools feels thin for a carbon-tracking surface. It is borderline: enough for a minimal monitor, but likely too few to support broader tracking workflows.

Completeness3/5

The tools cover current-session and current-project energy/CO2eq, but notable gaps remain. There is no way to query historical trends, other projects, or account-wide totals, which limits the surface for a carbon-tracking server.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Real-time Claude.ai subscription awareness for AI coding assistants. Surfaces live utilization, forecasts limits, gates expensive operations, and measures real per-task cost.
    5
    17 npm
    6
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    Estimates the environmental footprint of your AI use — energy (kWh), miles driven, water used for cooling, and CO₂ — plus a prompt-efficiency score, working with any AI client by measuring token usage.
    9
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Estimates data-center water consumption from local Claude Code transcripts, providing water usage summaries, breakdowns by model/project, and a compact HTML widget.
    MIT