Skip to main content
Glama
README.md
# AI Race Engineer

An AI race engineer for Formula 1: reads telemetry, models tyre degradation, evaluates pit
strategy, and delivers calls the way a real engineer would — short, timely, and grounded in data.

Built on [LangGraph](https://docs.langchain.com/oss/python/langgraph/overview) for orchestration
and [MCP](https://modelcontextprotocol.io/) for the tool surface, over
[FastF1](https://docs.fastf1.dev/) telemetry.

![Verstappen's fuel-corrected pace across four stints at the 2024 Austrian GP, with fitted degradation lines](plots/2024_Austria_VER.png)

Every point is a real, cleaned, fuel-corrected lap. Every line is the model fitted to it.

Three things are worth noticing:

- **The HARD line is shallower than the MEDIUM lines** (+0.100 vs +0.135 s/lap). That ordering
  is nowhere in the code — it falls out of the fit, and it's the evidence the lap cleaning and
  fuel correction are working.
- **The gaps between stints** are the pit stops, plus every lap thrown out as unrepresentative:
  safety car, in-laps, out-laps, timing glitches. 8 of VER's 71 laps didn't survive.
- **The dashed red line slopes downward**, which is physically nonsense — tyres don't get faster
  as they age. That's a 5-lap end-of-race stint, and the system flags it as a weak fit (R²=0.31)
  rather than reporting it as a finding. Knowing when the data can't support an answer is the
  hard half of this problem.

### The same result on an independent sample

Norris, same race, fitted separately — three stints, no weak fits:

![Norris's fuel-corrected pace across three stints at the 2024 Austrian GP, with fitted degradation lines](plots/2024_Austria_NOR.png)

| | MEDIUM (stint 1) | HARD | MEDIUM (stint 3) |
|---|---|---|---|
| **Verstappen** | +0.135 s/lap | **+0.100** | +0.199 |
| **Norris** | +0.117 s/lap | **+0.097** | +0.238 |

Two different cars, two different drivers, fitted independently — and both put the HARD below
the MEDIUM by a similar margin. One fit could be luck; two matching fits on independent data is
the method working.

The third-stint figures are worth a look too: both drivers roughly double their opening MEDIUM
degradation on the same compound, consistent with a hotter track and older tyres late on.
Nothing in the code looks for that — it falls out of the fit.

Generate either for any race since 2018:

```bash
race-engineer plot --year 2024 --circuit Austria --driver VER
```

## Separating fuel burn from tyre wear

Both make lap times change over a stint, and getting this wrong invalidates everything
downstream. A car sheds ~100 kg of fuel across a race and gets quicker as it does; fit raw lap
times and the model concludes tyres get *faster* with age, because fuel burn outruns tyre wear.

The obvious fix — subtract a fuel correction — just relocates the problem. **How much** per kg
is circuit-specific: it scales with how much of the lap is spent accelerating, so a twisty
circuit is far more mass-sensitive than a flat-out one. One global constant is the crudest
possible assumption, and at a low-degradation circuit it can swamp the signal entirely.

### Why you can't just regress it out

Within a single stint, fuel load and tyre age are **both linear in lap number**. They are
perfectly collinear. No regression separates them.

What breaks the tie is the pit stop: **tyre age resets to zero, fuel does not.** So if a
compound runs at two different points in the race, the pace difference at equal tyre age is
attributable to fuel. That makes this identifiable:

```
lap_time = base[compound] + deg[compound] · tyre_age + k · fuel_kg
```

One intercept and one slope per compound, plus a single shared `k`. The compound intercepts
absorb pace differences, the slopes absorb degradation, and `k` is identified purely by the
resets. (Same structure as the state-space treatment in
[arXiv:2512.00640](https://arxiv.org/abs/2512.00640), reduced to ordinary least squares.)

```bash
race-engineer analyse --year 2024 --circuit Monza --drivers NOR LEC --fit-fuel
```
```
NOR: fuel effect 0.027 s/lap/kg (fitted, R²=0.75)
LEC: fuel effect 0.030 s/lap/kg (default — design is rank-deficient)
```

Leclerc's fallback is the estimator working, not failing. He ran MEDIUM then HARD, one stint
each — no compound repeats, so nothing separates fuel from wear. It refuses rather than
returning a confident wrong number.

### What this settled

Monza fits nearly flat — degradation of +0.000 to +0.015 s/lap, every stint flagged weak. That
left an open question: is Monza genuinely low-degradation, or is the global constant wrong
there? Fitting the coefficient answers it. At 0.027 s/kg the degradation is still ~zero, so
**it's the circuit, not the correction.**

### It's opt-in, for a stated reason

The estimator assumes **one degradation slope per compound**. Austria violates that: the two
MEDIUM stints genuinely differ (+0.135 vs +0.199 s/lap, from track evolution), which biases the
fit low. It's a real improvement in principle and not yet reliable enough to be the default, so
it sits behind `--fit-fuel` until the per-stint case is handled.

## The core design constraint

An LLM agent loop is far too slow to be a race engineer. Real calls are sub-second and mostly
reflexive. So the system is split by timescale:

```
┌── FAST LOOP (deterministic Python, no LLM) ────────────┐
│  telemetry tick → rule engine → alerts                 │
│  "box this lap", "blue flags", "P2 within DRS"         │
│  Latency budget: <100 ms                               │
└────────────────────┬───────────────────────────────────┘
                     │ writes to shared RaceState
┌────────────────────▼───────────────────────────────────┐
│  SLOW LOOP (LangGraph + LLM)                           │
│  strategy reasoning, undercut math, debriefs, Q&A      │
│  Latency budget: 2–30 s                                │
└────────────────────────────────────────────────────────┘
```

There are **zero LLM calls on the critical path**. The model supplies judgment, not reflexes.

## Graph shape

```
START → strategist ⇄ tools → radio → END
```

`strategist` loops against 7 MCP tools until it stops asking for them, then everything funnels
through `radio` — so the driver hears exactly one message per cycle, however much analysis
happened. `radio` is the only node that speaks to the driver, and its entire job is compression:
a race engineer says *"Box, box, undercut Norris"*, not three paragraphs.

Routing between the two is a plain function, not a model call. It's decidable from state, so the
round trip bought nothing. The original design had an LLM router picking between four
specialists; dropping it removed a failure mode and a few hundred milliseconds.

## It works on real races

```
$ race-engineer analyse --year 2024 --circuit Austria --drivers VER NOR

2024 Austrian Grand Prix — Race, 71 laps

VER: 63 clean laps (8 dropped as unrepresentative)
  Stint 1 MEDIUM    laps  1-23  deg +0.135 s/lap  base 66.34s  R²=0.98
  Stint 2 HARD      laps 24-51  deg +0.100 s/lap  base 67.09s  R²=0.89
  Stint 3 MEDIUM    laps 52-64  deg +0.199 s/lap  base 66.82s  R²=0.68
  Stint 4 SOFT      laps 65-71  deg -0.553 s/lap  base 73.19s  R²=0.31  ⚠ weak fit — indicative only

NOR: 58 clean laps (6 dropped as unrepresentative)
  Stint 1 MEDIUM    laps  1-23  deg +0.117 s/lap  base 66.82s  R²=0.87
  Stint 2 HARD      laps 24-51  deg +0.097 s/lap  base 67.32s  R²=0.93
  Stint 3 MEDIUM    laps 52-64  deg +0.238 s/lap  base 67.32s  R²=0.72
```

The HARD compound degrading slower than the MEDIUM isn't hard-coded — it falls out of the fit.
That ordering is the evidence the lap cleaning and fuel correction are working.

Stints flagged ⚠ have an R² too low to trust. A 5-lap end-of-race stint fits noise, and the
system says so rather than reporting a confident wrong number.

## Setup

```bash
uv sync
cp .env.example .env      # add ANTHROPIC_API_KEY — only needed for the agent
```

## Usage

```bash
# Degradation analysis — no LLM, no API key
race-engineer analyse --year 2024 --circuit Austria --drivers VER NOR

# Ask the engineer — needs ANTHROPIC_API_KEY
race-engineer ask "Box now for the hard, or stay out?" \
    --year 2024 --circuit Austria --drivers VER NOR

# Run either MCP server standalone (works in any MCP client)
race-engineer serve-data
race-engineer serve-strategy
```

**On Kaggle:** open [`notebooks/kaggle_demo.ipynb`](notebooks/kaggle_demo.ipynb). Set
**Internet → On**, add `ANTHROPIC_API_KEY` under **Add-ons → Secrets**, and set the accelerator
to **None** — this is all CPU work.

## Status

| Phase | State |
|---|---|
| 1. Data layer — FastF1 loading, lap cleaning, fuel correction | ✅ Working on real races |
| 2. Strategy engine — degradation, crossover, undercut, pit window | ✅ 35 tests passing |
| 3. Agent layer — LangGraph + 7 MCP tools | ✅ End to end |
| 4. Evaluation harness — replay and score the calls | ⬜ Next, and the one that matters |
| 5–7. Live timing, voice, sim racing | ⬜ |

See [ROADMAP.md](ROADMAP.md) for detail and [docs/architecture.md](docs/architecture.md) for
the design.

## Development

```bash
uv run pytest      # tests
uv run ruff check  # lint
```

## Data sources

| Source | Cost | Used for |
|---|---|---|
| [FastF1](https://docs.fastf1.dev/) | Free, no key | Telemetry, laps, stints (2018→) |
| [Jolpica-F1](https://github.com/jolpica/jolpica-f1) | Free, no key | Results and standings (1950→) |
| [OpenF1](https://openf1.org/) | Free historical / paid live | Real-time timing |
| [Anthropic API](https://platform.claude.com/) | Paid | The reasoning layer |

Full breakdown, including which are optional, in [docs/apis.md](docs/apis.md).

## License

MIT

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation4/5

The tools are largely distinct: session loading, stint data, degradation modeling, and multi-driver comparison. The only potential overlap is between get_degradation and compare_drivers, but the single-vs-multi-driver distinction is clear from descriptions.

Naming Consistency4/5

Three tools follow the 'get_' verb-noun pattern, while compare_drivers uses a different verb without 'get'. This is a minor deviation, but the names remain predictable and readable overall.

Tool Count4/5

With 4 tools, the server is well-scoped for its tyre-focused purpose. It covers the essential workflow without bloat, though it is on the smaller side, leaving room for additional specialized tools.

Completeness4/5

The server covers the full tyre degradation analysis lifecycle: session loading, stint retrieval, degradation fitting, and cross-driver comparison. Minor gaps like raw lap times or session context tools exist, but they are not critical for the stated purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues