Skip to main content
Glama
README.md
# lpagent: load profiles for power-systems research

`lpagent` generates yearly electricity load time series (P and Q) for every load of a
pandapower grid, shaped by the German **BDEW standard load profiles** (1999 and 2025 sets),
with a documented stochastic model, and writes each run with everything needed to
regenerate and cite it.

You can drive it in four ways:

| Interface | Command | For |
|---|---|---|
| Public web app | `streamlit run streamlit_app.py` (hosting: see [DEPLOY.md](DEPLOY.md)) | a form (no LLM) and a chat assistant, multi-user, uploads, ZIP downloads |
| Chat (browser, local) | `streamlit run agent_app.py` | single-user; describing what you need in plain language; the LLM asks for missing information and calls the tools |
| Chat (terminal) | `python -m lpagent.cli` | the same, in a terminal |
| Scripted (no LLM) | `python -m lpagent.replay <manifest.json>` or the Python API | batch work, exact regeneration, no API key |

The LLM only chooses tools and parameters. Every number is produced by deterministic,
tested Python code, and each run's `manifest.json` records all settings, the per-load
assignment, assumptions, software versions and SHA-256 checksums of the outputs.

## Install

```bash
conda env create -f environment.yml && conda activate lpagent      # or any Python >= 3.10 env
pip install -e .                                                    # or: pip install -r requirements-dev.txt
python -m pytest                                                    # 73 tests, ~1 min
```

For bit-identical reproduction of someone else's run use `pip install -r requirements-lock.txt`
(random streams are only guaranteed identical for the same numpy version; the manifest records it).

**LLM for the chat interfaces** (environment variables; never put keys in code):

| Provider | Variables | Notes |
|---|---|---|
| Google Gemini | `GEMINI_API_KEY`, optional `GEMINI_MODEL` | free tier; the newest stable Flash model is picked automatically, with fallback when overloaded |
| Groq | `GROQ_API_KEY`, optional `GROQ_MODEL` (default `openai/gpt-oss-120b`) | free tier: 7,000–8,000 tokens per request and about 200,000 per model per day (≈30 agent requests); the agent keeps requests under the cap and switches to `gpt-oss-20b` / `qwen3.8-27b` when a model's quota is used up |
| OpenAI / compatible | `OPENAI_API_KEY`, optional `OPENAI_MODEL`, `OPENAI_BASE_URL` | Azure, vLLM etc. via base URL |
| Anthropic | `ANTHROPIC_API_KEY`, optional `ANTHROPIC_MODEL` | |
| Ollama (local) | `LLM_PROVIDER=ollama`, `OLLAMA_MODEL` | needs a model with reliable tool calling and a GPU; not recommended on laptops |

Select the provider in the app's sidebar or with `--provider`.

**What the agent does to stay reliable with any model** (all covered by tests):

- The LLM never produces numbers; it only chooses tools and parameters. Answers may report only
  values present in tool results, and the manifest is the source of truth.
- A *session state* note (grid, assignment, registered profiles, recent runs and their inputs) is
  maintained from tool results and never trimmed, so follow-ups work even in long conversations.
- Per-request token budgets are self-calibrating from the provider's own token counts; older
  turns are shortened before anything in the current turn.
- Rate limits are waited out, overloaded or exhausted models fall back to alternatives
  (Gemini Flash generations, Groq's tool-capable models), and malformed tool calls rejected by a
  provider are sent back to the model for correction.
- A tool call repeated with identical arguments after two identical failures ends the turn with a
  message about what is missing, instead of looping.

Stronger models (Gemini Flash, gpt-oss-120b, Claude, GPT) follow the policy more reliably than
small ones; the guardrails above are what make small free models usable.

## Public web app

`streamlit_app.py` is the version for a public service: every visitor gets an isolated session
(private temporary folder, own MCP server instance, deleted after 2 h idle), grids and profile
CSVs are uploaded instead of referenced by path, results are delivered as one ZIP per run, runs
are size-capped, the chat is rate-limited and falls back from Gemini to Groq, and a **form** tab
generates profiles without any LLM. Uploaded pandapower files are scanned before loading, because
pandapower's JSON/Excel readers can execute code named in a file. Deployment to Streamlit
Community Cloud (free): [DEPLOY.md](DEPLOY.md).

## Using the results

Outputs are written to `outputs/<run_id>/`:

```
manifest.json      all settings, per-load assignment, assumptions, statistics, checksums, versions
METHODS.md         methods description with equations and references, generated from the manifest
load_info.csv      load index, bus, base p/q, category, profile spec, noise sigma
P_<year>.csv       active power [MW], one column per pandapower load index, timestamp = interval start
Q_<year>.csv       reactive power [Mvar]
scenario_<k>/...   the same per scenario when n_scenarios > 1 (Monte Carlo)
profiles/          copies of user-supplied profiles used by the run
```

**pandapower time-series power flow** in three lines:

```python
from lpagent.timeseries import attach_run, run_power_flow
res = run_power_flow("outputs/ieee33_20260929-170946_d25d", year=2012, time_steps=range(0, 168))
res["res_bus.vm_pu"]          # DataFrame indexed by timestamp; also res_line.loading_percent, res_ext_grid.p_mw
# or attach controllers to your own net and run pandapower.timeseries.run_timeseries yourself:
net = pandapower.networks.case33bw(); attach_run(net, "outputs/ieee33_...", year=2012)
```

**Regenerate and verify a run** (no LLM, no key):

```bash
python -m lpagent.replay outputs/ieee33_20260929-170946_d25d          # exit code 0 = all files identical
```

**Scripted generation:** copy a manifest, edit `config` and/or `assignment.method`, delete
`assignment.loads` and `checksums`, and replay it. Or use the Python API:

```python
from lpagent import grids, assignment, generator, runs
h = grids.load_grid("ieee33")
a = assignment.build_assignment(h, "a1", distribution={"residential": 80, "commercial": 20}, basis="power")
cfg = generator.GenerationConfig(years=[2025, 2026], resolution_min=15, n_scenarios=10)
gen = generator.generate(h, a, cfg)
runs.RunStore("outputs").save("ieee33_mc", h, a, cfg, gen)
```

## Method (summary; each run's METHODS.md has the full text with references)

For load *i* at time *t* in year *y*:

    P_i(t) = B_i · s_p(i)(t − τ_i) · κ_i · g(y) · n_i(t)          Q_i(t) = P_i(t) · tan φ_i

| Symbol | Meaning | Options |
|---|---|---|
| s | BDEW profile of the load's category, 1,000 kWh/a normalised, quarter-hourly, holidays as Sundays, Dec 24/31 as Saturdays, household profiles dynamised; identical to the R reference implementation to 1e-9 W | 2025 set (H25, G25, L25, P25, S25) default; 1999 set (H0, G0–G6, L0–L2); **mixes** `0.7*H25+0.3*G25` (energy-weighted); **user profiles** from CSV (daily, weekly or yearly pattern) |
| B | base value | grid `p_mw` = annual **peak** (default) or annual **mean**; or **annual energy** in MWh per load/category |
| τ_i, κ_i | per-load time shift and scaling factor (customer diversity) | off by default |
| g | load growth | %/a |
| n_i | temporal noise, Gaussian **AR(1)** (default, φ = 0.8 at 1 h) or white or none; σ_i = σ_ref·√(P_ref/P̄_i): **larger loads fluctuate less** (a load is an aggregate of customers); optional **common component** shared by all loads | σ_ref = 10 % at the grid's median load; all parameters adjustable |
| tan φ | grid's q/p per load (default) or power factor per category | BDEW has no Q profiles |
| categories | residential, commercial, industrial (G3 proxy: BDEW has no industrial profile), agricultural, residential_pv, residential_pv_battery, mixed | assigned from grid metadata (CIGRE), an explicit per-load/bus mapping, and/or a distribution by count or power, by size or at random |

Timestamps are local standard time (no DST), 96 quarter-hours per day, leap years with 366 days;
values are interval averages. Noise parameters are modelling assumptions unless you calibrate them.

Built-in grids: CIGRE LV and MV (load types known), IEEE 33-bus, MV Oberrhein, Kerber Dorfnetz,
LV Schutterwald (types unknown; the agent asks), plus any pandapower `.json`/`.xlsx` file.

## MCP tools (also usable from Claude Desktop / Claude Code)

`list_grids`, `load_grid`, `get_grid_loads`, `list_bdew_profiles`, `get_bdew_curve`,
`register_profile`, `assign_load_types`, `get_default_settings`, `generate_profiles`,
`plot_run`, `get_run` (sections: summary, methods, assumptions, config, assignment, scenarios,
verify), `list_runs`. Server: `python -m lpagent.server` (stdio). The agent's behaviour policy is
the server's MCP `instructions`, so every client follows the same rules (ask for unknown load
types, show the plan, report only tool numbers).

## Project layout

```
lpagent/
  bdew.py          BDEW engine (port of standardlastprofile), holidays, resampling
  profiles.py      profile specs: BDEW ids, custom CSV profiles, energy-weighted mixes
  grids.py         pandapower grids, load-type detection
  assignment.py    category / profile assignment
  noise.py         AR(1)/white noise, size-dependent sigma, common component, per-load scaling and shift
  generator.py     P/Q generation, scenarios, statistics
  runs.py          run folder: CSVs, manifest with checksums, METHODS.md
  report.py        methods text from a manifest
  replay.py        regenerate + verify a run (CLI)
  timeseries.py    pandapower DFData/ConstControl helper and time-series power flow
  server.py        MCP server (tools + agent policy); Session = one user's state and limits
  agent.py         LLM tool-calling loop over MCP (Gemini, Groq, OpenAI, Anthropic, Ollama)
  cli.py           terminal chat
  webapp.py        multi-user sessions, uploads, cleanup, chat quotas (public app)
  data/slp_bdew.csv  all 16 BDEW profiles (built by scripts/build_bdew_table.py)
streamlit_app.py   public web app (form + chat); deployment: DEPLOY.md
agent_app.py       local single-user chat UI
tests/             73 tests incl. comparison with the R reference (runs if R is installed)
tests/live/        reliability battery against real LLM providers (needs API keys)
standardlastprofile/  vendored R package (CC0): source data and reference implementation
legacy/            earlier prototype, not used
```

## Licence

MIT (see LICENSE). The BDEW profile data come from the CC0-licensed R package standardlastprofile.

## Data sources

- VDEW (1999), *Repräsentative VDEW-Lastprofile*; VDEW (2000), *Anwendung der repräsentativen VDEW-Lastprofile step-by-step*; BDEW (2025), *Standardlastprofile Strom*: https://www.bdew.de/energie/standardlastprofile-strom/
- Döring, M., *standardlastprofile* R package (CC0), https://doi.org/10.32614/CRAN.package.standardlastprofile
- Thurner et al. (2018), pandapower, IEEE Trans. Power Systems 33(6), https://doi.org/10.1109/TPWRS.2018.2829021

## Limitations

- BDEW profiles describe German customer groups; other countries only change the holiday calendar.
- No industrial standard profile exists; supply a measured one with `register_profile` if you have it.
- Noise parameters are defaults, not fitted values. Heat pumps and EVs are not represented.
- Reactive power uses a constant power factor per load.

TDQS

A3.5/5.0

Scored across 12 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: listing/loading grids, inspecting loads, managing BDEW profiles, assigning load types, generating/plotting/retrieving runs. The few closely related tools (list_grids vs load_grid, load_grid vs get_grid_loads) are differentiated by their descriptions (summary vs detailed table vs listing). No two tools appear to do the same thing.

Naming Consistency5/5

All 12 tools use snake_case with a consistent verb_noun pattern (list_*, load_*, get_*, register_*, assign_*, generate_*, plot_*). The verbs are appropriate and predictable, making the set easy to navigate.

Tool Count5/5

12 tools is well within the ideal 3–15 range for a specialized load-profile agent. Each tool covers a distinct step in the workflow, and there is no obvious redundancy or missing core operation that would inflate the count.

Completeness4/5

The surface covers the end-to-end workflow: list/load grids, inspect loads, manage BDEW profiles, assign types, generate profiles, and retrieve/plot runs. Minor gaps exist—no update/delete for registered profiles and no per-load assignment retrieval—but these are not likely to block typical agent tasks.

Maintenance

ActivityMaintained
ResponsivenessNo issues