Skip to main content
Glama

lpagent: load profiles for power-systems research

lpagent generates yearly electricity load time series (P and Q) for every load of a pandapower grid, shaped by the German BDEW standard load profiles (1999 and 2025 sets), with a documented stochastic model, and writes each run with everything needed to regenerate and cite it.

You can drive it in four ways:

Interface

Command

For

Public web app

streamlit run streamlit_app.py (hosting: see DEPLOY.md)

a form (no LLM) and a chat assistant, multi-user, uploads, ZIP downloads

Chat (browser, local)

streamlit run agent_app.py

single-user; describing what you need in plain language; the LLM asks for missing information and calls the tools

Chat (terminal)

python -m lpagent.cli

the same, in a terminal

Scripted (no LLM)

python -m lpagent.replay <manifest.json> or the Python API

batch work, exact regeneration, no API key

The LLM only chooses tools and parameters. Every number is produced by deterministic, tested Python code, and each run's manifest.json records all settings, the per-load assignment, assumptions, software versions and SHA-256 checksums of the outputs.

Install

conda env create -f environment.yml && conda activate lpagent      # or any Python >= 3.10 env
pip install -e .                                                    # or: pip install -r requirements-dev.txt
python -m pytest                                                    # 73 tests, ~1 min

For bit-identical reproduction of someone else's run use pip install -r requirements-lock.txt (random streams are only guaranteed identical for the same numpy version; the manifest records it).

LLM for the chat interfaces (environment variables; never put keys in code):

Provider

Variables

Notes

Google Gemini

GEMINI_API_KEY, optional GEMINI_MODEL

free tier; the newest stable Flash model is picked automatically, with fallback when overloaded

Groq

GROQ_API_KEY, optional GROQ_MODEL (default openai/gpt-oss-120b)

free tier: 7,000–8,000 tokens per request and about 200,000 per model per day (≈30 agent requests); the agent keeps requests under the cap and switches to gpt-oss-20b / qwen3.8-27b when a model's quota is used up

OpenAI / compatible

OPENAI_API_KEY, optional OPENAI_MODEL, OPENAI_BASE_URL

Azure, vLLM etc. via base URL

Anthropic

ANTHROPIC_API_KEY, optional ANTHROPIC_MODEL

Ollama (local)

LLM_PROVIDER=ollama, OLLAMA_MODEL

needs a model with reliable tool calling and a GPU; not recommended on laptops

Select the provider in the app's sidebar or with --provider.

What the agent does to stay reliable with any model (all covered by tests):

  • The LLM never produces numbers; it only chooses tools and parameters. Answers may report only values present in tool results, and the manifest is the source of truth.

  • A session state note (grid, assignment, registered profiles, recent runs and their inputs) is maintained from tool results and never trimmed, so follow-ups work even in long conversations.

  • Per-request token budgets are self-calibrating from the provider's own token counts; older turns are shortened before anything in the current turn.

  • Rate limits are waited out, overloaded or exhausted models fall back to alternatives (Gemini Flash generations, Groq's tool-capable models), and malformed tool calls rejected by a provider are sent back to the model for correction.

  • A tool call repeated with identical arguments after two identical failures ends the turn with a message about what is missing, instead of looping.

Stronger models (Gemini Flash, gpt-oss-120b, Claude, GPT) follow the policy more reliably than small ones; the guardrails above are what make small free models usable.

Related MCP server: PowerFactory MCP Automation

Public web app

streamlit_app.py is the version for a public service: every visitor gets an isolated session (private temporary folder, own MCP server instance, deleted after 2 h idle), grids and profile CSVs are uploaded instead of referenced by path, results are delivered as one ZIP per run, runs are size-capped, the chat is rate-limited and falls back from Gemini to Groq, and a form tab generates profiles without any LLM. Uploaded pandapower files are scanned before loading, because pandapower's JSON/Excel readers can execute code named in a file. Deployment to Streamlit Community Cloud (free): DEPLOY.md.

Using the results

Outputs are written to outputs/<run_id>/:

manifest.json      all settings, per-load assignment, assumptions, statistics, checksums, versions
METHODS.md         methods description with equations and references, generated from the manifest
load_info.csv      load index, bus, base p/q, category, profile spec, noise sigma
P_<year>.csv       active power [MW], one column per pandapower load index, timestamp = interval start
Q_<year>.csv       reactive power [Mvar]
scenario_<k>/...   the same per scenario when n_scenarios > 1 (Monte Carlo)
profiles/          copies of user-supplied profiles used by the run

pandapower time-series power flow in three lines:

from lpagent.timeseries import attach_run, run_power_flow
res = run_power_flow("outputs/ieee33_20260929-170946_d25d", year=2012, time_steps=range(0, 168))
res["res_bus.vm_pu"]          # DataFrame indexed by timestamp; also res_line.loading_percent, res_ext_grid.p_mw
# or attach controllers to your own net and run pandapower.timeseries.run_timeseries yourself:
net = pandapower.networks.case33bw(); attach_run(net, "outputs/ieee33_...", year=2012)

Regenerate and verify a run (no LLM, no key):

python -m lpagent.replay outputs/ieee33_20260929-170946_d25d          # exit code 0 = all files identical

Scripted generation: copy a manifest, edit config and/or assignment.method, delete assignment.loads and checksums, and replay it. Or use the Python API:

from lpagent import grids, assignment, generator, runs
h = grids.load_grid("ieee33")
a = assignment.build_assignment(h, "a1", distribution={"residential": 80, "commercial": 20}, basis="power")
cfg = generator.GenerationConfig(years=[2025, 2026], resolution_min=15, n_scenarios=10)
gen = generator.generate(h, a, cfg)
runs.RunStore("outputs").save("ieee33_mc", h, a, cfg, gen)

Method (summary; each run's METHODS.md has the full text with references)

For load i at time t in year y:

P_i(t) = B_i · s_p(i)(t − τ_i) · κ_i · g(y) · n_i(t)          Q_i(t) = P_i(t) · tan φ_i

Symbol

Meaning

Options

s

BDEW profile of the load's category, 1,000 kWh/a normalised, quarter-hourly, holidays as Sundays, Dec 24/31 as Saturdays, household profiles dynamised; identical to the R reference implementation to 1e-9 W

2025 set (H25, G25, L25, P25, S25) default; 1999 set (H0, G0–G6, L0–L2); mixes 0.7*H25+0.3*G25 (energy-weighted); user profiles from CSV (daily, weekly or yearly pattern)

B

base value

grid p_mw = annual peak (default) or annual mean; or annual energy in MWh per load/category

τ_i, κ_i

per-load time shift and scaling factor (customer diversity)

off by default

g

load growth

%/a

n_i

temporal noise, Gaussian AR(1) (default, φ = 0.8 at 1 h) or white or none; σ_i = σ_ref·√(P_ref/P̄_i): larger loads fluctuate less (a load is an aggregate of customers); optional common component shared by all loads

σ_ref = 10 % at the grid's median load; all parameters adjustable

tan φ

grid's q/p per load (default) or power factor per category

BDEW has no Q profiles

categories

residential, commercial, industrial (G3 proxy: BDEW has no industrial profile), agricultural, residential_pv, residential_pv_battery, mixed

assigned from grid metadata (CIGRE), an explicit per-load/bus mapping, and/or a distribution by count or power, by size or at random

Timestamps are local standard time (no DST), 96 quarter-hours per day, leap years with 366 days; values are interval averages. Noise parameters are modelling assumptions unless you calibrate them.

Built-in grids: CIGRE LV and MV (load types known), IEEE 33-bus, MV Oberrhein, Kerber Dorfnetz, LV Schutterwald (types unknown; the agent asks), plus any pandapower .json/.xlsx file.

MCP tools (also usable from Claude Desktop / Claude Code)

list_grids, load_grid, get_grid_loads, list_bdew_profiles, get_bdew_curve, register_profile, assign_load_types, get_default_settings, generate_profiles, plot_run, get_run (sections: summary, methods, assumptions, config, assignment, scenarios, verify), list_runs. Server: python -m lpagent.server (stdio). The agent's behaviour policy is the server's MCP instructions, so every client follows the same rules (ask for unknown load types, show the plan, report only tool numbers).

Project layout

lpagent/
  bdew.py          BDEW engine (port of standardlastprofile), holidays, resampling
  profiles.py      profile specs: BDEW ids, custom CSV profiles, energy-weighted mixes
  grids.py         pandapower grids, load-type detection
  assignment.py    category / profile assignment
  noise.py         AR(1)/white noise, size-dependent sigma, common component, per-load scaling and shift
  generator.py     P/Q generation, scenarios, statistics
  runs.py          run folder: CSVs, manifest with checksums, METHODS.md
  report.py        methods text from a manifest
  replay.py        regenerate + verify a run (CLI)
  timeseries.py    pandapower DFData/ConstControl helper and time-series power flow
  server.py        MCP server (tools + agent policy); Session = one user's state and limits
  agent.py         LLM tool-calling loop over MCP (Gemini, Groq, OpenAI, Anthropic, Ollama)
  cli.py           terminal chat
  webapp.py        multi-user sessions, uploads, cleanup, chat quotas (public app)
  data/slp_bdew.csv  all 16 BDEW profiles (built by scripts/build_bdew_table.py)
streamlit_app.py   public web app (form + chat); deployment: DEPLOY.md
agent_app.py       local single-user chat UI
tests/             73 tests incl. comparison with the R reference (runs if R is installed)
tests/live/        reliability battery against real LLM providers (needs API keys)
standardlastprofile/  vendored R package (CC0): source data and reference implementation
legacy/            earlier prototype, not used

Licence

MIT (see LICENSE). The BDEW profile data come from the CC0-licensed R package standardlastprofile.

Data sources

Limitations

  • BDEW profiles describe German customer groups; other countries only change the holiday calendar.

  • No industrial standard profile exists; supply a measured one with register_profile if you have it.

  • Noise parameters are defaults, not fitted values. Heat pumps and EVs are not represented.

  • Reactive power uses a constant power factor per load.

Available Tools

12 tools
assign_load_typesB

Assign a category (residential, commercial, industrial, agricultural, residential_pv, residential_pv_battery, mixed) and profile to every load. Precedence: mapping > grid metadata > distribution > default_category. Returns assignment_id and a per-category summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
basisNoShares refer to number of loads or to powercount
grid_idYes
mappingNoCategory per key, e.g. {'5': 'commercial'}; keys per mapping_key
strategyNosize: largest loads -> industrial, mixed, commercial, agricultural, residential; random: seededsize
mapping_keyNoload_index
distributionNoPercent per category for the remaining loads, e.g. {'residential': 60, 'commercial': 40}
include_loadsNoAlso return the per-load table (large)
load_profilesNoProfile per load/bus (keys per mapping_key), e.g. {'7': 'G5'}
default_categoryNoCategory for loads still unassigned
category_profilesNoProfile per category: BDEW id, custom name or mix, e.g. {'mixed': '0.6*H25+0.4*G25'}
use_grid_metadataNoUse the grid's own load types (CIGRE)
profile_generationNo2025

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses the resolution precedence (mapping > grid metadata > distribution > default_category) and the return shape (assignment_id plus per-category summary), but says nothing about permissions, whether assignments overwrite existing data, or reversibility for what is clearly a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the core action and category vocabulary, with the precedence rule second. No filler, though the parenthetical category list is somewhat long.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter mutation tool with no annotations and no output schema, the description covers the core purpose, precedence, and return value, but does not give enough guidance on the many optional knobs (strategy, basis, distribution, include_loads) to be fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 69%, so the schema already documents most parameters individually. The description adds relational value by tying mapping, grid metadata, distribution and default_category together via the precedence chain, but leaves several params (seed, strategy, basis, profile_generation, include_loads) unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Assign') and resource ('category ... and profile to every load'), and enumerates the concrete category values an agent can expect. It is clearly distinct from profile-generation siblings like generate_profiles, though it does not name an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the tool is for assigning load categories before downstream profile work. There is no explicit when-to-use statement, no exclusions, and no named alternative among the many siblings (generate_profiles, get_grid_loads, etc.).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_profilesB

Generate P/Q time series per load and year; saves CSVs and a manifest (settings, assumptions, checksums); returns per-year and ensemble statistics.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
yearsYesCalendar years
grid_idYes
scalingNogrid p_mw = annual peak or mean; or scale to annual energypeak
reactiveNogrid_ratio: q/p from the grid; power_factor: use power_factorsgrid_ratio
clip_sigmaNo
n_scenariosNoMonte Carlo scenarios, seeds seed..seed+n-1
noise_modelNoar1
noise_sigmaNorelative std at the reference size (0.1 = 10 %)
assignment_idYes
power_factorsNocos(phi) per category, inductive
ar1_phi_hourlyNo
resolution_minNo
growth_rate_pctNogrowth %/a from the first year
holiday_countryNoDE
load_scaling_stdNostd of constant per-load factor
max_time_shift_minNomax per-load shift [min]
noise_common_shareNoshare of noise variance common to all loads (0 = independent)
holiday_subdivisionNoGerman state code, e.g. BY
noise_sigma_scalingNosqrt_size: sigma*sqrt(P_ref/P_i); constant: same for all loadssqrt_size
noise_reference_p_mwNosize [MW] where noise_sigma applies; default median load
annual_energy_mwh_by_loadNoannual_energy: MWh per load index (overrides category)
annual_energy_mwh_by_categoryNoannual_energy: MWh per load, per category

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful side effects: it writes CSVs plus a manifest containing settings, assumptions, and checksums, and it returns per-year and ensemble statistics. However, it omits whether the run is idempotent, what auth/permissions are needed, runtime cost for a 23-parameter Monte Carlo call, and whether existing artifacts are overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that is front-loaded with the core action and packs outputs into a parenthetical. It wastes no words, though the semicolon-chained clauses demand careful parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since there is no output schema, the description's note that CSVs, a manifest, and per-year/ensemble statistics are returned is genuinely useful. But for a 23-parameter, no-annotation, write-producing tool, the absence of usage guidance, artifact-overwrite behavior, and precedence among the many scaling/noise options leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 65%, so the schema documents most parameters (scaling, reactive, noise_model, n_scenarios, etc.). The description adds only the 'per load and year' framing and says nothing about how the required grid_id, assignment_id, and years combine or how the 20 optional knobs interact, which is the baseline-3 level when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate P/Q time series per load and year,' which is unambiguous. It reasonably distinguishes the tool from siblings like register_profile, get_bdew_curve, assign_load_types, and plot_run, though it never names an alternative to force the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no comparison to alternatives such as register_profile or load_grid. The agent must infer from the name alone that this is the generation step; nothing in the text routes it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_bdew_curveC

One BDEW profile for a year (W per 1,000 kWh/a): CSV, plot and statistics.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearYes
profile_idYesH0, H25, G0-G6, G25, L0-L2, L25, P25 or S25
resolution_minNo
holiday_countryNoDE
holiday_subdivisionNoGerman state code, e.g. BY, BW

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the output forms (CSV, plot, statistics) and the normalization unit (W per 1,000 kWh/a), but says nothing about auth, rate limits, or what the returned series represents beyond the unit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though its terseness veers toward under-specification rather than true completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Five parameters, no annotations, and no output schema, yet the description documents neither the input parameters nor the structure of the returned CSV/statistics. It is too thin for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40%; profile_id and holiday_subdivision are documented in the schema, while year, resolution_min, and holiday_country are not. The description adds no parameter meaning at all, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: retrieving one BDEW profile for a given year, and notes the output forms (CSV, plot, statistics). The singular 'one ... profile' implicitly distinguishes it from list_bdew_profiles, but no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, prerequisites, or alternatives. An agent must infer from 'one' that this is the single-profile variant versus list_bdew_profiles, with nothing stated to confirm that routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_default_settingsC

Default settings with explanations.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It hints that the return includes explanations, but says nothing about what defaults are covered, whether any state is required, or the read-only nature beyond the implied 'get'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, but it is under-specified rather than genuinely concise; there is no front-loaded statement of purpose beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with no output schema, the description should at least indicate what is returned and in what form. 'Default settings with explanations' is too thin to tell an agent what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document. Per the rubric this is the baseline 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the name ('Default settings with explanations') without a clear verb or resource scope, and gives no way to distinguish it from siblings like list_grids or get_bdew_curve. An agent cannot tell what domain these defaults belong to.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to call this versus the many other getters in the sibling list, nor any prerequisites or conditions. Usage is only inferable from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_grid_loadsB

Page through the load table (index, name, bus, p_mw, q_mvar, detected type).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
grid_idYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the operation is paginated and lists the returned columns, which is genuinely useful, but says nothing about permissions, the meaning of the detected 'type' field, or what happens at page boundaries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the verb and returned fields front-loaded; no filler. It is perhaps too terse to be fully useful, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, naming the returned columns is the right move and partially compensates. However, the grid_id parameter, paging semantics, and the fact that 'detected type' is inferred rather than declared are left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters, so the description must compensate and largely does not. It lists output columns rather than explaining limit, offset, or grid_id, leaving the paging defaults (limit capped at 40) and the grid identifier semantics undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Page through the load table') and enumerates the returned columns, so an agent knows exactly what data comes back. It does not differentiate itself from siblings like load_grid, but the resource is distinctive enough to be unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as load_grid or get_grid_curve, and no prerequisites. The only usage hint is the word 'Page', which implicitly signals limit/offset paging but is never spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_runC

A run's results, methods text, a manifest section, or checksum verification.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
sectionNosummary: statistics and files; methods: methods text with references; verify: recompute checksumssummary

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It hints at read-style retrieval modes but never states that the operation is read-only, what happens for an invalid run_id, whether 'verify' performs expensive recomputation, or how output differs per section. Most behavioral traits remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no padding, which is good, but it is under-specified rather than truly concise — the terseness leaves the reader without a verb or any front-loaded statement of what the call returns by default.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, an enum largely undocumented, and a required run_id with no description, the definition is not complete enough for reliable invocation. It should at minimum say that it fetches one run by id, that section defaults to summary, and roughly what each section returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% and the enum has ten values, yet the schema documents only summary, methods, and verify. The description partially compensates by naming results, methods text, and checksum verification, but the remaining sections (assumptions, config, assignment, scenarios, resolved, holidays, all) are undocumented in both places, and run_id has no explanation at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates what get_run can return (results, methods text, a manifest section, checksum verification), which conveys the resource, but it is a noun phrase with no verb and never states that it retrieves a run by id. It also does not distinguish itself from siblings like plot_run or list_runs, so the purpose is inferable but not explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no routing to alternatives such as list_runs (to find a run) or plot_run (to visualize one). An agent must guess that this is the retrieval tool and that run_id must come from elsewhere.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_bdew_profilesA

List BDEW profiles, categories with default profiles, registered custom profiles and the mix syntax.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses the scope of content returned (categories, custom profiles, mix syntax), but says nothing about it being a read-only, side-effect-free operation, nor about result size, ordering, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the verb and resource front-loaded and no filler. The trailing enumeration of content is slightly dense but every clause earns its place by describing the returned data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description partially compensates by enumerating what the listing contains. For a zero-parameter read tool this is nearly sufficient, though a note on read-only behavior or result shape would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; the sentence correctly implies no filtering arguments are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (BDEW profiles), and enumerates what is covered: profile categories with default profiles, registered custom profiles, and the mix syntax. It is clear what the tool returns, though it does not explicitly distinguish itself from siblings like register_profile or get_bdew_curve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance is given, and no sibling is named as an alternative. Usage is only implied: this is an obviously safe discovery/browse call for profile data, which an agent can infer but not confirm from the text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_gridsA

List the built-in benchmark grids and whether their load types are known. Custom pandapower grids (.json/.xlsx files) can also be loaded with load_grid.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the shape of the result (built-in grids plus load-type status), but says nothing about ordering, permissions, or caching. For a simple zero-arg read-only enumeration the gaps are minor, but the disclosure is still thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what is listed and immediately followed by the alternative for custom grids. No filler, no restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must cover the return content, and it does specify the two pieces of information returned. It could say more about how the grid list is structured or empty-state behavior, but for a zero-argument enumeration it is nearly sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters and the schema is an empty object, so there is nothing for the description to disambiguate. Baseline 4 applies; the absence of a parameter-syntax discussion is appropriate rather than a gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the built-in benchmark grids') plus the returned attribute (whether their load types are known), and it distinguishes itself from load_grid, which handles custom pandapower files. An agent can tell what this returns without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence routes the agent to load_grid for custom .json/.xlsx grids, giving a clear alternative and the condition that selects it. It stops short of an explicit when-not statement, but the boundary between built-in and custom grids is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsA

List the runs of this session (newest first).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It does disclose ordering (newest first) and session scoping, which is real value, but says nothing about pagination, result limits, whether runs can be missing, or what a run entry contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence that front-loads the action and appends the two useful qualifiers (scope, ordering). Nothing is wasted and nothing extraneous is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-argument list tool with no output schema, the description covers the essential facts an agent needs to decide to call it. It stops just short of describing the shape of a returned run or any result-size limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify; the baseline for a no-parameter tool applies. Schema coverage is 100% trivially.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List the runs") and adds scope ("of this session"), which distinguishes it from a global run listing. It does not explicitly contrast with the sibling get_run, but list-vs-get is strongly implied by the verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The session scope and newest-first ordering imply when this is useful (browsing recent runs), but there is no explicit guidance on when to prefer it over get_run or plot_run. Usage must be inferred from the verb and from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_gridB

Load a grid: summary of loads, total power, voltage levels and whether load types are known.

ParametersJSON Schema
NameRequiredDescriptionDefault
gridYesBuilt-in grid id, or a pandapower .json/.xlsx file (uploaded file name)
include_loadsNoAlso return the per-load table (large)

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It discloses the returned summary fields, which is useful since there is no output schema, but it does not state whether the operation is read-only, whether it caches or mutates state, or any performance/auth considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the verb and returned content are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully names the summary contents, but it omits usage routing against sibling tools and any behavioral safety profile. For an agent needing to choose correctly among grid-related tools, these gaps matter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters (grid id/file and include_loads). The description adds no parameter semantics beyond the schema, warranting the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Load') and resource ('grid'), and enumerates the returned summary: loads, total power, voltage levels, and load-type status. It does not differentiate itself from siblings like get_grid_loads or list_grids, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus get_grid_loads, list_grids, or get_default_settings. The description only says what it does, leaving the agent to guess the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plot_runC

Plot a run (PNG): year, week, day, load-duration curve or daily means.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoweek
yearYes
loadsNoload indices for group_by=load
run_idYes
group_byNocategory
scenarioNoscenario index
start_dateNoYYYY-MM-DD for week/day (default: at the peak)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose that the return is a PNG image (useful, since there is no output schema), but says nothing about permissions, cost, whether the run must exist, how group_by/loads interact, or pagination/error behavior for an operation with seven parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb, resource, output format, and the option list — no wasted words. It is terse to the point of under-specification, but structurally efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a seven-parameter plotting tool with no annotations and no output schema, the description covers only the view options. It omits how required run_id/year behave, when group_by=load needs the loads array, and what start_date does beyond the schema's own hints, leaving real gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 43%, so the description partially compensates by mapping view values to human-readable labels (duration -> "load-duration curve", daily_mean -> "daily means"). But it says nothing about run_id, year, group_by, scenario, or how the required year interacts with optional start_date for week/day plots.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Plot") and resource ("a run") and immediately names the output format (PNG), which distinguishes it from the sibling read/list tools. It also enumerates the available views. However, it does not explicitly contrast itself with any sibling, so it stops short of the 5 bar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives among the many sibling tools. The only implied usage is that run_id and year are required, which the schema already enforces. An agent must infer which view to pick and when.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_profileA

Register a user-supplied profile: CSV with 1 day, 7 days (Mon-Sun) or a full year of values, any units (rescaled to 1,000 kWh/a). Use its name in category_profiles / load_profiles.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName to use in profile specs
csv_pathYesCSV (uploaded file name): one value column, optional timestamp column first
step_minNoInferred from timestamps if omitted
descriptionNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses important behavior: accepted CSV formats (1 day, 7 days, full year), unit rescaling to 1,000 kWh/a, and optional timestamp column. However, it doesn't mention validation rules, error handling, whether registration is persistent, or any rate limits. Given no annotations, this is partial but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and format details. No wasted words; the second sentence efficiently links to downstream usage. Could be slightly more structured but is clear and compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters, no output schema, and no annotations, the description covers the main input formats and integration point but omits behavioral details like validation, persistence, or error handling. It's adequate but leaves gaps that an agent might need to infer or discover via trial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema documents most parameters. The description adds some semantics: CSV format expectations for csv_path, and implied use of name. However, it doesn't explain step_min (inferred from timestamps if omitted) or the description parameter beyond what the schema provides. Baseline 3 is appropriate when schema does most of the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Register a user-supplied profile') and adds key detail about accepted CSV formats and unit rescaling. Distinguishes it from siblings like load_grid or list_bdew_profiles, which handle built-in/BDew data. However, it doesn't explicitly distinguish 'register' as the setup step versus generate_profiles as the usage step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly indicates the tool's purpose: adding custom profiles. Mentions how the registered name should be used ('Use its name in category_profiles / load_profiles'), which provides downstream context. But it doesn't state when NOT to use it (e.g., use built-in profiles instead) or compare with sibling alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv1.1.0
    • First observedassign_load_types
    • First observedgenerate_profiles
    • First observedget_bdew_curve
    • First observedget_default_settings
    • First observedget_grid_loads
    • First observedget_run
    • First observedlist_bdew_profiles
    • First observedlist_grids
    • First observedlist_runs
    • First observedload_grid
    • First observedplot_run
    • First observedregister_profile

TDQS

A3.5/5.0

Scored across 12 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: listing/loading grids, inspecting loads, managing BDEW profiles, assigning load types, generating/plotting/retrieving runs. The few closely related tools (list_grids vs load_grid, load_grid vs get_grid_loads) are differentiated by their descriptions (summary vs detailed table vs listing). No two tools appear to do the same thing.

Naming Consistency5/5

All 12 tools use snake_case with a consistent verb_noun pattern (list_*, load_*, get_*, register_*, assign_*, generate_*, plot_*). The verbs are appropriate and predictable, making the set easy to navigate.

Tool Count5/5

12 tools is well within the ideal 3–15 range for a specialized load-profile agent. Each tool covers a distinct step in the workflow, and there is no obvious redundancy or missing core operation that would inflate the count.

Completeness4/5

The surface covers the end-to-end workflow: list/load grids, inspect loads, manage BDEW profiles, assign types, generate profiles, and retrieve/plot runs. Minor gaps exist—no update/delete for registered profiles and no per-load assignment retrieval—but these are not likely to block typical agent tasks.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    87+ specialized tools for German and European energy data. Direct AI access to Marktstammdatenregister (MaStR), ENTSO-E, Redispatch 2.0, and Grid Operations for utilities and datacenters.
    2
    GPL 3.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables access to European electricity data including day-ahead prices, probabilistic forecasts, carbon intensity, and cheapest-window optimization for 43 bidding zones.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables querying European electricity generation, prices, and capacity data from Fraunhofer ISE's Energy-Charts platform.
    148 npm
    MIT