nempulse
Server Details
Australian NEM battery (BESS) data for AI agents: revenue, dispatch, FCAS, events. Free, no auth.
- Status
- Healthy
- Uptime
- 99.1% over 55 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 11 tools
The tool set is mostly distinct: each tool has a clear domain, such as events, regional prices, fleet summary, rankings, simulator, and ad hoc query. However, several tools (get_battery_detail, get_battery_optimal, get_battery_revenue) all provide battery-level revenue metrics with subtle differences, so an agent could misselect without reading the detailed caveats. Cross-references in descriptions mitigate this, but the overlap remains.
All tool names use snake_case with a consistent verb_noun pattern: get_*, list_*, query_*. Minor variations like get_rankings (noun without object) are still predictable and readable. No mixed conventions.
11 tools is well-scoped for a NEM battery analytics server, covering reference, detail, revenue, optimal dispatch, events, prices, rankings, simulator, and ad hoc query. Each tool appears to earn its place without redundant duplicates.
The surface covers the core analytics lifecycle: fleet listing, per-unit detail, revenue, optimal dispatch, event analysis, regional prices, rankings, and hypothetical simulation. Minor gaps exist for raw per-interval dispatch or SOC time series, which are only indirectly available via query_nem_data, but the main workflows are well supported.
Available Tools
11 toolsget_battery_detailBattery detailARead-onlyIdempotentInspect
Deep-dive metrics for one battery by DUID (e.g. HPR1 = Hornsdale): revenue, dispatch, SOC, FCAS. WARNING: rev_today, energy_rev_today, fcas_rev_today and contingency_fcas_rev_today are MONTH-TO-DATE by default, not daily (matching the rev_mtd keys in fcas_breakdown) — do not report them as 'today's revenue'. throughput_cycles, throughput_mwh, avg_dispatch_price, avg_charge_price and efficiency_pct cover the same window. Pass date_from and date_to (both required together) to scope this window explicitly, e.g. to a single day, instead of relying on the month-to-date default. For a true daily time series use get_battery_revenue. rte_pct is NOT window-scoped: it is the unit's latest measured round-trip efficiency (fitted from AEMO's reported energy storage against its dispatch over a trailing 30 days, refreshed weekly), and is null for units whose fit has not cleared its acceptance checks — null means 'not measured', never 'inefficient'. Also returns commercial_context (e.g. TOLLED, CONTRACTED) and commercial_note — ALWAYS check commercial_context before comparing this unit's revenue against another unit's: tolled/contracted units do not trade merchant and their spot figures are not comparable.
| Name | Required | Description | Default |
|---|---|---|---|
| duid | Yes | Battery DUID, e.g. HPR1. Use list_bess_units to find it from a station name. | |
| date_to | No | Optional end date YYYY-MM-DD (requires date_from too). | |
| date_from | No | Optional start date YYYY-MM-DD (scopes revenue/throughput stats; requires date_to too). Omit both for the month-to-date default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint/idempotentHint already declaring the safety profile, the description still adds substantial non-obvious behavior: rev_today/energy_rev_today/fcas_rev_today/contingency_fcas_rev_today are month-to-date, rte_pct is not window-scoped and null means 'not measured', and the date_from/date_to pair scopes the window. These are exactly the traps an agent would otherwise fall into.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but front-loaded and entirely load-bearing: purpose first, then the MTD warning, then window scoping, then the rte_pct caveat, then the commercial_context comparison rule. The all-caps emphasis is used sparingly on the two genuinely dangerous mistakes. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden, and it does: it names returned field groups, explains the MTD vs windowed semantics, defines null rte_pct, and describes commercial_context/commercial_note. Nothing an agent needs to call and interpret this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), but the description goes well beyond it: it explains the coupling constraint that date_from and date_to are 'both required together', states what the window actually scopes (revenue/throughput stats) versus what it does not (rte_pct), and clarifies the month-to-date default when omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (deep-dive metrics) and resource (one battery by DUID), enumerates the metric families returned (revenue, dispatch, SOC, FCAS), and explicitly distinguishes itself from the sibling get_battery_revenue for daily time series. An agent can tell it apart from the other 10 siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('for a true daily time series use get_battery_revenue'), how to resolve input ('use list_bess_units to find it from a station name'), and a when-not-to-be-misled rule ('ALWAYS check commercial_context before comparing this unit's revenue against another'). Alternatives and exclusions are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_battery_optimalActual vs optimal dispatch revenueARead-onlyIdempotentInspect
Actual vs LP-optimal dispatch revenue, per-DUID summary, over a date range (energy-only, perfect-foresight benchmark). NOT a revenue-total source — use get_battery_revenue for that. Both the 'actual' AND the 'optimal' figures here are MLF-adjusted (get_battery_revenue's is gross) — the LP's objective is solved on MLF-adjusted prices, not just settled at them afterward — and both cover solved LP days only (days where the solver failed are dropped from both), so the two tools' totals will not match even for the same DUID and date range. The requested date_to may also be silently truncated to the latest date with sufficient fleet-wide LP coverage. Pass duid to restrict to one battery — omitting it scans every DUID and can time out even on a ~3-week range; even a single-DUID, single-month scan has been observed to time out, so keep date ranges short and retry narrower on a timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| duid | No | Battery DUID, e.g. HPR1 (optional — omit for all DUIDs) | |
| date_to | Yes | End date, YYYY-MM-DD (inclusive). | |
| date_from | Yes | Start date, YYYY-MM-DD (inclusive). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive), but the description adds substantial behavior beyond them: MLF-adjusted vs gross figures, both totals restricted to solved LP days, the two tools' totals intentionally not matching, silent truncation of date_to, and observed timeout behavior even for single-DUID single-month scans.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose before the caveats, and nearly every clause carries non-obvious information. The heavy em-dash parentheticals make it dense and slightly harder to parse, but there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, 3-param tool with no output schema, the description covers purpose, exclusions, comparison caveats, and failure modes well. It is slightly thin on the actual return shape (per-DUID summary is named but not itemized), which is the only notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: omitting duid scans every DUID and can time out, while passing it restricts to one battery. Date parameters are only constrained generically ('keep date ranges short'), and the date_to truncation caveat is a useful semantic addition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and comparison (actual vs LP-optimal dispatch revenue) with a defined scope (per-DUID, date range, energy-only, perfect-foresight benchmark). It explicitly names the sibling it is NOT (get_battery_revenue) and why, so an agent can disambiguate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when not to use it ('NOT a revenue-total source — use get_battery_revenue') and when to pass duid (omit for fleet scan, but risks timeouts). It gives concrete operational guidance: keep date ranges short, retry narrower on timeout.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_battery_revenueBattery daily revenueARead-onlyIdempotentInspect
Daily gross-spot revenue by market (energy + FCAS) for one battery (DUID) over a date range. Daily grain only. This is the tool for total revenue questions — use get_battery_optimal only for the actual-vs-perfect-foresight benchmark, not as a revenue source (its 'actual' figure is MLF-adjusted and solved-days-only, so it will not match this tool's totals). Each day also carries energy_rev_mlf_adjusted (null if the LP backcast hasn't run for that day yet, not zero) alongside the gross energy_rev, so MLF-adjusted figures are available here too without switching tools.
| Name | Required | Description | Default |
|---|---|---|---|
| duid | Yes | Battery DUID, e.g. HPR1. Use list_bess_units to find it from a station name. | |
| date_to | Yes | End date, YYYY-MM-DD (inclusive). | |
| date_from | Yes | Start date, YYYY-MM-DD (inclusive). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent safety profile, but the description adds genuinely useful behavior beyond them: daily-grain-only restriction, the fact that optimal's actual figure is MLF-adjusted and solved-days-only so totals won't match, and that energy_rev_mlf_adjusted is null (not zero) until the LP backcast runs. That null-vs-zero distinction is exactly the kind of trap an agent would otherwise misread.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and grain, then routing guidance, then the MLF caveat. Dense but every sentence carries routing or semantics payload; the long parenthetical about get_battery_optimal is the only spot that could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so reasonably: names the revenue components (energy, FCAS), the daily grain, and the MLF-adjusted field with its null semantics. The full shape/naming of returned fields is not enumerated, but enough is given to call and interpret it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds the daily-grain constraint that shapes how date_from/date_to can be used, plus returned-field nuance tied to the duid/lookup flow. It does not add syntax or format detail beyond the schema, so it sits just above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: daily gross-spot revenue by market (energy + FCAS) for one battery DUID over a date range, with daily grain. It explicitly distinguishes itself from sibling get_battery_optimal, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly claims the revenue-question territory ('This is the tool for total revenue questions') and gives a when-not clause for get_battery_optimal, including why its 'actual' figure is not a valid revenue source. This is the when/when-not/alternatives pattern at full strength.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_event_detailPrice event detailARead-onlyIdempotentInspect
One price event by event_id (take the id from list_events), with the response of every battery that was online, ranked by est_revenue. Event fields are as in list_events. Each battery row has duid, station, region, response (DISCH = discharging, CHRG = charging, IDLE, OFFLINE = no availability), avg_mw (positive is discharge), peak_mw, est_revenue (energy only, sum of MW times price over the event, gross, not MLF-adjusted; excludes FCAS), soc_pct at start and soc_pct_end, availability_mw, fcas_mw (average FCAS commitment), plus is_commissioning, unit_class and peer_comparable. Check those labels before comparing units. Per-interval dispatch during the event is not exposed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Event id, as returned in event_id by list_events. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower, yet the description adds substantial context: est_revenue is gross, energy-only, excludes FCAS, and is not MLF-adjusted, and per-interval dispatch is deliberately not exposed. That is real behavioral disclosure beyond the annotations, though it doesn't discuss auth or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and scope are front-loaded in the first sentence, then the dense field glossary follows. It is long, but with no output schema each clause about return fields and their semantics earns its place; only the packed listing of unit_class/peer_comparable feels slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return-value burden, and it does: it enumerates every battery row field, defines the response code meanings, sign conventions, and caveats on revenue and FCAS. An agent has everything needed to interpret the response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single required id parameter, so the baseline is 3. The description goes further by telling the agent where the value comes from ('take the id from list_events'), adding sourcing semantics the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'One price event by event_id', immediately distinguishing it from the sibling list_events by being the single-event lookup. The scope (all online batteries, ranked by est_revenue) is also made explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear sourcing guidance — 'take the id from list_events' — which directs the agent to the prerequisite tool, and warns to check the field labels before comparing units. It stops short of stating when NOT to use this tool versus e.g. get_battery_detail, so it lacks full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fleet_summaryNEM battery fleet summaryARead-onlyIdempotentInspect
Fleet-wide snapshot for the NEM battery fleet: unit_count, active_unit_count, total_mw, total_mwh, fleet_rev_mtd, avg_efficiency_pct, last_interval and spot_prices (latest price per region, $/MWh). Covers month-to-date unless date_from and date_to are given. fleet_rev_mtd is gross spot revenue (energy + FCAS + FPP) over that window, whatever its length, despite the name. avg_efficiency_pct is dollar-weighted actual energy revenue as a share of the perfect-foresight optimum (a capture rate), null if the optimisation has not run for the window. Commissioning units are excluded from the efficiency figure unless include_commissioning is true. For one battery use get_battery_detail; for a ranking use get_rankings.
| Name | Required | Description | Default |
|---|---|---|---|
| date_to | No | Optional. End date YYYY-MM-DD (requires date_from). | |
| date_from | No | Optional. Start date YYYY-MM-DD (requires date_to). Omit both for month-to-date. | |
| include_commissioning | No | Include units still commissioning in the efficiency figure (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful non-obvious behavior: fleet_rev_mtd is gross spot revenue over any window despite its name, avg_efficiency_pct is a capture rate and returns null when optimisation hasn't run, and commissioning units are excluded unless include_commissioning is true. It stops short of describing caching or volume but adds substantial context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the return payload, then caveats, then sibling routing in a single dense paragraph. Every clause carries information, though the field enumeration is long enough to be slightly heavy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must and does explain the return fields and their edge cases (null efficiency, gross-revenue naming caveat). Combined with the date-window defaults and sibling routing, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains the month-to-date default, that fleet_rev_mtd spans whatever date_from/date_to window is set, and that include_commissioning toggles the efficiency figure specifically. This goes beyond the schema's terse field notes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (fleet-wide snapshot for the NEM battery fleet) and enumerates the exact metrics returned, so the agent knows precisely what this tool yields. It also names the two sibling tools it is not, distinguishing it from get_battery_detail and get_rankings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'For one battery use get_battery_detail; for a ranking use get_rankings.' The default window behavior ('covers month-to-date unless date_from and date_to are given') further clarifies when the tool applies without arguments.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rankingsBattery capture-rate rankingsARead-onlyIdempotentInspect
League table of merchant NEM batteries over a date range, ranked by capture rate by default (actual revenue as a share of the forecast-informed benchmark; capture_perfect_pct is the share of the perfect-foresight ceiling). Each row has rank, duid, station, region, capacity, duration_band, capture_fc_pct, capture_perfect_pct, rev_per_mw_day, rev_per_mwh_day, total_rev, cycles_per_day and movement against the previous period of equal length. Units still commissioning and CONTRACTED or TOLLED units are listed under excluded, not ranked, unless include_arrangement is true (their spot revenue is not their commercial return). Check unit_class and peer_comparable before reading a gap as performance. It scans the whole fleet twice, so keep the range to a month or less; it can be slow or time out on long ranges. For one unit's revenue use get_battery_revenue.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | NEM region code: NSW1, QLD1, SA1 (South Australia), TAS1 or VIC1. | |
| date_to | Yes | End date, YYYY-MM-DD (inclusive). | |
| sort_by | No | Ranking metric (default capture_fc). | |
| date_from | Yes | Start date, YYYY-MM-DD (inclusive). | |
| duration_band | No | Optional. Under 2 hours, 2 to 4 hours, or over 4 hours. | |
| include_arrangement | No | Rank CONTRACTED and TOLLED units too (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this read-only, idempotent, non-open-world, but the description adds material behavior beyond them: a double fleet scan that is slow and can time out, exclusion of commissioning/CONTRACTED/TOLLED units from ranking (with the commercial-return rationale), movement against an equal-length prior period, and the caveat to check unit_class and peer_comparable before treating a gap as performance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and metric definition, then layers caveats, exclusions and the sibling routing in a logical order. It is dense and sentence-heavy, but virtually every clause carries actionable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return contract itself, naming every row field and the movement column, and it covers exclusion semantics, performance caveats, and runtime limits. An agent has everything needed to call and interpret this tool correctly without opening anything else.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further: it explains what capture_fc_pct vs capture_perfect_pct actually mean, and clarifies that include_arrangement changes whether CONTRACTED/TOLLED units are ranked at all. It does not restate enum/format details already in the schema, so the added meaning is real but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('League table of merchant NEM batteries over a date range, ranked by capture rate by default') and defines the two headline metrics precisely (capture_fc_pct as revenue share of forecast-informed benchmark, capture_perfect_pct as share of the perfect-foresight ceiling). It also enumerates the row shape, so the agent knows exactly what it gets back and how it differs from the unit-level sibling get_battery_revenue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'For one unit's revenue use get_battery_revenue,' which disambiguates against the closest sibling. It also states the practical operating constraint (keep the range to a month or less; can time out on long ranges) and the exclusion condition plus its override (include_arrangement), covering when-not as well as when.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_regional_pricesRegional daily price summaryARead-onlyIdempotentInspect
Daily minimum, maximum and mean regional reference price (RRP, $/MWh) for one NEM region over a date range. Daily summary only: the raw 5-minute price series is not exposed (it is published by AEMO on NEMWEB). Use it for price context around revenue or events. For individual price spikes use list_events.
| Name | Required | Description | Default |
|---|---|---|---|
| region | Yes | NEM region code: NSW1, QLD1, SA1 (South Australia), TAS1 or VIC1. | |
| date_to | Yes | End date, YYYY-MM-DD (inclusive). | |
| date_from | Yes | Start date, YYYY-MM-DD (inclusive). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description adds meaningful behavioral context beyond them: it is a daily summary only and the raw 5-minute series is deliberately not exposed. This granularity limitation is not derivable from the annotations or schema and helps the agent avoid misusing the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose before the granularity caveat and the sibling redirect. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only, 3-param tool, the description covers what is returned (daily min/max/mean), the unit, the scoping constraint, and the data-granularity limit. With annotations covering the safety profile and no output schema needed, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with all three parameters described (region enum, inclusive date range), so the schema does the heavy lifting. The description adds the $/MWh unit and 'one region' scoping but no syntax details beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (daily min/max/mean RRP for one NEM region over a date range), including units ($/MWh) and the exact metrics returned. An agent can distinguish this from sibling tools like get_battery_revenue or get_rankings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it 'for price context around revenue or events' and routes the spike-detection case to list_events, which is a clear alternative. It gives when-to-use context and one concrete exclusion, though it doesn't cover other possible siblings (e.g., query_nem_data).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_simulator_revenueHypothetical battery revenueARead-onlyIdempotentInspect
Modelled revenue for a hypothetical 1 MW battery of a given duration in one region (not a real unit), from NEMPulse's optimal-dispatch model on real prices. Returns P10/P50/P90 revenue for three tiers: realistic (forecast-informed result scaled by the capture ratio real batteries achieve, the only tier to use as a project case), forecast_informed (idealised execution on AEMO's public price forecast) and perfect (an unreachable ceiling). Also energy_share and fcas_share, a monthly series, and the real-unit cohorts behind the FCAS cap and the capture ratio. Gross spot revenue only, before costs, degradation and financing. It is a screening estimate, not a forecast of future revenue.
| Name | Required | Description | Default |
|---|---|---|---|
| region | Yes | NEM region code: NSW1, QLD1, SA1 (South Australia), TAS1 or VIC1. | |
| duration | Yes | Battery duration in hours: 1, 2 or 4. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read-only idempotent profile, and the description adds substantial context beyond them: the gross-revenue boundary (before costs, degradation, financing), the tier definitions and their reliability, the limitations of the FCAS cap and capture-ratio cohorts, and the screening (not forecasting) nature. This is exactly the extra behavioral disclosure an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: what it is, whose data, then the return shape, then caveats. Every sentence carries information, though the tier explanation is somewhat compressed and the block could be broken into shorter units for scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of explaining returns, and it does: tiers, P10/P50/P90, energy_share/fcas_share, monthly series, and supporting cohort data. An agent know exactly what comes back and how much to trust it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with enums on both parameters, so region codes and duration values are already fully documented in the schema. The description adds only a light restatement ('of a given duration in one region') without extra syntax or constraint detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (modelled revenue for a hypothetical 1 MW battery) and immediately scopes it as 'not a real unit', which distinguishes it from the sibling get_battery_revenue. The modelled/optimal-dispatch framing is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong guidance on the returned tiers ('the only tier to use as a project case', 'an unreachable ceiling') and when not to rely on the result ('a screening estimate, not a forecast'). It stops short of explicitly naming the sibling to use for real units, but the 'not a real unit' framing routes the agent away from get_battery_revenue clearly enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_bess_unitsList NEM battery unitsARead-onlyIdempotentInspect
Reference list of every NEM-registered grid-scale battery. Call this first to turn a station name into the DUID the other tools need. Returns the whole fleet in one response (no filters): DUID, station, region, MW/MWh, MLF, coordinates, is_commissioning, commercial_context, unit_class, peer_comparable, rev_per_mw_yr. Caveats on the fields follow. rev_per_mw_yr is trailing 365-day energy + FCAS revenue (NOT FPP), gross (no MLF), divided by Max Cap MW, then annualised over the days the unit actually had dispatch data — not over a fixed 365-day denominator. This is a rough simulator guide, NOT a performance ranking. It is distorted for any unit with is_commissioning=true or commissioned within the last 365 days, because that span-annualisation extrapolates a few months of ramp-up behaviour out to a full year (scale-up factors of 1.5x-2.4x are live in the current data), which magnifies both weak and negative figures rather than diluting them — do NOT try to 'correct' it by rescaling to the unit's operating span, as that double-counts the annualisation. It is also inflated for small FCAS-primary units, and structurally biased against longer-duration units (shorter-duration units can concentrate power into the highest-price intervals). Do not use it to compare units, rank performance, or answer 'which battery earns most' — use get_battery_revenue over a matched window, or get_battery_optimal for capture. Before comparing any two units' revenue, check BOTH labels on each. commercial_context (TOLLED/CONTRACTED) means the unit does not trade merchant, so its spot revenue is not comparable; a null here means no publicly documented arrangement was found, NOT that the unit is confirmed merchant — only an explicit MERCHANT value means that, and most units are null. unit_class is the structural label (STANDALONE/HYBRID/NETWORK-SUPPORT/MICRO): peer_comparable is false for the latter three, whose dispatch answers to a co-located generator, a non-market obligation, or sub-10MW FCAS granularity rather than to price. Never pool a peer_comparable=false unit into a cross-unit statistic or ranking. Null until the first background refresh completes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive), and the description goes well beyond them: it discloses that the whole fleet returns in one response with no filters, enumerates returned fields, and warns about distortion behavior (span-annualisation inflating commissioning units 1.5x-2.4x, FCAS-primary inflation, duration bias), null semantics for commercial_context, and peer_comparable pooling constraints. This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: purpose, then the DUID/usage directive, then the field list, then caveats. However it is very long and some caveat passages (the double-counting/anti-rescaling warning) are dense and repetitive; it could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return-value burden, and it does: it lists every returned field and explains the two most misusable ones (rev_per_mw_yr, commercial_context/peer_comparable) in depth. An agent has everything needed to call and correctly interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description contributes field-level semantics (DUID, commercial_context, unit_class, rev_per_mw_yr), but those are return-value concerns rather than parameter meaning, so it does not push above the zero-parameter baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource and scope: 'Reference list of every NEM-registered grid-scale battery.' It also distinguishes itself from siblings by positioning itself as the resolution step ('turn a station name into the DUID the other tools need'), so an agent can tell clearly what this tool is for versus get_battery_detail or get_rankings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('Call this first to turn a station name into the DUID') and when NOT to use it, naming alternatives: 'use get_battery_revenue over a matched window, or get_battery_optimal for capture' and 'Do not use it to compare units, rank performance.' This is textbook alternative-routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_eventsList NEM price eventsARead-onlyIdempotentInspect
List NEM spot-price events, newest first. An event is a run of 5-minute intervals in one region where the price stays above $500/MWh or below -$200/MWh (short gaps are bridged), typed by its peak price: negative (below -$200), elevated (peak $500 to $1,000), spike ($1,000 to $5,000) or extreme ($5,000 and above). Each event has event_id, region, event_type, started_at, ended_at, duration_min, peak_rrp, avg_rrp and is_open (still ongoing). Filter by region, event_type and a date range. By default only event summaries are returned; set include_responses to true to also get how each battery behaved during each event, in which case keep limit small (response rows make the payload large and it is truncated at about 100 KB). Use get_event_detail for one event's battery responses.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max events to return (default 10, max 500). | |
| offset | No | Skip this many events, for paging (default 0). | |
| region | No | NEM region code: NSW1, QLD1, SA1 (South Australia), TAS1 or VIC1. | |
| date_to | No | Optional. Only events starting on or before this date, YYYY-MM-DD. | |
| date_from | No | Optional. Only events starting on or after this date, YYYY-MM-DD. | |
| event_type | No | Optional event type filter. | |
| include_responses | No | Also return per-battery responses for each event (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the description adds orthogonal value: default ordering (newest first), the ~100 KB truncation behavior when include_responses is true, and the meaning of is_open. That is exactly the payload-risk context an agent needs before choosing include_responses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the definition, then filtering, then the include_responses caveat and the sibling hand-off in a logical order. It is dense and somewhat long due to the event-type taxonomy, but each clause carries information an agent would otherwise have to guess.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description enumerates the returned fields (event_id, region, event_type, started_at, ended_at, duration_min, peak_rrp, avg_rrp, is_open) and warns about payload size, so nothing an agent needs to call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond it by defining the numeric price bands behind each event_type enum value and stating that filters combine (region, event_type, date range), which the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list NEM spot-price events) and immediately defines what an 'event' is, including the threshold logic and the four type buckets. This distinguishes it from siblings like query_nem_data and get_regional_prices without requiring schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: set include_responses only when battery behavior is needed, keep limit small in that case, and use get_event_detail for a single event's responses. It names the alternative tool and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_nem_dataAsk a question of the NEM dataARead-onlyIdempotentInspect
Last resort for ad hoc questions the other tools cannot answer. Takes a plain-English question, generates SQL over NEMPulse's tables, and returns the SQL, the result rows and a short explanation. It is the slowest tool (can take close to a minute) and it is AI-generated, so check the returned SQL before relying on a number. Scope each question to about one region-month or less: wider aggregates (a full year by region, or per-day top-N across all regions) risk exceeding the 8 second SQL limit, and a longer wait does not help. For per-day top-N or bottom-N questions, ask for a window function (ROW_NUMBER or RANK) rather than a per-day subquery, which has been seen to silently return all-null rows. Queryable data: dispatch prices, daily revenue, optimal dispatch, battery price profiles and market events. The market analysis tables (monthly spreads, regressions, correlations) are not reachable here.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | Plain-English question, max 500 characters. Name the region and date range, e.g. 'Which SA1 battery earned the most FCAS revenue in August 2026?' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly/idempotent safety, but the description adds the costly behavioral facts: slowest tool (~1 minute), AI-generated so the SQL should be verified, an 8-second SQL execution limit, and a known failure mode where per-day subqueries silently return all-null rows. That is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the identity ('last resort'), then proceeds to cost, scope limits, a specific failure-mode warning, and finally the queryable data boundary. Every sentence carries operational information; nothing is redundant with the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating the return payload (SQL, result rows, short explanation). Combined with the data-scope list and the query-shaping warnings, an agent has everything needed to call this correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the single 'question' param with an example, so the baseline is 3. The description then adds a real constraint on the parameter's content — keep scope to roughly one region-month because wider aggregates blow the 8-second limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource chain: takes a plain-English question, generates SQL over NEMPulse tables, and returns SQL, rows and an explanation. The opening clause 'Last resort for ad hoc questions the other tools cannot answer' explicitly differentiates it from all ten sibling getters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('other tools cannot answer'), when-not (the market analysis tables are unreachable), and alternative-avoidance guidance ('For per-day top-N ... ask for a window function rather than a per-day subquery'). Performance and scope conditions are also given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
- Changed
get_battery_detail2 fields changed- added
Input schema / properties / date_from / formatAdded value: +"date" - changed
Input schema / properties / duid / descriptionPrevious value: -"Battery DUID, e.g. HPR1"New value: +"Battery DUID, e.g. HPR1. Use list_bess_units to find it from a station name."
- Changed
get_battery_optimal4 fields changed- changed
Input schema / properties / date_from / descriptionPrevious value: -"Start date YYYY-MM-DD"New value: +"Start date, YYYY-MM-DD (inclusive)." - added
Input schema / properties / date_from / formatAdded value: +"date" - changed
Input schema / properties / date_to / descriptionPrevious value: -"End date YYYY-MM-DD"New value: +"End date, YYYY-MM-DD (inclusive)." - added
Input schema / properties / date_to / formatAdded value: +"date"
- Changed
get_battery_revenue5 fields changed- changed
Input schema / properties / date_from / descriptionPrevious value: -"Start date YYYY-MM-DD"New value: +"Start date, YYYY-MM-DD (inclusive)." - added
Input schema / properties / date_from / formatAdded value: +"date" - changed
Input schema / properties / date_to / descriptionPrevious value: -"End date YYYY-MM-DD"New value: +"End date, YYYY-MM-DD (inclusive)." - added
Input schema / properties / date_to / formatAdded value: +"date" - changed
Input schema / properties / duid / descriptionPrevious value: -"Battery DUID, e.g. HPR1"New value: +"Battery DUID, e.g. HPR1. Use list_bess_units to find it from a station name."
- Changed
get_event_detail1 field changed- changed
Input schema / properties / id / descriptionPrevious value: -"Event id"New value: +"Event id, as returned in event_id by list_events."
- Changed
get_fleet_summary3 fields changed- added
Input schema / properties / date_fromAdded value: +{ + "description": "Optional. Start date YYYY-MM-DD (requires date_to). Omit both for month-to-date.", + "format": "date", + "type": "string" +} - added
Input schema / properties / date_toAdded value: +{ + "description": "Optional. End date YYYY-MM-DD (requires date_from).", + "format": "date", + "type": "string" +} - added
Input schema / properties / include_commissioningAdded value: +{ + "description": "Include units still commissioning in the efficiency figure (default false).", + "type": "boolean" +}
- Added
get_rankings - Added
get_regional_prices - Added
get_simulator_revenue - Changed
list_events10 fields changed- added
Input schema / properties / date_fromAdded value: +{ + "description": "Optional. Only events starting on or after this date, YYYY-MM-DD.", + "format": "date", + "type": "string" +} - added
Input schema / properties / date_toAdded value: +{ + "description": "Optional. Only events starting on or before this date, YYYY-MM-DD.", + "format": "date", + "type": "string" +} - added
Input schema / properties / event_typeAdded value: +{ + "description": "Optional event type filter.", + "enum": [ + "negative", + "elevated", + "spike", + "extreme" + ], + "type": "string" +} - added
Input schema / properties / include_responsesAdded value: +{ + "description": "Also return per-battery responses for each event (default false).", + "type": "boolean" +} - changed
Input schema / properties / limit / descriptionPrevious value: -"Max events to return (default 50, max 500)"New value: +"Max events to return (default 10, max 500)." - added
Input schema / properties / limit / maximumAdded value: +500 - added
Input schema / properties / limit / minimumAdded value: +1 - added
Input schema / properties / offsetAdded value: +{ + "description": "Skip this many events, for paging (default 0).", + "minimum": 0, + "type": "integer" +} - changed
Input schema / properties / region / descriptionPrevious value: -"NEM region, e.g. SA1 (optional)"New value: +"NEM region code: NSW1, QLD1, SA1 (South Australia), TAS1 or VIC1." - added
Input schema / properties / region / enumAdded value: +[ + "NSW1", + "QLD1", + "SA1", + "TAS1", + "VIC1" +]
- Changed
query_nem_data2 fields changed- changed
Input schema / properties / question / descriptionPrevious value: -"Plain-English question (max 500 chars)."New value: +"Plain-English question, max 500 characters. Name the region and date range, e.g. 'Which SA1 battery earned the most FCAS revenue in August 2026?'" - added
Input schema / properties / question / maxLengthAdded value: +500
1 tool update
- Changed
get_battery_detail2 fields changed- added
Input schema / properties / date_fromAdded value: +{ + "description": "Optional start date YYYY-MM-DD (scopes revenue/throughput stats; requires date_to too). Omit both for the month-to-date default.", + "type": "string" +} - added
Input schema / properties / date_toAdded value: +{ + "description": "Optional end date YYYY-MM-DD (requires date_from too).", + "type": "string" +}
1 tool update
- Changed
get_battery_optimal1 field changed- added
Input schema / properties / duidAdded value: +{ + "description": "Battery DUID, e.g. HPR1 (optional — omit for all DUIDs)", + "type": "string" +}
8 tool updates
- First observed
get_battery_detail - First observed
get_battery_optimal - First observed
get_battery_revenue - First observed
get_event_detail - First observed
get_fleet_summary - First observed
list_bess_units - First observed
list_events - First observed
query_nem_data
Related MCP Connectors
Query Australia's electricity market (NEM/AEMO): prices, generation, FCAS, interconnectors, bids.
Live US power market prices, load, generation, weather and permits for AI agents.
Real-time electricity price signals for AI agents. Spot prices, cheapest hours, and contract recommendations. 31 countries across Europe and Oceania. No authentication required.
Real-time electricity prices for AI agents. 40+ countries, 100+ zones. No auth required.
Related MCP Servers
- AlicenseAqualityBmaintenanceOne-call Australian energy-market plumbing via AEMO — cited, structured responses for market data and analysis, not a data broker.556 PyPIMIT
- AlicenseNot gradedqualityAmaintenance87+ specialized tools for German and European energy data. Direct AI access to Marktstammdatenregister (MaStR), ENTSO-E, Redispatch 2.0, and Grid Operations for utilities and datacenters.2GPL 3.0
- AlicenseBqualityDmaintenanceConnects AI agents to energy infrastructure with 30+ tools for managing sites, assets, dispatch, settlements, compliance, and carbon tracking.3438 npm1MIT
- AlicenseNot gradedqualityCmaintenanceAccess verified historical market data with quality flags, funding rates, and more, supporting micropayments for AI agents and trading bots.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.