openmep
OpenMEP
Open MEP building state that a language model can reason about and edit correctly — grounded in deterministic loads and duct sizing, graded by a deterministic checker.
The bet: the useful unit of AI-assisted HVAC design is not a Revit plugin but a small, inspectable data model (rooms, loads, terminals, ducts, connectivity, constraints) that deterministic engines compute over and agents edit. This repo is that model, the engines, and the harness that measures whether an agent can edit it without breaking it.
What is here
Workspace | What it does |
| Dependency-free equal-friction round-duct sizing, supply/return/exhaust network CFM propagation, port-coincidence connectivity, ASHRAE equivalent diameters, and actual-vs-recommended size grading. |
| MCP server ( |
| Sizes and checks HVAC in Pascal Editor scenes, live through Pascal's MCP ( |
| Agent skill (Pascal skills.sh format) that drives the adapter from an MCP-connected agent: gather CFM, set |
| Synthetic T1–T6 mechanical-plan editing benchmark (single + branching fixture sets), deterministic grader, |
| IFC ingestion (web-ifc): storeys, spaces, elements, space boundaries, geometric envelope recovery; 2D plan extraction (Clipper2 footprints); design-day cooling/heating loads (ASHRAE 62.1 ventilation, 90.1 assemblies); mechanical plan layers with CFM-grounded terminals and inferred duct connectivity; plan validator; headless eval runner. |
| Next.js 16 viewer: 3D (That Open / Three.js) and storey-driven 2D plan mode with the load dashboard. |
|
|
Real-model tests and evals run on public IFC fixtures (buildingSMART
Duplex and WBDG Office ARCH + MEP pairs, CC BY 4.0) fetched by
fixtures/fetch-public-ifc.sh; see fixtures/README.md for provenance and the
exporter quirks the pipeline handles. No IFC is committed to this repository.
Related MCP server: Vensim MCP
Quick start
Give an agent a duct-sizing tool
claude mcp add openmep -- npx -y -p @openmep/mcp openmep-mcpAny MCP client works ({"command": "npx", "args": ["-y", "-p", "@openmep/mcp", "openmep-mcp"]}).
Describe the system in plain language, with the CFM each terminal needs, and the
agent calls the engine; see packages/mcp/README.md.
Size a duct network from JSON (no editor required)
npx -p @openmep/hvac-domain openmep-hvac size packages/hvac-domain/examples/furnace-two-registers.items.json --pretty
npx -p @openmep/hvac-domain openmep-hvac duct --cfm 250 --role main --diameter-in 6From a checkout: npm ci && npm run build --workspace @openmep/hvac-domain, then
node packages/hvac-domain/bin/openmep-hvac.js ....
The network is a list of equipment, fittings, terminals with engineer-supplied
requiredCfm, and segments linked by connectedItemRefs. The result lists each
run's system, CFM, role, recommended standard round diameter, velocity and friction
rate, grades any existing section as ok/undersized/oversized, and reports
terminals with no airflow or no path to equipment. Any host that can produce this
document gets the same engine; openmep-pascal items produces it from a Pascal
scene.
Try airflow design
Change a room's supply airflow, preview the connected duct sizes on the plan, apply or undo the change, and export the edited plan JSON.
npm ci
npm run dev:design
# Open http://localhost:3001/designRequires Node.js 24+. The first launch fetches the public Office IFC pair and
prepares Level 1; no database, .env, or model credentials are needed. Pass
-- --refresh to rebuild the example, or -- --port 3002 to use another port.
Edits are session-local manual airflow scenarios with round-duct proposals;
export saves the applied plan JSON, not a rewritten IFC. The public example
labels its inferred connections and assumed initial demands. Existing model
viewers also link to Airflow design for their selected storey.
Full viewer and IFC pipeline
npm ci
fixtures/fetch-public-ifc.sh # duplex + wbdg_office (~55 MB)
npm run build --workspace @openmep/hvac-domain --workspace @openmep/agent-benchmark
npm testMaterialize an editable HVAC plan from the public office model:
npm run eval:materialize -- \
--arch fixtures/public/wbdg_office/arc.ifc \
--mech fixtures/public/wbdg_office/mep.ifc \
--model-id wbdg_office
# eval-runs/wbdg_office/Level_1/{architecture,loads,mechanicalVisual2D,mechanicalEdit2D,baseline}.jsonRun the synthetic agent benchmark: npm run benchmark:verify, or see
packages/agent-benchmark/README.md for grading your own candidates.
Web app
Copy
.env.exampleto.env; start PostgreSQL and create the database.npm run db:initnpm run dev:web— seeds the Duplex pair as a stable dev fixture at/models/00000000-0000-4000-8000-000000000001?mode=plan(MEP_SKIP_DEV_FIXTURE=1to skip;IFC_FIXTURE_DIRto relocate fixtures).npm run dev:workerin a second terminal.
Status and limits
Loads are design-day, steady-state (no RTS/CTS lag, latent, or schedules) and not yet validated against a reference case; envelope U-values come from ASHRAE 90.1 prescriptive tables, climate from 15 city anchors. Treat outputs as preliminary sizing inputs, not a sealed design.
Roof loads are zero for Revit exports whose ceilings are room-bounding (all seen so far); plenum inference is open work.
Duct connectivity for exports without
IfcRelConnectsPortsis inferred geometrically and flagged as such on every item.
License
MIT — see LICENSE. Fixture data keeps its upstream licenses (fixtures/README.md).
Release and provenance rules for the npm packages: docs/public-artifact-release.md.
Available Tools
2 toolssize_ductSize one ductARead-onlyIdempotent
Recommend a standard round diameter for one duct run carrying a given CFM, with velocity and friction at that size. Optionally grade an existing round, rect, or oval section as ok, undersized, or oversized.
| Name | Required | Description | Default |
|---|---|---|---|
| cfm | Yes | Airflow the run carries, CFM. | |
| role | No | main (at equipment), branch (serves several terminals), or runout (one terminal). Default main. | |
| existing | No | Existing section to grade: { diameterIn } for round, or { shape: rect|oval, widthIn, heightIn }. | |
| airflowType | No | System, which sets the velocity cap. Default supply. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description still adds real value by disclosing output content — recommended diameter, velocity, friction, and the ok/undersized/oversized grade — which is not available elsewhere since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with no filler. The primary recommend behavior is front-loaded, and the optional grading capability follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining what comes back, and it does so (diameter, velocity, friction, grade verdict). Defaults for role and airflowType live in the schema, so nothing critical is missing, though the sizing standard/basis is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, including the role and airflowType enums and the existing-section shapes. The description adds the grading vocabulary (ok/undersized/oversized) but no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (recommend/grade) and resource (one duct run) with clear scope. The phrase 'one duct run' implicitly contrasts with the sibling size_duct_network, but the sibling is never named, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description signals two modes of use: recommend a size from CFM, or grade an existing section via the 'existing' input. This implies when the grading path applies, but there is no explicit when-not guidance or named alternative (e.g., use size_duct_network for a whole system).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
size_duct_networkSize a duct networkARead-onlyIdempotent
Size every supply, return, exhaust, and outside-air segment of a connected duct network. Each terminal's requiredCfm is carried along its shortest same-system path to equipment and summed per segment; roles (main/branch/runout) are inferred. Returns per-segment recommendations, grades for segments with an existing section, and findings: missing-required-cfm, no-equipment-path, dangling-reference, mixed-system-segment, undersized, oversized.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Every terminal, segment, fitting, and equipment item in the network. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, so safety is covered; the description adds real behavioral detail beyond that — the shortest same-system path algorithm, per-segment summation, inferred roles, and the exact finding codes emitted. It stops short of stating assumptions (e.g. what happens when no equipment path exists beyond naming the finding).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the verb and scope lead, followed by the algorithm and then the return contents. Every clause carries information, though the findings list is a long run-on that could be broken up.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned recommendations, grades, and finding codes, which is exactly what an agent needs to interpret results. The main missing piece is any statement of limitations or how it differs from size_duct.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes further by explaining how item fields are consumed — requiredCfm carried along paths and summed per segment, existing sections graded against the recommendation, and roles inferred. That adds operational meaning not present in the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Size every supply, return, exhaust, and outside-air segment of a connected duct network'), which is far more informative than the title alone. It does not, however, distinguish itself from the sibling size_duct tool, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Scope implies usage (a connected network with terminals/equipment that must be sized together), but there is no explicit when-to-use guidance and no mention of the sibling size_duct or when that alternative would be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
size_duct - First observed
size_duct_network
TDQS
Scored across 2 tools
The two tools have related but distinguishable scopes: size_duct handles one run (with optional grading), while size_duct_network sizes a whole connected system. There is mild overlap since both can grade existing sections, but the single-run vs. network boundary is clear enough for correct selection.
Both names use a consistent snake_case noun_verb pattern (size_duct, size_duct_network), with the network variant being a clean, predictable extension. No mixing of conventions.
Two tools is thin for a duct-design server; the single-run case is arguably a degenerate case of the network tool, so it borders on a minimal surface. The network tool is substantial, but the overall set feels sparse.
Coverage of the core duct-sizing workflow (single run plus full multi-system network with grading and findings) is solid for a focused calc tool. Minor gaps exist, e.g. no standalone grading without sizing or explicit design-parameter/friction-rate configuration, but no major dead ends.
Maintenance
Related MCP Connectors
Heating, ventilation and energy engineering calculations and diagnostics for AI agents.
Design, solve and simulate HVAC systems from real components, weather years and buildings.
Construction takeoff and estimating for AI agents. Measure a drawing PDF, export a priced estimate.
Solar, weatherization, EV charging, battery and heat-pump decision tools for AI agents.
Related MCP Servers
FlicenseNot gradedqualityCmaintenanceProvides deterministic, standards-based calculations for data center critical power infrastructure. Enables site selection, generator sizing, UPS sizing, NFPA 110 compliance, and more via 50+ AI agents and 8 compound chains.-- AlicenseAqualityCmaintenanceConverts structured system dynamics specifications into Vensim .mdl files with layout, SVG preview, static audit, and native Vensim integration. It can be used as an MCP server or from the command line.7MIT
- AlicenseAqualityCmaintenanceEnables AI assistants to load, simulate, inspect, and edit EPANET hydraulic/water-quality models locally via stdio.33MIT
- FlicenseNot gradedqualityAmaintenanceFloor plans, elevations and sections for agents. Draw a room in real millimetres, place doors, windows and tables, validate with typed errors, render deterministic SVG. Free key, no approval.-