CityDPC-MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CityDPC-MCP@CityDPC-MCP list the available CityJSON files"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CityDPC-MCP
A Model Context Protocol (MCP) server that lets LLM agents such as Claude Code or Codex inspect, analyse and edit CityJSON and CityGML building models. It is built on CityDPC, the Python library for 3D city models maintained by RWTH-E3D.
Companion repository to the paper CityDPC-MCP: Enabling 3D city model processing via Model Context Protocol. It contains the server, the benchmark used in the paper and the per-run results.
Installation
You need uv. It fetches Python 3.12+ and all dependencies, CityDPC included.
uv tool install git+https://github.com/HamerCode/CityDPC-MCPThis puts a citydpc-mcp command on your PATH. The server communicates over stdio and works with the *.json / *.gml files in one dataset directory, which you set with --dataset-dir (default: the current directory).
Note:
save_datasetoverwrites files in that directory, so work on copies of important data.
Claude Code
claude mcp add citydpc -- citydpc-mcp --dataset-dir /absolute/path/to/your/data
claude mcp list # should show: citydpc: citydpc-mcp ... - ✔ ConnectedAdd --scope user to make the server available in all projects. If you leave out --dataset-dir, the server uses the project directory Claude Code was started in.
Codex
codex mcp add citydpc -- citydpc-mcp --dataset-dir /absolute/path/to/your/dataOr add it to ~/.codex/config.toml yourself:
[mcp_servers.citydpc]
command = "citydpc-mcp"
args = ["--dataset-dir", "/absolute/path/to/your/data"]Other MCP clients (Claude Desktop, Cursor, …)
{
"mcpServers": {
"citydpc": {
"command": "citydpc-mcp",
"args": ["--dataset-dir", "/absolute/path/to/your/data"]
}
}
}Desktop apps often don't search ~/.local/bin. If yours can't find the command, use the absolute path that which citydpc-mcp prints. To run the server without installing it, use uvx --from git+https://github.com/HamerCode/CityDPC-MCP citydpc-mcp as the command. The first start takes a while because uv installs the dependencies then.
Try it
Point the server at the sample dataset (evaluation/data/evaluation.city.json, 36 LoD2 buildings) and ask something like:
Load evaluation.city.json. Which building has the highest measuredHeight, and how many party walls does the dataset have?
Related MCP server: ifc-geometry-mcp
Tools
Group | Tools |
Datasets |
|
Analysis |
|
Editing |
|
Versioning |
|
All edits happen in memory until save_dataset is called. The paper evaluated the other 18 tools; create_dataset was added afterwards for quick testing. Their names and (German) descriptions are unchanged, including the historical spelling get_buiding_Id_list. See src/citydpc_mcp/server.py for signatures.
Evaluation
The benchmark asks an LLM to solve five tasks on the 36-building sample, in three setups:
MCP: the model uses this server.
No Tools: the dataset is pasted into the prompt.
Sandbox: the dataset is in the prompt and the model also has a Python code sandbox.
Task | Category | Checked against |
| query | all 36 building IDs and the total |
| query | the ID and height of the tallest building |
| analytic | the number of party walls (CityDPC: 10) |
| state change | the saved file: measuredHeight +2 m |
| state change | the saved file: new building, attributes and valid geometry, importable by CityDPC |
The paper tested GPT-OSS-120B, Mistral Small 4 and GPT-5.4 Mini, all with high reasoning effort. Each combination of model, setup and task ran 100 times, for 4,500 runs in total.
Results
Averages across the five tasks (full tables):
Setup | GPT-OSS success | Mistral success | GPT-5.4 Mini success | Avg. tokens (GPT-OSS / Mistral / GPT-5.4 Mini) |
MCP | 93% | 80% | 97% | 14,865 / 25,755 / 15,429 |
No Tools | 40% | 34% | 61% | 62,188 / 80,633 / 71,879 |
Sandbox | 34% | 30% | 64% | 236,803 / 83,517 / 117,703 |
evaluation/results/paper_runs.csv has one row per run: tokens, tool calls, the sequence of tool calls, duration and success. The paper's Tables 2 and 3 are computed from this file, and all 300 numbers match the paper:
git clone https://github.com/HamerCode/CityDPC-MCP && cd CityDPC-MCP
uv sync --extra eval
uv run python evaluation/analyze.py tablesRunning the benchmark yourself
cp evaluation/.env.example evaluation/.env # add your endpoint and API key
uv run python evaluation/run.py --setup mcp,no_tools --model gpt-oss-120b \
--api completions --reasoning-effort high --runs 100
uv run python evaluation/analyze.py extract evaluation/logs/<run> -o my_runs.csv
uv run python evaluation/analyze.py tables my_runs.csvThe runner works with any OpenAI-compatible endpoint, through either the Responses or the Chat Completions API. The paper ran GPT-OSS and Mistral through KI:Connect NRW Chat Completions, and GPT-5.4 Mini through the Responses API of ChatGPT's Codex backend. Prompts, task specifications, validation and success criteria are the ones used for the paper. Current models will not reproduce the historical runs exactly.
The code_interpreter (Sandbox) setup also needs a Docker code-sandbox MCP server (CODE_SANDBOX_MCP_BINARY, e.g. code-sandbox-mcp). Leave that setup out if you don't have one.
The following check needs no API key. It exercises the server over stdio, solves all five tasks with the reference tool calls and confirms that the evaluation's validation accepts each solution:
uv run python tests/smoke_test.pyRepository layout
src/citydpc_mcp/server.py MCP server (19 tools)
evaluation/
cases.py, schemas.py task prompts and answer schemas
run.py, client.py benchmark runner and OpenAI-compatible tool-calling client
validation.py checks answers and saved datasets
analyze.py logs -> per-run CSV -> result tables
ground_truth.py regenerates data/ground_truth.json with CityDPC
data/ sample dataset and ground truth
results/ per-run results and tables of the paper
tests/smoke_test.py offline end-to-end checkCredits and license
CityDPC is developed by Maxim Shamovich and Simon Raming at RWTH-E3D under the Apache 2.0 license. This server installs CityDPC v0.1.22 pinned to commit 4ede6d7. This repository is licensed under Apache 2.0 as well (LICENSE).
If you use this work, please cite the paper CityDPC-MCP: Enabling 3D city model processing via Model Context Protocol and link a specific commit of this repository.
Available Tools
19 toolsanalyse_datasetB
Analysiert das Dataset und gibt detaillierte Informationen über die Gebäude zurück.
Returns: dict: Analyse-Ergebnisse mit Informationen über GML-Version, CRS, LoD, Anzahl Gebäude etc.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only inspection, but does not state that it has no side effects, whether a dataset must be loaded first, or any cost/performance traits. The 'Returns' block largely restates what an output schema would already cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the core action. The explicit 'Returns' section is somewhat redundant since an output schema exists, but it costs only one line and does not obscure the main point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter inspection tool with an output schema, the description covers what the tool does and what comes back. The main omission is the precondition/side-effect profile, which matters more because no annotations exist to cover it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters (empty object schema), so there is nothing to document or compensate for. Baseline 4 applies; the description introduces no parameter-related ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Analysiert das Dataset') and enumerates the returned content (GML-Version, CRS, LoD, Anzahl Gebäude), which distinguishes it reasonably well from siblings like number_of_buildings. However, it does not explicitly contrast itself with overlapping siblings such as list_datasets or number_of_buildings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, nor prerequisites such as whether a dataset must first be loaded via load_dataset. The only implied usage is that it inspects a dataset, which the reader must infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calculate_roof_volume_by_idB
Berechnet das Dachvolumen für ein spezifisches Gebäude.
Args: building_id: Die gml:id des Gebäudes
Returns: float: Das berechnete Dachvolumen in Kubikmetern
| Name | Required | Description | Default |
|---|---|---|---|
| building_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the return type and unit (float, cubic meters), but says nothing about dependencies (e.g. whether the building must be loaded first), permissions, or error behavior for a computation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The docstring is short, front-loaded, and uses a clean Args/Returns structure with no wasted prose. Slightly boilerplate, but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the Returns line is somewhat redundant and return values need not be re-explained. However, with no annotations and a dependency-sensitive computation, the description leaves behavioral gaps that an agent would need to resolve.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it clarifies that building_id is "Die gml:id des Gebäudes," adding meaning beyond the bare string type in the schema. Only one parameter, and its semantics are adequately explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ("Berechnet") and resource ("Dachvolumen") scoped to a single building, which is clearly distinct from retrieval siblings like get_building_by_id. It does not explicitly name alternatives, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus siblings, no prerequisites, and no exclusions. The agent must infer usage purely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_buildingA
Erstellt ein neues Gebäude mit wählbarem Detaillierungsgrad (LoD 0/1/2).
LoD: 0=2D-Grundfläche | 1=3D-Quader | 2=3D mit Dachgeometrie (Standard)
Das Gebäude wird direkt zum aktiven Dataset hinzugefügt (Single Source of Truth). Verwende take_snapshot() vor dem Erstellen und save_dataset() zum Speichern.
Args: id: Eindeutige Gebäude-ID groundsCoordinates: 2D-Grundriss als Liste von Koordinaten [[x1,y1], [x2,y2], ..., [x1,y1]] groundSurfaceHeight: Bodenhöhe in Metern geometryHeight: Gebäudehöhe in Metern (erforderlich für LoD 1+2) roofType: Dachtyp (für LoD 2): "1000"=Flach, "1010"=Pult, "1020"=Pult versetzt, "1030"=Sattel, "1040"=Walm, "1070"=Zelt roofHeight: Dachhöhe in Metern (für LoD 2, außer "1000") roofOrientation: Dachausrichtung als Koordinaten-Index (für "1010","1020","1030") lod: Detaillierungsgrad 0-2 (Standard: 2) isRoofEdge: Für LoD 0, ob Koordinaten Dachkante sind (Standard: False)
Returns: dict: Das erstellte Gebäude mit allen Attributen und Bounding Box
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| lod | No | ||
| roofType | No | ||
| isRoofEdge | No | ||
| roofHeight | No | ||
| geometryHeight | No | ||
| roofOrientation | No | ||
| groundsCoordinates | Yes | ||
| groundSurfaceHeight | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the building is added directly to the active dataset (single source of truth), that a snapshot should precede creation and save_dataset persist it, and what the return value contains (all attributes plus bounding box). It does not discuss permissions or failure/reversibility semantics, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and LoD semantics before the parameter block, and the Args entries each earn their place. The Returns section partially duplicates the existing output schema, which is minor waste and keeps it just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters at 0% schema coverage and no annotations, the definition supplies the parameter grammar, defaults, conditional requirements, side effects on the active dataset, and workflow routing. Nothing essential for a correct invocation is missing; return details are covered by the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate fully, and it does: every one of the 9 parameters is documented with format (e.g. groundsCoordinates as closed [[x1,y1],...] ring), units (Meter), defaults (lod=2, isRoofEdge=False), conditionality (geometryHeight for LoD 1+2, roofType/roofHeight for LoD 2), and enumerated roofType codes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Erstellt ein neues Gebäude') plus the operative scope (wählbarer Detaillierungsgrad LoD 0/1/2). It further distinguishes itself from siblings by naming the take_snapshot/save_dataset workflow, so an agent can tell what this tool does and how it fits alongside create_dataset and enrich_building.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit procedural guidance: use take_snapshot() before creating and save_dataset() to persist, which routes the agent through sibling tools. There is no explicit 'when not to use' or disambiguation against enrich_building/remove_building_from_dataset, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_datasetA
Erstellt ein neues, leeres Dataset als GML- oder JSON-Datei und lädt es.
Praktisch zum Testen: Danach können direkt mit create_building() Gebäude hinzugefügt und mit save_dataset() gespeichert werden. Bestehende Dateien werden nicht überschrieben.
Args: filename: Name der neuen Datei, Endung '.json' (CityJSON) oder '.gml' (CityGML), z.B. 'test.city.json' title: Optionaler Titel des Datasets epsg: EPSG-Code des Koordinatenreferenzsystems (Standard: 25832 = ETRS89 / UTM 32N)
Returns: str: Erfolgsmeldung oder Fehlermeldung
| Name | Required | Description | Default |
|---|---|---|---|
| epsg | No | ||
| title | No | ||
| filename | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It usefully discloses the non-overwrite guarantee and that the dataset is loaded, but omits permissions, whether it becomes the active dataset, and concrete error behavior beyond 'Fehlermeldung'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then a practical note, then Args/Returns. Because the schema carries no descriptions, the Args block earns its place rather than repeating structured data; the padding is minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be detailed, and the description covers the create-and-load behavior plus the non-overwrite constraint. For a creation tool with zero annotations it could say more about failure modes or whether it replaces the current dataset.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: it explains the '.json' (CityJSON) / '.gml' (CityGML) extension convention with an example ('test.city.json'), the optional title, and the epsg default meaning (25832 = ETRS89 / UTM 32N). It does not mark which parameter is required, a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Erstellt ein neues, leeres Dataset als GML- oder JSON-Datei und lädt es'), including the accepted formats and the load side effect. It also distinguishes itself from siblings by naming the follow-on tools create_building() and save_dataset().
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context ('Praktisch zum Testen') and an explicit guardrail ('Bestehende Dateien werden nicht überschrieben'). It does not name an alternative such as load_dataset for opening existing files, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enrich_buildingA
Reichert ein Gebäude mit semantischen Informationen an.
Die Änderungen werden direkt im Dataset vorgenommen (Single Source of Truth). Verwende take_snapshot() vor wichtigen Änderungen und save_dataset() zum Speichern.
Args: building_id: Die gml:id des Gebäudes measured_height: Gemessene Gebäudehöhe in Metern roof_type: Dachtyp-Code (z.B. "1000"=Flach, "1030"=Sattel, "1040"=Walm, "3100"=andere) roof_height: Dachhöhe in Metern function: Gebäudefunktion-Code nach CityGML (z.B. "31001_1010"=Wohngebäude, "31001_2000"=Gewerbe) usage: Nutzung des Gebäudes (z.B. "residential", "commercial", "industrial") year_of_construction: Baujahr (z.B. 2020) storeys_above_ground: Anzahl der Stockwerke über dem Boden storeys_below_ground: Anzahl der Stockwerke unter dem Boden (Keller) creation_date: Erstellungsdatum des Gebäudes (ISO-Format: "2020-01-15") Returns: dict: Das angereicherte Gebäude mit allen Attributen
| Name | Required | Description | Default |
|---|---|---|---|
| usage | No | ||
| function | No | ||
| roof_type | No | ||
| building_id | Yes | ||
| roof_height | No | ||
| creation_date | No | ||
| measured_height | No | ||
| storeys_above_ground | No | ||
| storeys_below_ground | No | ||
| year_of_construction | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that changes are made directly in the dataset, recommends snapshotting before important changes, and notes save_dataset() is needed to persist. It does not cover permissions, validation failures, or rollback behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and mutation workflow are front-loaded, and the Args list is necessary because the schema lacks property descriptions. The Returns line is somewhat redundant given the output schema exists, but overall the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with no annotations and 0% schema coverage, the description is complete enough: it explains the mutation, persistence workflow, and all parameter meanings. It could say more about invalid inputs or permission requirements, but the output schema handles return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all 10 parameters. It describes every parameter, including units, code examples, and intended meaning for roof_type, function, usage, and date formats, adding substantial value beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: enriching a building with semantic information. It is clearly distinct from create_building, remove_building_attributes, and read-only siblings, but it does not explicitly name alternatives or contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear workflow guidance: use take_snapshot() before important changes and save_dataset() to persist. This establishes context and prerequisites, though it does not explicitly say when to choose enrich_building over create_building or manual attribute editing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filter_datasetA
Filtert das Dataset nach Adressen oder Koordinaten und setzt das gefilterte Dataset als aktives Dataset.
Das gefilterte Dataset wird zum neuen aktiven Dataset. Alle nachfolgenden Tool-Aufrufe arbeiten dann nur mit den gefilterten Gebäuden. Mit load_dataset() kann jederzeit das Original neu geladen werden.
WICHTIG: Nach dem Filtern können keine Änderungen mehr gespeichert werden, um das Original zu schützen. Die History wird geleert!
Args: addressRestriciton: Dictionary mit Adressbeschränkungen (z.B. {"thoroughfareName": "Stakenholt"}) borderCoordinates: Liste von Koordinaten für einen rechteckigen Bereich (z.B. [[x1,y1], [x2,y2], [x3,y3], [x4,y4]])
Returns: dict: Informationen zum gefilterten Dataset (building_count, minimum, maximum, filepath, title)
| Name | Required | Description | Default |
|---|---|---|---|
| borderCoordinates | No | ||
| addressRestriciton | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the active-dataset side effect, that the history is cleared, and that saving changes is blocked afterwards to protect the original. These are non-obvious, high-impact consequences an agent must know before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: purpose sentence first, then behavioral warnings, then Args, then Returns. Some redundancy (the filtered-dataset concept is restated across sentences) and the Returns block duplicates the existing output schema, costing a little tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the Returns block is optional, and the destructive/warning behavior is covered. Given nested-object parameters with zero schema description coverage, the definition is largely complete, though it omits how the two optional filters combine and the exact lifetime of the no-save restriction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it supplies a concrete address example ({"thoroughfareName": "Stakenholt"}) and explains borderCoordinates as a rectangular area with a coordinate-list example. It leaves some ambiguity (whether the two filters are exclusive, what happens when neither is given), which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Filtert das Dataset nach Adressen oder Koordinaten') and the side effect that distinguishes it ('setzt das gefilterte Dataset als aktives Dataset'). It also names load_dataset as the way back to the original, so an agent can separate it from sibling reload/create tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: subsequent calls operate only on filtered buildings, and load_dataset() restores the original. The warning that no changes can be saved after filtering implicitly tells the agent not to filter when edits are needed. It stops short of explicitly stating when to choose filtering over analyse_dataset or how the tool relates to save/snapshot siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_all_buildingsB
Gibt Informationen zu allen Gebäuden im Dataset zurück.
Returns: list[dict]: Liste mit Dictionaries aller Gebäude und deren Attributen
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says what is returned but omits whether the operation is read-only, requires a loaded dataset, may be expensive for large collections, or has any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose. The second sentence restates return type, which is redundant given the output schema exists, but it is not excessive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument read tool with an output schema, the return values are covered elsewhere. However, the description does not state that a dataset must be loaded first, which is important context in a dataset-oriented toolset.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds no parameter information because there is none to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('Gibt Informationen zu allen Gebäuden im Dataset zurück') and scope ('allen Gebäuden'). The scope distinguishes it implicitly from siblings like get_building_by_id or number_of_buildings, but it does not name an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are given. The agent must infer from the name alone that this is for retrieving all buildings rather than a count, an ID list, or a single building.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_buiding_Id_listB
Gibt eine Liste aller Gebäude-IDs im Dataset zurück.
Returns: list[str]: Liste aller gml:id Werte der Gebäude
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It implies a read operation but never states it, and the critical scope question — which dataset the IDs come from, given no dataset parameter — is unexplained. Return type is covered by the output schema, so the description adds little behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but the second half ('Returns: list[str]...') restates what the first sentence and the output schema already convey, so it is partially redundant rather than maximally efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers the return value, so that needn't be explained. However, for a 0-parameter getter in a toolset with list_datasets/load_dataset, the description omits which dataset scope applies, leaving a real ambiguity an agent would want resolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline 4 per the rubric. Nothing is required of the description here, though it could have clarified that the operation implicitly targets the currently loaded dataset.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns a list of all building IDs (gml:id values) in the dataset. This is distinguishable from get_all_buildings (full objects) or get_building_by_id (single object), but the distinction is left implicit rather than called out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this over the many ID-related siblings (get_all_buildings, get_building_by_id, number_of_buildings). No prerequisites, no exclusions, no alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_building_by_idB
Gibt alle Informationen zu einem spezifischen Gebäude zurück.
Args: building_id: Die gml:id des Gebäudes
Returns: dict: Dictionary mit allen Gebäude-Attributen
| Name | Required | Description | Default |
|---|---|---|---|
| building_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It mentions that all building attributes are returned, but says nothing about behavior when the id does not exist, permissions, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded, with the purpose in the first sentence. The Args/Returns blocks are standard docstring boilerplate, and the Returns block is largely redundant with the existing output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter lookup with an output schema, the essential facts (what it returns, what the id is) are covered. The main omission is not-found/error behavior, which matters little for this simple read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it specifies that building_id is the gml:id of the building, information the schema (plain string) does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns all information for a specific building. The word 'spezifischen' implicitly distinguishes it from the list-style siblings (get_all_buildings, get_buiding_Id_list), though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the lookup-by-id tool, but the description never states when to prefer it over get_all_buildings or get_buiding_Id_list, nor any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dataset_historyB
Zeigt die Historie aller Snapshots und Speicherungen.
Returns: dict: Übersicht über Snapshots, Speicherungen und aktuellen Zustand
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read ('Zeigt') and names the returned contents (snapshots, saves, current state), but never states that the operation is side-effect free, nor mentions permissions, cost, or ordering of the history. That is thin for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no padding. The trailing 'Returns:' block restates information the output schema already carries, which is mild redundancy but not harmful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and it correctly avoids needing to since there are no inputs. For a simple parameterless read tool the description is essentially complete, though it could say a bit more about when the history is useful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description adds nothing parameter-related, which is correct since there is nothing to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it shows the history of all snapshots and saves. The scope ('aller Snapshots und Speicherungen') is concrete enough to distinguish it from the many dataset/building tools, though it does not explicitly differentiate itself from snapshot siblings like take_snapshot or rollback_to_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as take_snapshot or rollback_to_snapshot, which are the natural adjacent tools. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_party_wallsB
Findet alle angrenzenden Wände (Party Walls) zwischen Gebäuden im Dataset.
Returns: list: Liste aller erkannten Party Walls als [id of b0, id of w0, id of b1, id of w1, area, collision coordinates]
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only lookup and describes the return tuple, but says nothing about whether a dataset must be loaded, whether the scan is expensive on large datasets, or how empty results are handled. The return-value explanation is largely redundant given an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is short and front-loads the core purpose in one sentence. The 'Returns' block is verbose and duplicates what the output schema already provides, costing some efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and a separate output schema, the description covers the essentials of what the tool does. However, it omits the key operational precondition — that the tool acts on the loaded dataset — which an agent needs in order to invoke it correctly in sequence with load_dataset.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document and the baseline of 4 applies. Schema coverage is 100% with an empty object schema, so no parameter semantics are missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Findet alle angrenzenden Wände (Party Walls) zwischen Gebäuden im Dataset'), which is unambiguous and matches the tool name. It does not explicitly contrast itself with siblings like analyse_dataset or get_all_buildings, but the resource is distinctive enough to disambiguate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any stated prerequisite such as needing a dataset loaded first. The only contextual hint is the phrase 'im Dataset', which implies it operates on the currently active dataset but never says so explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_datasetsB
Listet alle verfügbaren GML- oder JSON-Datasets im Data-Verzeichnis auf.
Returns: list[str]: Liste aller verfügbaren Dataset-Dateinamen
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden. 'Listet ... auf' plus the Returns block effectively communicates that this is a non-mutating enumeration and reveals which dataset formats exist (GML or JSON), but it does not state ordering, whether the directory is fixed, or what an empty result means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and resource, no padding. The 'Returns: list[str]' block is largely redundant given that a structured output schema exists, costing it a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing whose return type is covered by the output schema, the description covers what is needed to call it. Minor gaps remain around ordering, the location of the 'data directory', and behavior on an empty directory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify; the baseline for a no-argument tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — listing all available GML/JSON datasets in the data directory — which an agent can read unambiguously. It does not, however, distinguish itself from the many dataset-oriented siblings (load_dataset, filter_dataset, analyse_dataset, create_dataset), so the agent must infer which one to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites, and no reference to an alternative tool. With siblings like load_dataset and get_dataset_history in the same family, the description should say something like 'use to discover datasets before loading one', but says nothing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_datasetB
Lädt ein Dataset aus einer GML- oder JSON-Datei.
Erstellt automatisch einen Snapshot des ursprünglichen Zustands.
Args: filename: Der Name der GML- oder JSON-Datei (z.B. 'EssenExample.gml')
Returns: str: Erfolgsmeldung oder Fehlermeldung
| Name | Required | Description | Default |
|---|---|---|---|
| filename | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses an important side effect: an automatic snapshot of the original state is created. However, it omits whether loading overwrites an existing dataset, what permissions are needed, and how failures are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and side effect, then terse Args/Returns. Slightly wasteful that the Returns line restates something the output schema already covers, but the overall length is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. The key side effect (snapshot) is documented, but for a state-changing load operation the description leaves open whether existing data is replaced and what the snapshot implies for rollback_to_snapshot.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: filename is identified as the GML or JSON file name, with a concrete example ('EssenExample.gml') and the accepted formats. This meaningfully clarifies the single parameter beyond the bare string type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: loads a dataset from a GML or JSON file, and notes the auto-snapshot. An agent can distinguish it from create_dataset/save_dataset reasonably well. It stops short of explicitly naming the siblings it differs from, so not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus create_dataset, list_datasets, or save_dataset. Usage is only weakly implied by 'loads from a file'. No prerequisites (file location, source of filename) or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
number_of_buildingsA
Gibt die Anzahl der Gebäude im aktuellen Dataset zurück.
Returns: int: Anzahl der Gebäude
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses the return type (int) and that the count is scoped to the 'current dataset', which is useful context. It does not mention behavior when no dataset is loaded or any error conditions, leaving a modest gap for what is a trivially safe read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste. The purpose is front-loaded and the return type follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool, the definition covers what is needed. Since an output schema exists, the explicit Returns block is mildly redundant, but the scoping phrase 'im aktuellen Dataset' adds genuinely useful context about which dataset is counted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to compensate for. The baseline for a no-parameter tool is a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-and-resource combination: returns the count of buildings in the current dataset. The purpose is unambiguous. It does not, however, differentiate itself from the sibling get_all_buildings, which an agent could otherwise use to achieve the same count.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as get_all_buildings, nor any stated precondition (e.g. that a dataset must already be loaded). Usage must be inferred entirely from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_building_attributesA
Entfernt Attribute eines Gebäudes (setzt sie auf None).
Args: building_id: Die gml:id des Gebäudes attributes: Liste von Attributnamen (gleiche Namen wie in enrich_building, z.B. "measuredHeight", "roofType"). Unbekannte Namen werden ignoriert.
Returns: dict: Das aktualisierte Gebäude mit allen Attributen
| Name | Required | Description | Default |
|---|---|---|---|
| attributes | Yes | ||
| building_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key trait that removal means setting values to None rather than deleting the keys, plus that unknown names are silently ignored. It does not state whether changes persist to the dataset, whether a save is required, or any permission requirements — significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in one sentence, followed by a compact Args/Returns block. Every line adds information; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Parameters and return shape are covered, and an output schema exists so return details are a bonus rather than a necessity. The gap is the mutation's lifecycle semantics (persistence, snapshot interaction, auth) which, absent any annotations, leaves the agent guessing about side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: building_id is identified as the gml:id, and attributes is described as a list of names using the same vocabulary as enrich_building, with concrete examples (measuredHeight, roofType) and the unknown-name fallback.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it removes (sets to None) attributes of a building. Referencing enrich_building for attribute naming implicitly signals this is the inverse operation, giving some sibling differentiation, though the contrast is not stated explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the link to enrich_building (same attribute names), so an agent can infer this is the undo of enrichment. However, there is no explicit when-to-use or when-not-to-use guidance, nor any mention of prerequisites such as having the building loaded in a dataset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_building_from_datasetA
Entfernt ein Gebäude aus dem Dataset.
Die Änderung wird direkt im Dataset vorgenommen (Single Source of Truth). Verwende take_snapshot() vor wichtigen Änderungen und save_dataset() zum Speichern.
Args: building_id: Die gml:id des zu entfernenden Gebäudes
Returns: dict: Bestätigungsmeldung mit Anzahl verbleibender Gebäude
| Name | Required | Description | Default |
|---|---|---|---|
| building_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose real behavioral traits: the change is applied directly in-place (Single Source of Truth) and is not persisted without save_dataset(), plus snapshotting enables recovery. It omits permissions/auth needs and whether the removal is reversible by other means, but the persistence and mutation semantics are well conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by behavioral context and then Args/Returns. Appropriately sized for a one-parameter mutation tool, though the Returns section is slightly redundant given an output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation with no annotations, the description covers purpose, in-place persistence behavior, snapshot/save workflow, and the parameter's meaning. The output schema makes the Returns note unnecessary, and permissions are unaddressed, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it does explain that building_id is the gml:id of the building to remove. This adds meaningful type/identity semantics beyond the bare 'string' schema, fully covering the single required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Entfernt ein Gebäude aus dem Dataset'), so an agent immediately knows this is a removal operation on a building. It does not, however, distinguish itself from the sibling 'remove_building_attributes', leaving a minor ambiguity between the two removal tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit workflow guidance: use take_snapshot() before important changes and save_dataset() to persist, naming two sibling tools directly. It stops short of a full when/when-not contrast (e.g., versus remove_building_attributes), but the surrounding context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rollback_to_snapshotA
Stellt einen früheren Dataset-Zustand wieder her.
Wie Git-Reset: Kehrt zu einem gespeicherten Zustand zurück. Alle Änderungen nach diesem Snapshot gehen verloren!
Args: snapshot_index: Index des Snapshots (von get_dataset_history() erhalten)
Returns: dict: Informationen über den Rollback
| Name | Required | Description | Default |
|---|---|---|---|
| snapshot_index | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does the most important part: 'Alle Änderungen nach diesem Snapshot gehen verloren!' clearly marks this as a destructive, irreversible-in-practice mutation. It does not cover permissions/authorisation needs or whether the rollback can itself be reverted, so it falls short of complete disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and the loss warning in the first three lines; the Args block adds real value. The 'Returns: dict: Informationen über den Rollback' section is redundant given an output schema exists and is the only wasted sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with an output schema (so return values need no prose), the description supplies purpose, destructive consequence, and parameter provenance. Minor gaps remain around authorisation and the state of the snapshot list after rollback, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: snapshot_index is defined as the snapshot's index and, critically, its provenance is given (obtained from get_dataset_history()). It does not describe indexing base or ordering, so it is not fully self-sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Stellt einen früheren Dataset-Zustand wieder her') and reinforces it with the 'wie Git-Reset' analogy, which cleanly separates it from the neighbouring take_snapshot (create), get_dataset_history (read history) and save_dataset (persist) tools. An agent can pick this tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use and a prerequisite: the snapshot_index comes from get_dataset_history(), so the agent knows what to call first. It also flags the condition of data loss. It stops short of explicit when-not guidance or naming an alternative tool for undoing the rollback itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_datasetB
Speichert das Dataset dauerhaft in die GML/JSON-Datei.
Übernimmt alle aktuellen Änderungen permanent in die Datei.
Args: description: Beschreibung des Speicherns target_format: Optionales Zielformat ("gml" oder "json"). Wenn None, wird das ursprüngliche Format beibehalten. Returns: dict: Informationen über den Speichervorgang
| Name | Required | Description | Default |
|---|---|---|---|
| description | No | ||
| target_format | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that saving is permanent ('dauerhaft', 'permanent') and writes all current changes to the file, but says nothing about overwrite behavior, permissions, or reversibility, which matters for a persistence/mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first two sentences with no padding. The Args/Returns blocks re-state structured info, but the overall size remains tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described. For a 2-parameter persistence tool with no annotations, the description covers purpose and both parameters adequately but omits contrast with snapshot/history siblings and any risk or overwrite context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains target_format well, naming the accepted values ('gml' or 'json') and the None default (keep original format), but the 'description' parameter is only glossed as 'Beschreibung des Speicherns', which is thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (speichern/save) and resource (dataset) plus the destination artifact (GML/JSON file), so the agent knows it persists data. It does not differentiate itself from related state-management siblings like take_snapshot or rollback_to_snapshot, so it stays below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use save_dataset versus take_snapshot, get_dataset_history, or create_dataset beyond the implicit notion of persisting changes. No prerequisites or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_snapshotA
Erstellt einen Snapshot des aktuellen Dataset-Zustands.
Wie ein Git-Commit: Speichert den aktuellen Zustand für spätere Rollbacks.
Args: description: Beschreibung des Snapshots
Returns: dict: Informationen über den erstellten Snapshot
| Name | Required | Description | Default |
|---|---|---|---|
| description | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully communicates that the operation captures current state (non-destructive, rollback-oriented), but omits overwrite semantics, snapshot limits, persistence, or permissions. Adequate but incomplete for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior, and the Args/Returns block is tidy in the conventional docstring style. The Returns line is somewhat redundant given an output schema exists, but the overall size is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation tool with no annotations, the description explains the essential effect and purpose, and an output schema already covers return values. The remaining gaps are the undocumented parameter semantics and the absence of any when-to-use guidance relative to save_dataset/rollback_to_snapshot siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter has 0% schema description coverage, and the description only says 'description: Beschreibung des Snapshots', which merely restates the parameter name. No format, length, or whether it is optional/default behavior is conveyed, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Erstellt einen Snapshot des aktuellen Dataset-Zustands') and reinforces it with a concrete analogy ('Wie ein Git-Commit'). An agent can immediately distinguish this from siblings like rollback_to_snapshot or save_dataset without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'für spätere Rollbacks' hints at when snapshots are used, but the description never says when to call this vs alternatives such as save_dataset, get_dataset_history, or rollback_to_snapshot. The agent must infer the pairing with rollback_to_snapshot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
v1.0.0- First observed
analyse_dataset - First observed
calculate_roof_volume_by_id - First observed
create_building - First observed
create_dataset - First observed
enrich_building - First observed
filter_dataset - First observed
get_all_buildings - First observed
get_buiding_Id_list - First observed
get_building_by_id - First observed
get_dataset_history - First observed
get_party_walls - First observed
list_datasets - First observed
load_dataset - First observed
number_of_buildings - First observed
remove_building_attributes - First observed
remove_building_from_dataset - First observed
rollback_to_snapshot - First observed
save_dataset - First observed
take_snapshot
TDQS
Scored across 19 tools
Most tools have distinct resource/action boundaries: dataset lifecycle, building CRUD, semantic enrichment, spatial analysis, and snapshot history are clearly separated. However, number_of_buildings and analyse_dataset overlap somewhat, and get_all_buildings/get_buiding_Id_list/get_building_by_id require reading descriptions to distinguish intent.
Tool names mostly follow a snake_case, verb-oriented pattern (list_datasets, load_dataset, create_building, remove_building_attributes). Minor deviations include the noun-only number_of_buildings and the typo/camelCase segment in get_buiding_Id_list.
At 19 tools, the set is borderline heavy for the domain. Most tools earn their place, but there is some redundancy around dataset/building summaries that could be consolidated.
The surface covers dataset loading, creation, filtering, saving, history, and building CRUD, enrichment, removal, and analysis. Minor gaps remain, such as no explicit delete/update dataset operation and no general building geometry update tool.
Maintenance
Related MCP Connectors
Create, inspect, validate, and save editable 3D building scenes with Pascal's hosted MCP server.
MCP server for progressive tool usage at any scale (see https://klavis.ai)
MCP server for the Seline Analytics API
MCP server for Product Management
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceBrowser-based geometry processing server that enables AI agents to create geometry, run analysis and operators, inspect results, and take screenshots via MCP.93 npm6MIT
- AlicenseAqualityBmaintenanceMCP server for IFC geometric audit, providing tools to detect space clashes, extract space inventories, compute surface losses, check boundaries, and verify opening correspondences in BIM models.6Apache 2.0
- AlicenseCqualityCmaintenanceMCP server for COBie Excel validation, updates, and PDF extraction.631MIT
- AlicenseBqualityBmaintenanceMCP server for reading, viewing, and editing IFC models using IfcOpenShell and Bonsai. Enables querying model data, viewing via Blender's viewport, and making IFC-semantic edits without raw mesh operations.28MIT