Skip to main content
Glama

elster-mcp-server

Originally forked from lukasschwarz/elster-mcp-server (MIT). This version replaces its click automation with an HTTP form engine that drives any ELSTER form by Kennzahl, reads back the exact data that would be transmitted, and by design never submits — the taxpayer clicks "Absenden". See CLAUDE.md for the architecture.

A Model Context Protocol (MCP) server that lets Claude (or any MCP-capable client) drive the German tax portal ELSTER via Puppeteer.


English

  • This project is an experimental, community-built tool. It is not affiliated with, endorsed by, or supported by the Bundesministerium der Finanzen, the ELSTER project, or any tax authority.

  • The official, supported way to submit tax data programmatically is the ERiC library (registration as a software vendor required). This tool instead automates the public ELSTER web portal with a real user session — the same path a human user takes — using credentials YOU provide.

  • The official ELSTER terms of use ("Nutzungsbedingungen") may restrict automated access to the portal. Whether your specific use is permitted is your responsibility to verify before running this software.

  • Use at your own risk. The author(s) provide this software AS IS, WITHOUT WARRANTY OF ANY KIND (see LICENSE). The author(s) accept NO liability for incorrect tax submissions, account suspensions, missed deadlines, lost data, or any other consequences arising from the use of this software.

  • This project is not tax advice (no "Hilfeleistung in Steuersachen" in the sense of § 2 StBerG). If you are unsure whether a submission is correct, consult a Steuerberater.

  • Operators using this software in a commercial context (e.g. submitting on behalf of third parties) may be subject to the German Steuerberatungsgesetz and must verify their own licensing situation.

Deutsch

  • Dieses Projekt ist ein experimentelles, von der Community gebautes Werkzeug. Es ist weder vom Bundesministerium der Finanzen noch vom ELSTER-Projekt noch von einer Finanzbehörde unterstützt, autorisiert oder geprüft.

  • Der offizielle, vom BMF unterstützte Weg zur programmatischen Übermittlung von Steuerdaten ist die ERiC-Bibliothek (Registrierung als Softwarehersteller erforderlich). Dieses Tool nimmt stattdessen den Weg über das öffentliche ELSTER-Webportal — denselben Weg, den ein menschlicher Nutzer per Browser geht — mit Zertifikatsdaten, die DU bereitstellst.

  • Die offiziellen ELSTER-Nutzungsbedingungen können automatisierten Zugriff auf das Portal einschränken oder verbieten. Es liegt in deiner alleinigen Verantwortung zu prüfen, ob dein konkreter Anwendungsfall erlaubt ist, bevor du dieses Tool nutzt.

  • Nutzung auf eigenes Risiko. Die Autor:innen stellen die Software OHNE JEGLICHE GEWÄHRLEISTUNG bereit (siehe LICENSE). Die Autor:innen übernehmen keine Haftung für fehlerhafte Steuerübermittlungen, gesperrte Konten, versäumte Fristen, Datenverluste oder sonstige Folgen aus der Nutzung dieser Software.

  • Dieses Projekt ist keine Steuerberatung im Sinne des § 2 StBerG. In Zweifelsfällen ist ein:e Steuerberater:in zu konsultieren.

  • Wer diese Software gewerblich einsetzt (z.B. Übermittlung im Auftrag Dritter), unterliegt unter Umständen dem Steuerberatungsgesetz und muss seine Berechtigung selbst sicherstellen.

Practical safeguards built into the tool

  • The only tool that actually transmits data is elster_ustva_confirm — it requires an explicit second call after elster_ustva_start has paused at AWAITING_CONFIRM. Nothing is sent without that second confirmation.

  • The EÜR and ESt tools never submit. They only fill the form up to "Prüfen" and stop, so you review and submit yourself in the ELSTER portal.

  • All sync / history / inbox tools are read-only and never modify state on the ELSTER side.


Related MCP server: elster-mcp-server

Features

Tool

What it does

Submits?

elster_login_test

Verifies your certificate + password can log in

No

elster_config_show

Shows the loaded config (secrets redacted)

No

elster_kennziffern_list

Returns the supported UStVA Kennziffern with descriptions

No

elster_ustva_generate_xml

Generates a UStVA XML snapshot (archive only)

No

elster_ustva_detect_reverse_charge

Detects §13b reverse-charge suppliers

No

elster_datenuebernahme_list

Lists earlier submissions ELSTER offers to carry over into a new form

No

elster_ustva_start

Logs in, fills, runs Prüfung, then pauses for confirmation

Pauses

elster_ustva_confirm

Clicks "Absenden" after you reviewed

Yes

elster_eur_start

Fills Anlage EÜR up to Prüfung, then "Speichern und Verlassen"

No

elster_est_start

Opens ESt 1 A, fills basics, runs Prüfung, keeps browser open 30 min

No

elster_submissions_list

Lists submitted forms with their nachrichtId / aufgabeId

No

elster_submission_protocol

Reads the full field-level content of past submissions

No

elster_sync_history

Reads "Übermittelte Formulare" (optionally with PDFs)

No

elster_sync_inbox

Reads ELSTER inbox (optionally with PDFs)

No

elster_session_status / _list / _cancel

Session management

No

elster_drafts_list

Lists saved drafts with their aufgabeId

No

elster_form_new / elster_form_open

Starts a new form (any type) or opens a draft in a persistent session

No

elster_form_page / elster_form_crawl

Reads a page / walks a whole form: fields by Kennzahl, repeat groups, navigation

No

elster_form_set / _add_row / _delete_row

Fills fields and table rows by Kennzahl

No

elster_form_press

Escape hatch for other form commands (send/delete refused)

No

elster_form_check / elster_form_save

Runs "Prüfen" (incl. provisional tax result) / saves the draft

No

Requirements

  • Node.js ≥ 18

  • An ELSTER certificate file (.pfx) — get it from https://www.elster.de → "Mein ELSTER" → "Mein Benutzerkonto" → "Zertifikat verlängern"

  • The certificate password

  • Your Steuernummer and Bundesland-Code

Install

git clone https://github.com/Infraviored/elster-mcp-server.git
cd elster-mcp-server
npm install
npm run build

Puppeteer will install a bundled Chromium on first install (~150 MB).

Configuration

cp config.example.json config.json
$EDITOR config.json

All keys in config.json can be overridden by environment variables (ELSTER_PFX_PATH, ELSTER_PASSWORD, ELSTER_TAX_NUMBER, ELSTER_STATE_CODE, ELSTER_NAME, ELSTER_FIRST_NAME, ELSTER_STREET, ELSTER_HOUSE_NUMBER, ELSTER_ZIP, ELSTER_CITY, ELSTER_COUNTRY, ELSTER_DOWNLOAD_DIR, ELSTER_SCREENSHOT_DIR, ELSTER_HEADLESS, ELSTER_EST_SKIP_EUR). Env vars win over the file.

You can also point the loader at a different config file via ELSTER_CONFIG_PATH=/path/to/your/config.json.

The two-digit stateCode for your Finanzamt is published by ELSTER — look up the current value in the official ELSTER documentation.

Reverse-Charge supplier list

Add your §13b UStG suppliers under ustva.reverseChargeSuppliers in config.json. Patterns are case-insensitive regexes matched against the voucher's contactName or description. Example entry:

{ "pattern": "your-supplier\\s+ireland", "region": "EU", "name": "Your Supplier Ireland" }

Use with Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "elster": {
      "command": "node",
      "args": ["/absolute/path/to/elster-mcp-server/dist/index.js"],
      "env": {
        "ELSTER_CONFIG_PATH": "/absolute/path/to/elster-mcp-server/config.json"
      }
    }
  }
}

See examples/claude_desktop_config.json for the template.

Use with any MCP client

Run the server in stdio mode:

node dist/index.js

Then connect via your client's MCP transport.

Typical UStVA flow

1. elster_login_test                          → { ok: true }
2. elster_kennziffern_list                    → reference for valid codes
3. elster_ustva_start({                       → { sessionId: "ustva-..." }
     year: 2026,
     period: "Q1",
     report: { "81": 12000, "86": 300, "66": 1845.30 }
   })
4. elster_session_status({ sessionId })       → poll until status == AWAITING_CONFIRM
   (open the screenshot at screenshotPath to verify)
5. elster_ustva_confirm({ sessionId })        → { success: true, ticket: "..." }

Typical EÜR flow

1. elster_login_test
2. elster_eur_start({
     year: 2025,
     data: {
       betriebseinnahmen: 50000,
       fahrzeugkosten: 1200,
       afa: 800,
       homeOffice: 1260
     }
   })
3. elster_session_status (poll until SAVED or AWAITING_REVIEW)
4. open the ELSTER portal in your browser → "Meine Formulare" → review the draft → submit manually

Datenübernahme (Vorjahresdaten übernehmen)

ELSTER can copy an earlier submission of the same form into a new one — the "Datenübernahme" step it offers right after you pick the tax year. The three *_start tools expose it through an optional takeover argument:

takeover

Effect

omitted / "none"

Start from a blank form (default)

"latest"

Carry over the most recently sent submission

2024

Carry over that tax year's submission

"537842781"

Carry over that exact aufgabeId

Check what is on offer first — this never fills or submits anything:

elster_datenuebernahme_list({ form: "est", year: 2025 })
→ {
    "candidates": [
      { "aufgabeId": "537842781",
        "description": "ESt unbeschränkt (ESt 1 A) 2024, …",
        "sentAt": "31.07.2025 23:32 Uhr", "year": 2024 },
      …
    ]
  }

elster_est_start({ year: 2025, data: {}, takeover: 2024 })

If the requested submission is not on offer the call fails instead of quietly starting a blank form — carrying nothing over when you asked for last year's data would produce a wrong return. An empty candidate list is normal: ELSTER only offers a takeover once a matching form has actually been submitted.

Past submissions as context

elster_submission_protocol reads the "Übertragungsprotokoll" of returns you have already filed — every field that was actually transmitted, with its Zeile number, label, value and ELSTER field id, grouped by form and section:

elster_submissions_list({ formFilter: "ESt unbeschränkt" })
→ [{ nachrichtId: { type: "a", id: 178614775 }, year: 2024,
     ordnungskriterium: "…", aufgabeId: "537842781", abgabeXmlId: 51359059 }, …]

elster_submission_protocol({ formFilter: "ESt unbeschränkt", years: [2024] })
→ { title: "Übertragungsprotokoll …",
    meta: { Finanzamt, Transferticket, "Eingang auf Server", … },
    sections: [ { form: "Anlage N (…)",
                  heading: ["Werbungskosten", "Aufwendungen für Arbeitsmittel"],
                  rows: [ { zeile: "5", label: "…", value: "…",
                            fieldId: "id-N-ArbL-…-E0200204_usb1_1-1-1-1" } ] } ] }

The fieldId is ELSTER's own identifier, which is exactly what elster_est_start's data keys are matched against — so last year's protocol can be read, adjusted and fed back into this year's form.

Note that this returns real personal tax data (identification number, bank details, income). It is read-only and never leaves your machine, but treat the output like the tax return it is.

Security notes

  • Never commit your .env, config.json, or .pfx. They are gitignored by default.

  • The certificate password is read from env / config and passed to Puppeteer — make sure the host running this server is trusted.

  • Set ELSTER_HEADLESS=false once to watch the first run and confirm everything is wired correctly.

Limitations

  • The ELSTER portal selectors can change. If a flow breaks, run with ELSTER_HEADLESS=false and check the screenshots written to ./screenshots/.

  • The ESt tool is intentionally a thin wrapper — German income-tax forms (Anlage G, V, N, S, KAP …) are dozens of different forms with thousands of fields. This server provides the framework (login, open, fill-by-label-or-id, Prüfen) and leaves the field choices to you.

  • No XML submission path. Official programmatic submission requires the ERiC library (registration as a software vendor). This server uses the same Online-Formular path that any taxpayer uses.

License

MIT

Contributing

PRs welcome. The most useful additions are:

  1. More robust selectors for changed ELSTER pages

  2. Pre-filled Anlage G / V / N / S templates for ESt

  3. A typed report schema validator for elster_ustva_*

When opening an issue, please run with ELSTER_HEADLESS=false and attach the screenshot under ./screenshots/ that shows the failure.

Available Tools

32 tools
elster_belege_listA

Lists receipts in "Meine Belege" (id, label, year, Belegart, recognised values, status). Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNoOnly receipts for this tax year.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the key trait ('Read-only') plus a summary of returned fields. It omits pagination behavior, result limits, and any auth/session prerequisites, which for a list tool with no annotations leaves meaningful gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the verb and scope front-loaded and the field list appended; every clause earns its place and there is no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list with one optional parameter and no output schema, the description supplies the returned fields it would otherwise be missing. It could add pagination or scope notes, but nothing essential to invoking it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'year' parameter is documented in the schema as 'Only receipts for this tax year.' The description adds no filter syntax or semantics beyond that, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Lists receipts in Meine Belege' and enumerates the returned fields (id, label, year, Belegart, recognised values, status). It is clearly distinguishable from the write-side sibling elster_beleg_upload, but does not explicitly name or differentiate itself from the other list tools (submissions_list, drafts_list).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the resource ('Meine Belege' receipts, optional tax-year filter), so an agent can infer when it applies. However, there is no explicit when/when-not guidance and no routing to sibling list tools, leaving the selection context to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_beleg_uploadA

Uploads one receipt file (PDF, PNG or JPG, max 10 MB) to "Meine Belege" and tags it with a Belegart and its values, so it is linked to the matching form line. Belegart is a path like "N/Arbeitsmittel" (fields Art_der_Arbeitsmittel, Betrag), "N/Weitere_Wk/Sonst" (Bezeichnung, Betrag), "N/Fortb" (Bezeichnung, Betrag), "N/Dienstreise". Betrag as number or "12,99". Stores a document in the account; submits nothing. Receipts are optional for the ESt (Belegvorhaltepflicht) but can avoid queries.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute path to the receipt.
idNrNoSteuer-IdNr of the person; default: first person of the account.
yearYes
labelNoDisplay name; default: file name.
fieldsYes
belegartYes
stichwoerterNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses the accepted formats, the 10 MB cap, and crucially that it 'Stores a document in the account; submits nothing' – an important side-effect distinction given the sibling submission tools. However it omits session/auth requirements, duplicate-handling, and what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and constraints, then gives the belegart mapping examples. The example list is dense but earns its place by encoding the nested-key contract; nothing is gratuitously verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with a nested object and no output schema, the description covers the central belegart/fields mapping well but leaves required 'year' and 'stichwoerter' unexplained, and never states what the call returns or whether a login is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43% (low), so the description must compensate. It does meaningfully, explaining the belegart path syntax ('N/Arbeitsmittel') and which field keys map to each line (Art_der_Arbeitsmittel, Betrag, etc.) plus the Betrag numeric format. It does not clarify year, stichwoerter, or label beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Uploads one receipt file') plus the destination ('Meine Belege'), and its tagging purpose. An agent can distinguish it from the reading-oriented elster_belege_list sibling without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives usage context ('Receipts are optional for the ESt (Belegvorhaltepflicht) but can avoid queries'), which tells the agent why and when this tool is worth calling. It does not name alternative tools or explicit exclusions, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_config_showA

Shows the currently loaded ELSTER configuration (with secrets redacted) so you can verify env vars / config.json were picked up.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait: secrets are redacted, which is important for an agent to know. Without annotations, this adds value. However, it does not detail whether the tool reads from files, environment, or cache, but for a display tool this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence that is perfectly concise with no wasted words. It front-loads the action and adds necessary context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple display tool with no parameters and no output schema, the description fully covers the purpose, output behavior (redacted secrets), and usage context (verification of env vars/config).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so baseline is 4. The description adds no parameter info, but none is needed. It provides context about what is shown (config, secrets redacted) beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool shows the currently loaded ELSTER configuration with secrets redacted, specifying the action (show), resource (config), and a key detail (redaction). It distinguishes from sibling tools which are action-oriented (start, list, generate, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used to verify environment variables or config.json were picked up, but does not explicitly state when not to use it or mention alternatives. The purpose is clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_datenuebernahme_listA

Lists the earlier submissions ELSTER offers to carry over ("Datenübernahme") for a given form and tax year, without starting a filling session. Read-only. Feed an aufgabeId or year from the result into the "takeover" argument of elster_ustva_start / elster_eur_start / elster_est_start.

ParametersJSON Schema
NameRequiredDescriptionDefault
formYesWhich form to inspect.
yearYesTax year of the NEW form you intend to file.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries the full burden. It usefully states 'Read-only' and 'without starting a filling session', covering the main behavioral trait. However, it omits details on permissions, rate limits, or other side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scoping constraint. Every sentence earns its place by guiding the agent to the downstream usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists. The description implies the return contains an 'aufgabeId' or year, which is helpful, but does not fully describe the list's structure or pagination behavior. Adequate for a simple two-parameter list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds the context that 'form' selects which form to inspect and 'year' is the tax year of the new form, but provides no syntax or format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lists), resource (earlier submissions ELSTER offers to carry over), and scope (for a given form and tax year). Explicitly distinguishes from the start tools by saying 'without starting a filling session', allowing an agent to identify it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: use this to find carry-over data before starting a filling session. Names the downstream tools (elster_ustva_start / elster_eur_start / elster_est_start) and the argument ('takeover') to pass results into. No explicit when-not-to-use, but the alternative path is evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_drafts_listA

Lists saved drafts ("Meine Formulare → Entwürfe") with their aufgabeId, newest first. Uses the persistent engine session (logs in on first use, then stays logged in).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add real behavioral context: it discloses the persistent engine session, including that it logs in on first use and then stays logged in, which tells the agent about a side effect and session state. It omits failure modes and what happens on session expiry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool returns and scoped by the UI path before the session detail. No padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, no-output-schema read tool, the description covers scope, a returned identifier, ordering and session behavior, which is most of what an agent needs. It does not enumerate the full shape of a draft record, leaving some return-value uncertainty with no output schema to fall back on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4. The description still adds value by naming a returned field (aufgabeId) and the sort order (newest first), which is more than the empty schema conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists saved drafts') and pins down scope with the UI path 'Meine Formulare → Entwürfe', which separates it from submission- and Beleg-listing siblings. It does not explicitly name an alternative, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the read path for retrieving draft forms rather than submissions, but the description never states when to use it, when not to, or which sibling to prefer (e.g. elster_submissions_list). Adequate but with a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_edaten_fetchA

Retrieves the pre-filled tax data ("vorausgefüllte Steuererklärung" / eDaten) the tax authority already holds for a year: Lohnsteuerbescheinigung, Vorsorgeaufwendungen, Lohnersatzleistungen, Riester/Rürup. ELSTER only exposes these from inside the ESt form, so this walks the form up to the import step and reads the values back. The form is left without saving, so no draft is kept (ELSTER may offer it for recovery at the next login). Nothing is ever transmitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearYesTax year to retrieve.
anlagenNoCheckbox ids on the Anlagenauswahl page. Defaults to Hauptvordruck + Anlage N + Anlage Vorsorgeaufwand, which is what eDaten fills. Exact ids, e.g. "VAnlageN", "VAnlageG", "VAnlageKAP".

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that the form is left unsaved, no draft is kept, ELSTER may offer it for recovery at next login, and nothing is ever transmitted. It still doesn't state session/login prerequisites or expected runtime, so it is not fully complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and its German name, then the mechanism, then the side-effect caveats. Every sentence earns its place and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read operation with no output schema and no annotations, the description covers purpose, mechanism, side effects, and the categories of data returned. It does not sketch the exact return shape or error conditions, but it is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (year, anlagen) are already documented in the schema, including the default anlagen set and id examples. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (retrieves), a precisely named resource (pre-filled tax data / vorausgefüllte Steuererklärung / eDaten), and enumerates the exact data categories returned (Lohnsteuerbescheinigung, Vorsorgeaufwendungen, Lohnersatzleistungen, Riester/Rürup). It is clearly distinguishable from siblings like elster_drafts_list or elster_datenuebernahme_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains when this tool is needed by describing the constraint that forces it: ELSTER only exposes eDaten from inside the ESt form, so the tool walks the form to the import step. It gives clear context but never explicitly names an alternative tool or a when-not-to-use condition, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_est_startA

Starts an ESt 1 A (Einkommensteuererklärung) form-prep session. Opens the form, fills taxpayer basics from config + any extra fields you provide (by ELSTER input id/name hint), runs Prüfung, then waits 30 min for you to review in the portal. NEVER submits.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataNoOptional map of field-id hints → values. Each key is matched against ELSTER input id/name as substring. Use empty {} to only fill taxpayer basics from config.
yearYes
takeoverNoOptional "Datenübernahme" — carry the data of an earlier submission of this form into the new one. "none" (default) fills a blank form; "latest" takes the most recently sent one; a tax year (e.g. 2024) takes that year's submission; any other digit string is treated as an explicit aufgabeId. Fails loudly if the requested submission is not offered — list what is available with elster_datenuebernahme_list.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses the session lifecycle, that it runs Prüfung, waits 30 minutes for review, and emphatically never submits. It also mentions failing loudly when a requested takeover submission is unavailable. It omits auth requirements, what happens after the 30-minute window, or how to cancel a lingering session.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tightly constructed sentences front-loading purpose, then workflow and the critical 'NEVER submits' constraint. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description adequately covers the session-start workflow, the 30-minute review window, and the never-submit guarantee. It leaves some gaps around post-timeout behavior and how to check or cancel the session, but the core information an agent needs to call the tool correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with the required 'year' parameter having no description in either the schema or the tool description. The description restates what the schema already says about the data parameter (input id/name hints) and adds little for takeover. It does not compensate for the undocumented year parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Starts an ESt 1 A form-prep session'), and names the form type (ESt 1 A) which distinguishes it from sibling start tools like elster_ustva_start and elster_eur_start. An agent can identify the correct tool without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly outlines the workflow context (opens form, fills basics, runs Prüfung, waits 30 min for portal review) and the alternative for listing available Datenübernahme submissions within a parameter. However, it does not explicitly state when to use this tool versus other form-preparation or session tools, nor does it set exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_eur_startA

Starts an EÜR (Anlage Einnahmen-Überschuss-Rechnung) form-prep session. Fills the form up to Prüfung, then tries to "Speichern und Verlassen" so the draft survives in ELSTER. NEVER submits.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataYesMap of field names to numeric amounts. Supported fields: betriebseinnahmen, kfzPrivatNutzung, fahrzeugkosten, kfzSteuer, telekommunikation, versicherungen, bewirtung, reisekosten, bankgebuehren, fremdleistungen, software, buchfuehrung, beratung, werbung, gwg, steuern, uebrigeBA, afa, homeOffice, iabAbzug.
yearYes
takeoverNoOptional "Datenübernahme" — carry the data of an earlier submission of this form into the new one. "none" (default) fills a blank form; "latest" takes the most recently sent one; a tax year (e.g. 2024) takes that year's submission; any other digit string is treated as an explicit aufgabeId. Fails loudly if the requested submission is not offered — list what is available with elster_datenuebernahme_list.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the critical behavior: it stops at Prüfung, attempts 'Speichern und Verlassen' so a draft survives, and explicitly never submits. That is strong safety-relevant disclosure, though it omits auth requirements and failure handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, led by the core purpose, then the mechanism, then the hard constraint. No filler and no repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations exist, so the description must orient the agent on a mutating setup tool. It covers the central action and the no-submit guarantee well; only the return value (e.g. session handle) and follow-up routing are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% and the schema already documents data (with the full field list) and takeover in depth. The description adds nothing about year, data, or takeover semantics, so the baseline 3 applies: the schema, not the description, does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (starts) and resource (an EÜR form-prep session), expanding the acronym and describing the exact workflow outcome: fills the form up to Prüfung, then saves the draft. This clearly separates it from parallel siblings like elster_ustva_start and elster_est_start, which target different form types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its place in a workflow (start a session, become a surviving draft) and the 'NEVER submits' constraint bounds its role, but it names no explicit alternative or precondition, and never routes the agent to companion tools such as elster_form_set or elster_submissions_list. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_add_rowA

Adds one row to an inline repeat group ("Mzb", e.g. Arbeitsmittel "AufwendungenArbeitsmittel") by filling its new-row template and committing it ("Eintrag übernehmen"). Group names and template keys come from elster_form_page → groups. Same key and money rules as elster_form_set.

ParametersJSON Schema
NameRequiredDescriptionDefault
ridNo
groupYes
valuesYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load, and it does disclose meaningful behavior: it fills the new-row template and COMMITS the row ('Eintrag übernehmen'), signalling a persistent mutation rather than a local draft. However, it omits permission/auth needs, what happens on duplicate rows or missing groups, and any return/error behavior, leaving gaps for a mutation tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action ('Adds one row...'), then the mechanism, then the input sourcing. Dense and mostly earns its space, though the parenthetical German/English pairs add minor noise that a purely normative reader doesn't need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, nested-object, no-output-schema, no-annotation tool, the description covers the commit semantics and input sourcing but leaves 'rid' unaddressed and says nothing about failure modes or persistence details. Adequate but not complete for the complexity involved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It adds real meaning for 'group' (an inline repeat group / Mzb, with an example) and implies 'values' follows the same key and money rules as elster_form_set, but it never explains the shape of 'values' or the 'rid' parameter at all, leaving two of three parameters under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource: 'Adds one row to an inline repeat group', with a concrete domain example (Mzb / Arbeitsmittel 'AufwendungenArbeitsmittel'). It is clearly distinguishable from siblings like elster_form_delete_row, elster_form_set, and elster_form_save by naming the exact mutation it performs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent where inputs come from ('Group names and template keys come from elster_form_page → groups') and routes validation/format rules to elster_form_set. It stops short of stating explicit when-not-to-use conditions or contrasting with elster_form_set as an alternative action, so it's clear context rather than full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_checkA

Runs ELSTER's "Prüfen" on the open form. Returns ok, the headline (for ESt it includes the provisional "Erstattung/Nachzahlung"), and the error panel with the causing pages. Never sends.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers the key behavioral fact that the operation is non-mutating ('Never sends') plus the shape of the return value (ok, headline, error panel with causing pages). It omits any mention of auth/session requirements or latency, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed clauses covering action, return shape, and the critical no-send guarantee, with no filler and the primary action front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, the description usefully enumerates what comes back and flags the non-mutating behavior, which is exactly what an agent needs. It could add a note about what 'ok' vs an error panel implies for next steps, but it is otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so there is nothing to disambiguate; baseline for a no-parameter tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Runs ELSTER's "Prüfen" on the open form'), which is unambiguous and clearly distinct from siblings like elster_form_save, elster_form_press, or the submission tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'on the open form' implies a prerequisite (a form must already be open), and 'Never sends' hints this is the pre-submission validation step. However, it never explicitly says when to use this versus elster_form_save, elster_form_review, or elster_form_press, leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_crawlA

Walks every page below a RID (breadth-first, read-only) and returns each page's fields and repeat groups. Use it once per form/Anlage to learn which Kennzahl lives on which page.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootRidYes
maxPagesNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose read-only status, breadth-first ordering, and that it is intended as a one-time discovery pass. It omits cost/truncation behavior (what happens at the default maxPages=60), pagination, and any error handling for an operationally heavy crawl.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and finishing with the usage cue. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description must cover behavior and returns. It names what is returned (fields, repeat groups) but leaves the maxPages cap and truncation semantics unexplained for what is a potentially expensive crawl.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither parameter is documented in the schema. 'Below a RID' gives partial meaning to rootRid, but maxPages (default 60) is never explained — an agent cannot tell what the page cap means or how it affects results.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Walks every page below a RID') plus the traversal order (breadth-first, read-only) and the return content (fields and repeat groups). It implicitly contrasts with single-page siblings like elster_form_page by emphasizing 'every page', but it never names a sibling, so it stops short of full discrimination.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it once per form/Anlage to learn which Kennzahl lives on which page' gives a clear context and even a quantity hint ('once'). However, it does not name alternatives (e.g., elster_kennziffern_list, elster_form_page) or state when not to crawl.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_delete_rowC

Deletes row index (0-based) of an inline repeat group on a page.

ParametersJSON Schema
NameRequiredDescriptionDefault
ridNo
groupYes
indexYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a destructive mutation but does not disclose whether deletion is reversible, whether subsequent row indices shift, permission requirements, or what is returned. The '0-based' note is the only behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is tight, though the brevity edges toward under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter mutation tool with no annotations, no output schema, and zero schema description coverage, the description is too thin. Key details an agent needs to call it correctly, such as the meaning of `rid` and post-delete behavior, are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must document all three parameters. It explains `index` ('0-based') and hints at `group`, but leaves `rid` entirely unexplained, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Deletes') and resource ('row index of an inline repeat group'), which is enough to distinguish it from most siblings. However, it does not explicitly contrast with the obvious sibling elster_form_add_row or clarify its role in the form-editing workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are given. The description never mentions the related add/set page tools or when deleting a repeat-group row is appropriate versus other form mutations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_newA

Starts a NEW form: picks the year, handles Datenübernahme (none by default), Anlagenauswahl and the eDaten import, and stops on the form's Startseite. Returns the page like elster_form_page. Follow with elster_form_save to persist it as a draft.

ParametersJSON Schema
NameRequiredDescriptionDefault
formYesPortal slug from /eportal/formulare-leistungen/alleformulare/<slug>, e.g. "est", "euer", "ustvaeru".
yearYes
anlagenNoAnlagen checkbox values to select, e.g. ["VAnlageEUER"]. Omit to keep ELSTER's defaults.
takeoverNoaufgabeId of an earlier submission to carry over (see elster_datenuebernahme_list). Omit for none.
importEdatenNoAccept the eDaten (vorausgefüllte Steuererklärung) import if offered.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: Datenübernahme is 'none by default', the run halts at the form's Startseite (i.e., it does not persist or submit), and the return mirrors elster_form_page. It still omits session/auth prerequisites and what failure states look like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and ending on the next-step routing. Zero filler; every clause conveys new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-step initialization tool with no output schema and no annotations, the description covers the main stages, the stopping point, the return type by reference, and the required follow-up. Only the authentication/session preconditions are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already explains form, anlagen, takeover, and importEdaten. The description only reinforces defaults ('none by default') and does not add format or syntax detail beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Starts a NEW form') and enumerates the concrete flow it performs (year, Datenübernahme, Anlagenauswahl, eDaten import). The capitalized NEW implicitly separates it from elster_form_open, but it never names that sibling, so differentiation relies on inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear forward guidance ('Follow with elster_form_save to persist it as a draft') and notes defaults, but says nothing about when to choose this over the other start tools (elster_form_open, elster_est_start, elster_ustva_start) or which situations it is not suited for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_openA

Opens a saved draft in the engine session and returns its Startseite (fields, repeat groups, navigation RIDs). Without aufgabeId the newest draft is opened. If ELSTER still holds the form open from an earlier call, it is re-entered instead of failing. Only one form can be open at a time.

ParametersJSON Schema
NameRequiredDescriptionDefault
aufgabeIdNoFrom elster_drafts_list.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses several important behaviors: it returns the Startseite contents, falls back to the newest draft, re-enters an already-open form instead of failing, and enforces a one-form-open constraint. It does not cover error handling, permissions, or session prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences with the core action and return value front-loaded, followed by fallback and state-management details. Every sentence adds distinct information and none is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no annotations and no output schema, the description is nearly complete: it explains the action, the optional-parameter fallback, re-entry behavior, and the one-open-form constraint, and it sketches the return payload. Minor gaps remain around error cases and session prerequisites, but nothing critical for calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema by explaining what happens when aufgabeId is omitted (the newest draft is opened), which is not evident from the parameter definition alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Opens') and resource ('saved draft in the engine session'), and even summarizes the return payload (Startseite fields, repeat groups, navigation RIDs). It implicitly distinguishes itself from elster_form_new (new form) and elster_drafts_list (listing drafts), but does not explicitly name any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for the optional aufgabeId parameter (newest draft opened when omitted) and explains the re-entry behavior when a form is already open. It does not explicitly state when to prefer this tool over alternatives like elster_form_new, but the saved-draft focus and the parameter description 'From elster_drafts_list' make the intended usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_pageA

Reads a page of the open form. With rid, jumps there first (RIDs look like "FormData://est-2025-v1/Startseite[0]/MAVSAnlageN[0]/VAnlageN[0]/HomeofficePauschale[0]" and come from the nav list of any page). Returns plain fields (name, kennzahl, label, value, options), repeat groups (committed rows, the template fields for a new row, sub-page RIDs), non-navigation commands, validation errors and the nav RIDs visible from this page. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
ridNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the key safety trait ('Read-only.') and describes the return shape in useful detail, but omits pagination semantics, authentication/prerequisite state, and any limits or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the RID mechanic, then the return contents in a single dense paragraph. Every clause carries information, though the parenthetical RID example is long and slightly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does the heavy lifting by enumerating what is returned (plain fields, repeat groups, non-navigation commands, validation errors, nav RIDs), which is sufficient for an agent to invoke and interpret the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single rid parameter is undocumented in the schema, but the description compensates well: it gives a concrete RID example, explains the format ('FormData://...'), notes the source ('the nav list of any page'), and implies rid is optional ('jumps there first').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Reads a page of the open form,' and elaborates on the exact payload (fields, repeat groups, commands, validation errors, nav RIDs). It does not explicitly differentiate itself from siblings like elster_form_crawl or elster_form_check, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage ('a page of the open form') and explains that passing rid jumps there first, but never states when to prefer this tool over elster_form_crawl, elster_form_check, or elster_form_review. Usage is inferable rather than explicit, and no when-not conditions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_pressB

Escape hatch: presses a button by id or posts a raw reqCmd JSON from elster_form_page → commands (e.g. EditDetachedMzbSemIndex to add an Anlage for a person, FillInProfile, AddMzbItem). Sending, deleting drafts and logging out are refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
ridNoJump here first.
commandNoreqCmd JSON (string or object).
buttonIdNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the refusal policy (no sending, draft deletion, or logout) and that it mutates form state via commands. It omits auth requirements, idempotency, error behavior, and what the response looks like for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, but it is front-loaded with the core concept, then examples, then refusals. Each clause earns its place, with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-annotation, no-output-schema tool, the description covers identity, examples, and exclusions, but leaves gaps on return behavior, error handling, and when this outranks the sibling form tools. Adequate but not fully complete for a low-level escape hatch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so the schema documents most parameters. The description adds real value by giving concrete command examples (AddMzbItem, FillInProfile) and clarifying buttonId vs. raw command usage. The 'rid' parameter ('Jump here first') remains opaque in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('presses a button by id or posts a raw reqCmd JSON'), and the 'escape hatch' framing distinguishes it from structured siblings like elster_form_set or elster_form_add_row. The purpose is clear, though the description leans on internal jargon (reqCmd, EditDetachedMzbSemIndex) that assumes domain knowledge.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Escape hatch' label implies this is for cases the normal form tools can't handle, and the refusal list ('Sending, deleting drafts and logging out are refused') gives explicit when-not boundaries. However, it never names which sibling to prefer for standard edits, leaving the routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_reviewA

Final check before the user submits: runs "Prüfen" and, only if it is clean, reads ELSTER's "Formular absenden" overview — exactly the data that would be transmitted, including eDaten fields — as sections of rows (Zeile, label, value, Kennzahl, source). Returns to edit mode in the same call. Never transmits: "Absenden" stays refused, and no command but EINGABE/PRUEFEN is accepted while on the overview.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: state-machine behavior (returns to edit mode in the same call), safety guarantees ("Absenden" stays refused, never transmits), and accepted-command constraints (only EINGABE/PRUEFEN while on the overview). It even discloses the return shape, which is far beyond what any structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the pre-submit context are front-loaded, and the dense sentence still avoids filler. It is lengthy and information-packed rather than trimmed, but every clause adds a distinct behavioral fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing the returned structure as sections of rows with (Zeile, label, value, Kennzahl, source). Combined with the conditional flow and safety constraints, an agent has everything needed to invoke and interpret this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline there is nothing to document; the schema is empty and closed. No parameter meaning is needed or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific composite action: runs "Prüfen", then conditionally reads ELSTER's "Formular absenden" overview, exposing exactly what would be transmitted. This clearly separates it from the submission tools in the sibling set (e.g. elster_form_press) by declaring it never transmits, so an agent can tell its role apart without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: this is the final check performed before the user submits, and the overview read is gated on "Prüfen" being clean. It does not explicitly name an alternative tool (e.g. elster_form_check) or spell out when to prefer that instead, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_saveA

Saves the open form as a draft ("Speichern und Formular verlassen") and closes it. Reopen with elster_form_open. The server session times out after ~30 min idle and unsaved work is then only in ELSTER's auto-recovery — save before long pauses.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the form is saved as a draft, that it closes the form, and that the server session times out after ~30 minutes idle with unsaved work left only in ELSTER's auto-recovery. This is meaningful behavioral context beyond a bare mutation statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and immediately followed by the reopen path and the time-sensitive warning. Every sentence adds distinct value, and nothing is repeated or wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-annotation, no-output-schema tool, the description covers the essential behavior: what it does, how to reopen, and the session timeout risk. It does not mention return value or error handling, but those are secondary for a simple save-and-close operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document. Per the rubric, a 0-parameter tool earns a baseline of 4; the empty schema is fully consistent with the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (saves) and resource (open form as a draft) and notes that it closes the form. It also distinguishes the follow-up action by naming elster_form_open for reopening, so an agent can separate it from open/new tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: use it to save the open form before long pauses, and reopen with elster_form_open. It does not explicitly state when not to use it or name alternatives like elster_form_new, but the intended usage is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_form_setA

Sets plain fields on a page and saves them to the server-side form. Keys are a Kennzahl ("E0204507") or the field name ("eruNWkHomeofficeE0204507"); a key must match exactly one field on the page. Checkboxes take true/false, radios and selects take the option value. Money: fields labelled "(Euro)" accept whole euros only (no decimal separator — round expenses up, income down); "(Euro, Cent)" fields take "3,36". Validation errors come back in errors (HTTP is always 200). Sub-pages of detached repeat groups (e.g. ".../MZBErsteTaetigkeitsstaette[0]") are set with this tool too; jumping to index [n] of such a group creates entry n.

ParametersJSON Schema
NameRequiredDescriptionDefault
ridNoPage to set fields on. Omit for the current page.
valuesYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it discloses that validation errors return in `errors` while HTTP is always 200, defines exact value semantics per widget type, and covers money rounding/format rules. It omits auth/permission needs and whether a partial set is atomic or what happens to untouched fields, but the error-handling disclosure is unusually valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries operational content: purpose, key format, widget value rules, money formatting, error channel, and repeat-group edge case. It is dense but front-loaded and free of filler, and the money rules genuinely change how an agent formats input.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no annotations and no output schema, the description covers keys, value types, money formatting, error surface, and nested repeat-group sub-pages. It stops short of stating whether the whole form saves or only supplied fields, and there is no success-response detail, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (only `rid` is documented in the schema). The description compensates heavily for the undocumented `values` object by explaining that keys are a Kennzahl or field name, must match exactly one field, and what value types each widget expects. It materially exceeds the schema for the most important parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: setting plain fields on a page and saving them server-side. It distinguishes itself by covering plain fields, checkboxes, radios, selects, and detached repeat-group sub-pages, which implicitly separates it from siblings like elster_form_add_row or elster_form_press. However, it never names an alternative tool, so sibling differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives strong how-to guidance (key formats, checkbox/radio value conventions, repeat-group sub-page handling), but says nothing about when to reach for this tool versus elster_form_save, elster_form_press, or elster_form_add_row. Usage is strongly implied by scope statements rather than explicitly routed, landing at minimum-viable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_kennziffern_listA

Returns the list of supported UStVA Kennziffern (codes 81, 86, 66 etc.) with descriptions and whether they are NET (base amount) or TAX (tax amount).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must carry the burden. It states the return content but does not disclose side effects, authentication requirements, or idempotency. Minimal but sufficient for a simple list query.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, includes specific examples and categorization. No extraneous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, description adequately explains what is returned. Lacks mention of prerequisites like authentication, but for a simple list of codes it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters, so baseline is 4 per guidelines. Description adds no parameter info because none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states verb 'returns' and resource 'list of supported UStVA Kennziffern', includes examples and type distinction (NET/TAX). No sibling tool serves a similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance. Usage is implied by the tool's name and description, but no alternatives or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_login_testA

Verifies that the configured certificate + password can log into the ELSTER portal. Returns success and final URL or an error. Use this once before submitting anything.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description must disclose behavioral traits. It states returns success/final URL or error, which covers basic behavior. However, could provide more detail on side effects (likely none) or error format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose. Every sentence provides value: first states what it does, second gives usage guidance. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (no params, no output schema), the description covers purpose, return behavior, and usage guidance adequately. A bit more detail on error cases could improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, and schema coverage is trivially 100%. The description adds no param info, but with zero params, the baseline is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies login ability with certificate+password, specifying both verb and resource. It distinguishes from siblings like elster_ustva_start which are for actual submissions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use this once before submitting anything', providing clear context for when to use. Does not mention exclusions but given the tool's simplicity, the guidance is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_session_cancelC

Cancels a running session (closes the browser, marks status as CANCELLED).

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the main effects (closes browser, marks cancelled) but omits important details like irreversibility, required permissions, or potential side effects. The transparency is adequate for a simple operation but could be more explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, delivering the core action in a single sentence. Every word adds value, with no filler. It earns a 4 for efficiency, though a bit more context would not harm conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (simple), lack of output schema, and absent annotations, the description covers only the basic action. It does not address when to cancel, whether the session must be active, or what happens to other operations. The completeness is minimal for an agent to use it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description adds no meaning to the 'sessionId' parameter beyond what the name implies. The parameter's purpose is obvious from context, but the tool description should clarify its role (e.g., which session, as the identifier).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('cancels') and the resource ('running session'), and adds specific behavioral details like closing the browser and marking status as CANCELLED. However, it does not explicitly distinguish from sibling session management tools, though the difference is implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus related tools like elster_session_list or elster_session_status. The description lacks context for prerequisites, such as requiring an active session, or when cancellation is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_session_listA

Lists all currently tracked sessions (USTVA / EUR / EST / SYNC) with their status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must be transparent. It states sessions are listed with status, but does not disclose if the tool is read-only, requires authentication, or has any side effects. Adequate for a simple list, but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, efficient, front-loads purpose and scope. Could be slightly more informative about what 'status' entails, but no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description is mostly complete for a simple list tool. It mentions session types and status, though it could clarify the meaning of status or mention any limits. Still adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters, schema coverage 100% => baseline 4. Description does not need to add param info, so score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'Lists' and resource 'currently tracked sessions' with enumerated session types and status, making purpose immediately clear and distinguishing from sibling tools like elster_session_status which likely focus on single sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. The description implies it's for an overview, but does not state exclusions or conditions such as 'use this to get a list before canceling a session'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_session_statusB

Returns the current status, progress log, and any screenshot path for a session started via elster_ustva_start / elster_eur_start / elster_est_start.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. Only lists outputs (status, progress log, screenshot path) without disclosing whether the tool is read-only, requires authentication, or what happens on invalid sessionId. Minimal behavioral insight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with verb and outputs, zero extraneous words. Ideal conciseness for a simple status-check tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool is simple (1 param, no output schema), but description omits return value details (e.g., status codes, progress log structure) and does not compensate for missing schema documentation. Adequate but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with no description for sessionId. Description adds no meaning beyond the schema, failing to clarify format, constraints, or purpose of the single parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'Returns' and names three resources: status, progress log, screenshot path. It explicitly links to sibling start tools (elster_ustva_start, etc.), distinguishing from listing or cancellation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage after starting a session via named tools, but lacks explicit guidance on when to use versus alternatives like elster_session_list or elster_session_cancel. No when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_submission_protocolA

Reads the "Übertragungsprotokoll" of already-submitted returns — every field that was actually filed, with its Zeile number, label, value and ELSTER field id — so earlier years can be used as context. Select either by nachrichtId (from elster_submissions_list) or by formFilter/years. Read-only; returns real personal tax data.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoCap how many protocols to fetch (each is a round-trip). Default 3 when selecting by filter.
yearsNoRestrict to these tax years.
formFilterNoCase-insensitive regex on the form description, e.g. "ESt unbeschränkt".
nachrichtIdNoOne specific submission, as returned by elster_submissions_list, e.g. { "type": "a", "id": 178614775 }.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it declares 'Read-only' and warns that results contain 'real personal tax data'. The parameter note about each protocol being a round-trip with a default of 3 also discloses cost behavior. Missing only permission/auth specifics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what is read, then the selection modes and safety note. Dense but every clause carries information; only mildly long-winded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, so the description must cover returns and safety, and it does: it describes the returned fields, the sensitivity of the data, and the read-only nature. An agent has enough to invoke it correctly, though pagination/ordering behavior is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds the either/or grouping semantics (nachrichtId OR formFilter/years) that the flat schema does not express, which is meaningful selection guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (reads the Übertragungsprotokoll of already-submitted returns) and enumerates the returned content (Zeile number, label, value, ELSTER field id). It also distinguishes itself from the sibling elster_submissions_list by naming it as the source for nachrichtId.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the two selection modes explicitly — by nachrichtId (obtained from elster_submissions_list) or by formFilter/years — and states the purpose (earlier years as context). No explicit when-not guidance, but the routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_submissions_listA

Lists everything under "Meine Formulare" → "Übermittelte Formulare": what was filed, when, its Ordnungskriterium, plus the aufgabeId to reuse it as a Datenübernahme source and the nachrichtId needed by elster_submission_protocol. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
formFilterNoOptional case-insensitive regex on the form description, e.g. "ESt unbeschränkt" or "Umsatzsteuer".

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the burden and does reasonably well: it declares the operation read-only and describes exactly what the listing returns. It omits pagination/scope limits, permission requirements, or ordering behavior, so it stops short of full disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense, front-loaded sentence that lists scope then return payload then downstream consumers, with no filler. The German UI-path notation is heavy but earns its place as the exact locator.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by enumerating the returned fields (filing, date, Ordnungskriterium, aufgabeId, nachrichtId), and the sole parameter is covered by the schema. Adequate for a simple list tool, though read-only behavior and any result limits are only lightly covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single formFilter parameter is fully documented in the schema with an example. The description adds nothing about the filter, so the baseline of 3 applies — the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (lists filed submissions under 'Meine Formulare' → 'Übermittelte Formulare') and enumerates the returned fields, so an agent can distinguish it from siblings like elster_drafts_list or elster_belege_list without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains downstream use cases explicitly — the aufgabeId is for reuse as a Datenübernahme source and the nachrichtId is consumed by elster_submission_protocol — giving clear context. It does not, however, say when to prefer this over elster_drafts_list or state any exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_sync_historyC

Reads "Übermittelte Formulare" (transmission history) from ELSTER. Optionally downloads PDFs.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearsNo
downloadPdfsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must carry full burden. It indicates a read operation with optional download, but does not disclose side effects, permission requirements, or whether downloading modifies state. Minimal transparency beyond the obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise (one sentence) but sacrifices detail. It is front-loaded with the main action, but could usefully add more information without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 2 parameters, no output schema, and no annotations, the description is incomplete. It does not explain years, return format, or prerequisites (e.g., valid session). Among many sibling tools, more context would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema describes two parameters (years, downloadPdfs) with 0% description coverage. The tool description hints at 'downloads PDFs' but does not explain the years parameter or its purpose. No parameter-level documentation provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it reads transmission history from ELSTER and optionally downloads PDFs. The verb 'reads' and resource 'transmission history' are specific. However, it does not explicitly differentiate from sibling tools like elster_sync_inbox, though the name suggests history vs inbox.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. Does not mention prerequisites, ideal scenarios, or when to avoid using it. Sibling tools are listed but not addressed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_sync_inboxC

Reads ELSTER inbox messages ("Posteingang"). Optionally downloads each message as PDF.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxPagesNo
downloadPdfsNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits but only states two actions: reading and optionally downloading PDFs. It fails to mention authentication requirements, side effects (e.g., state changes from downloads), pagination behavior, or error handling, which are critical for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, properly front-loaded with the primary action. However, it could benefit from slightly more structure to separate reading from downloading, though overall it remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and only two parameters, the description lacks detail on the return format (e.g., message list with metadata), prerequisites, and edge cases. While the tool appears simple, the description omits information needed for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It only hints at 'downloadPdfs' via 'Optionally downloads each message as PDF', but provides no meaning for 'maxPages'. The parameter semantics for half the parameters are entirely unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the verb 'Reads' and the resource 'ELSTER inbox messages', with an optional action 'downloads each message as PDF'. It differentiates from sibling tools which cover config, sessions, forms, etc., making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., for specific inbox filtering or batch operations). The description does not mention prerequisites or when not to use it, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_ustva_confirmA

Confirms submission of a UStVA session that is in AWAITING_CONFIRM state. Triggers the final "Absenden" click. Returns the transmission ticket on success.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It mentions the final action and return value (transmission ticket), but lacks details on error behavior, idempotency, or side effects. The state requirement adds some transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero fluff. The first sentence states the purpose and required state, and the second adds the action and return value. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool, the description covers the core functionality and return value. It does not explain error conditions or idempotency, but given the tool's straightforward nature, it is largely complete. Minor gaps in parameter documentation reduce completeness slightly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter (sessionId) has no description in the schema (0% coverage) and the tool description does not mention it at all. The description fails to add meaning beyond the schema, leaving the agent to guess the expected format or role of sessionId.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool confirms submission of a UStVA session in a specific state (AWAITING_CONFIRM) and triggers the final 'Absenden' click. It identifies the specific resource and action, distinguishing it from sibling tools like elster_ustva_generate_xml or elster_ustva_start.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the required state (AWAITING_CONFIRM), implying when to use. It does not explicitly mention when not to use or provide alternatives, but the context from sibling tools makes the usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_ustva_detect_reverse_chargeC

Tests whether a voucher would be detected as Reverse-Charge (§13b UStG) based on the configured supplier patterns. Returns matched supplier and region (EU / NON_EU), or null.

ParametersJSON Schema
NameRequiredDescriptionDefault
contactNameNo
descriptionNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must cover behavior. It mentions returns but omits side effects, required permissions, error cases, or whether the tool is read-only. The phrase 'based on the configured supplier patterns' hints at external configuration not explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words. It front-loads the action ('Tests whether...') and quickly states output. However, it sacrifices completeness for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the domain-specific tax context, lack of output schema, and zero parameter explanations, the description fails to equip an agent for correct invocation. Essential details about input fields, configuration, and error conditions are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description does not explain the parameters contactName and description. Agents cannot determine what values to provide or how these relate to the voucher being tested. This is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool tests reverse-charge detection based on supplier patterns and specifies the output (matched supplier and region or null). It distinguishes from sibling tools like elster_ustva_confirm or elster_ustva_generate_xml by focusing on detection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs. alternatives. The description only states what it does, without clarifying prerequisites, scenarios, or exclusions. Agents must infer context from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_ustva_generate_xmlA

Generates an ELSTER UStVA XML snapshot for archiving. Does NOT submit (submission goes via elster_ustva_start). Useful for audit trails.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearYesTax year, e.g. 2026
periodYesMonth (1-12) or quarter as "Q1".."Q4"
reportYesMap of Kennziffer → amount in EUR. Keys are the bare 2-3 digit code (e.g. "81", "66"). Net amounts for NET-type codes, tax amounts for TAX-type codes.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of disclosure. It clarifies the non-submitting nature and archiving purpose, which is helpful, but lacks details on side effects, performance, or permissions. It does not contradict any annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, consisting of two sentences that front-load the purpose and key differentiator. Every sentence adds value without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 parameters with a nested object) and lack of output schema, the description adequately explains its primary function and distinction from siblings. However, it could be more complete by briefly mentioning what the output is (e.g., returned XML string or file), but overall it is sufficient for an expert user.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage, so parameters are well described in the schema. The description adds no additional information about the parameters beyond what the schema already provides, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('generates') and resource ('ELSTER UStVA XML snapshot'), clearly stating its archiving purpose and distinguishing itself from the sibling 'elster_ustva_start' by explicitly noting it does NOT submit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool ('for archiving', 'useful for audit trails') and when not to ('does NOT submit'), pointing to the correct sibling tool ('submission goes via elster_ustva_start').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

elster_ustva_startA

Starts a UStVA submission session. Opens a browser, logs in, fills the form, runs Prüfung, then PAUSES at AWAITING_CONFIRM. You must explicitly call elster_ustva_confirm to send. Returns a sessionId — poll status via elster_session_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
yearYes
periodYesMonth (1-12) or "Q1".."Q4"
reportYesMap of Kennziffer → amount in EUR (bare digit keys, e.g. "81":12345.67).
takeoverNoOptional "Datenübernahme" — carry the data of an earlier submission of this form into the new one. "none" (default) fills a blank form; "latest" takes the most recently sent one; a tax year (e.g. 2024) takes that year's submission; any other digit string is treated as an explicit aufgabeId. Fails loudly if the requested submission is not offered — list what is available with elster_datenuebernahme_list.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries full burden. It discloses the full workflow (opens browser, logs in, fills form, runs Prüfung, pauses), the required confirm step, and the return value. It omits error handling, auth details beyond login, and rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences/fragments, front-loaded with the core action and the critical pause constraint. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers the key workflow, the pause/confirm handshake, and polling. It lacks failure-mode guidance, but is otherwise complete for a multi-step session starter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%; the description adds no parameter meaning. The schema already documents period, report, and takeover, but 'year' remains undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Starts a UStVA submission session') and distinguishes from the confirm sibling by describing the pause point. Clearly differentiated from other start tools (eur_start, est_start) via the UStVA qualifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the follow-up call (elster_ustva_confirm) required to send, and the status-polling alternative (elster_session_status). Does not state when to prefer this over other form tools or prerequisites beyond login.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 32 tool updatesv0.1.0
    • First observedelster_beleg_upload
    • First observedelster_belege_list
    • First observedelster_config_show
    • First observedelster_datenuebernahme_list
    • First observedelster_drafts_list
    • First observedelster_edaten_fetch
    • First observedelster_est_start
    • First observedelster_eur_start
    • First observedelster_form_add_row
    • First observedelster_form_check
    • First observedelster_form_crawl
    • First observedelster_form_delete_row
    • First observedelster_form_new
    • First observedelster_form_open
    • First observedelster_form_page
    • First observedelster_form_press
    • First observedelster_form_review
    • First observedelster_form_save
    • First observedelster_form_set
    • First observedelster_kennziffern_list
    • First observedelster_login_test
    • First observedelster_session_cancel
    • First observedelster_session_list
    • First observedelster_session_status
    • First observedelster_submission_protocol
    • First observedelster_submissions_list
    • First observedelster_sync_history
    • First observedelster_sync_inbox
    • First observedelster_ustva_confirm
    • First observedelster_ustva_detect_reverse_charge
    • First observedelster_ustva_generate_xml
    • First observedelster_ustva_start

TDQS

B3.3/5.0

Scored across 32 tools

Disambiguation3/5

Most tools target clearly distinct actions (form_set vs form_add_row vs form_delete_row, ustva_start vs est_start vs eur_start), but several have overlapping read purposes: elster_sync_history and elster_submissions_list both read the same 'Übermittelte Formulare' area, elster_form_check is a subset of elster_form_review, and elster_form_crawl overlaps elster_form_page. Descriptions are detailed enough to mitigate most misselection, but the duplication is real.

Naming Consistency4/5

Names consistently use snake_case with an elster_ prefix and a domain segment (form_, ustva_, session_, sync_), making the pattern predictable. Minor deviations: singular vs plural of the same noun (beleg_upload / belege_list, submission_protocol / submissions_list).

Tool Count3/5

32 tools is on the heavy side even for a genuinely complex domain spanning ESt/EÜR/UStVA forms, browser sessions, receipts, drafts, and sync. Several tools look consolidatable (session_list vs session_status, submissions_list vs sync_history, check vs review), pushing it past a comfortably scoped surface.

Completeness4/5

Coverage is broad: login/config checks, form lifecycle (open/new/page/set/add/delete/save/check/review), UStVA-specific helpers, receipts list/upload, eDaten fetch, drafts, and submission history/protocol. Gaps are minor and partly by design (submission deliberately limited to UStVA; no receipt deletion/update, no explicit logout).

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    MCP server for DACH e-invoicing. Create XRechnung (UBL) and ZUGFeRD 2.3 (Factur-X CII) invoices, validate against EN 16931 rules, extract data from XML, and convert between UBL, CII and JSON formats.
    6
    30 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A local Windows automation MCP server that lets a KI-Agent safely control the SteuerSparErklärung tax software: inventory and open tax cases, read pages, compare receipts and values, and edit verified working copies, with all ELSTER transmission paths locked tight.
    MIT