Skip to main content
Glama

Siemens Docs MCP

An MCP server that lets an AI assistant search and read Siemens documentation live — any publication on docs.tia.siemens.cloud (TIA Portal, STEP 7, WinCC Unified, Openness, …) and docs.industrial-operations-x.siemens.cloud (Industrial Operations X) — plus a CLI that exports a whole publication to a tree of Markdown files.

Built to solve the problem of documentation portals that only offer low-quality PDF exports or JavaScript-rendered web views, making the content difficult to search, reference, or feed to AI tools.


Installation

PyPI

pip install siemens-docs-mcp

This installs two commands: siemens-docs-mcp (the MCP server) and siemens-docs-export (the Markdown exporter). Requires Python 3.11+.

From source

git clone https://github.com/Czarnak/siemens-docs-mcp
cd siemens-docs-mcp
python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -e .
# development (pytest, ruff, pip-audit; needs pip >= 25.1):
pip install -e . --group dev

Related MCP server: Jamf Docs MCP Server

MCP server

The server speaks MCP over stdio. Register it with Claude Code:

# from PyPI, no manual install (needs uv)
claude mcp add siemens-docs -- uvx siemens-docs-mcp

# or an installed copy
claude mcp add siemens-docs -- siemens-docs-mcp
# from a source checkout: <repo>/.venv/Scripts/siemens-docs-mcp  (Linux/macOS: <repo>/.venv/bin/siemens-docs-mcp)

Pass settings with -e, e.g. claude mcp add siemens-docs -e SIEMENS_DOCS_LOCALE=de-DE -- uvx siemens-docs-mcp.

Tools:

Tool

Purpose

search_docs

Full-text search (filter by product, version, locale, host).

read_page

Read one page as Markdown; page through long pages with offset.

get_toc

Table of contents of a publication or of the subtree under a topic URL.

list_publications

List publications; discover valid product / version values.

Any reader URL returned by a tool (or copied from the browser) is valid input to read_page and get_toc.

Environment variables:

Variable

Default

Meaning

SIEMENS_DOCS_HOSTS

(none)

Comma-separated extra hosts, appended to the built-in two (docs.tia.siemens.cloud, docs.industrial-operations-x.siemens.cloud).

SIEMENS_DOCS_LOCALE

en-US

Default locale for search and listings.

SIEMENS_DOCS_MIN_INTERVAL

0.3

Minimum seconds between requests to a host.

The publication catalog and TOCs are cached in memory (not on disk), so the first call per host takes a few seconds.


Markdown export (CLI)

Export a whole publication to Markdown — one file per page, folders mirroring the navigation hierarchy. Create a config with any reader URL of the publication (a topic URL exports the whole publication):

# my_docs.yaml
name: tia_openness_v21
url: "https://docs.tia.siemens.cloud/r/en-us/v21/tia-portal-openness-api-for-automation-of-engineering-workflows"
output_dir: "output/tia_openness_v21"
# Preview pages without writing files
siemens-docs-export my_docs.yaml --dry-run

# Export everything
siemens-docs-export my_docs.yaml

# Export a single page (for testing output quality)
siemens-docs-export my_docs.yaml --page cybersecurity-information

# Override the output directory / verbose logging
siemens-docs-export my_docs.yaml --output /tmp/docs --verbose

A ready-made config lives in configs/ in the repository.

Output structure

output/tia_openness_v21/
├── index.md                          ← root page
├── cybersecurity-information.md
├── what-s-new-in-tia-portal-openness.md
├── basics/
│   ├── basics.md
│   └── ...
├── tia-portal-openness-api/
│   ├── tia-portal-openness-object/
│   │   └── ...
│   └── ...
└── ...

Configuration reference

Key

Required

Description

name

No

Human-readable label shown in log output.

url

Yes*

Any reader URL of the publication (resolved to its map automatically).

api_base

Yes*

Legacy: root URL of the Fluidtopics instance (use with map_id).

map_id

Yes*

Legacy: Fluidtopics map identifier (use with api_base).

output_dir

Yes

Directory where Markdown files will be written.

* Either url, or api_base + map_id. If both are present, url wins.


How it works

Fluidtopics (the platform behind both portals) exposes a REST API that the browser SPA uses internally. This package calls that API directly — no browser automation:

  1. Catalog — GET /api/khub/maps lists every publication; reader URLs are resolved against it.

  2. Search — POST /api/khub/clustered-search with product/version/locale filters.

  3. TOC — GET /api/khub/maps/{mapId}/pages returns the navigation tree.

  4. Content — GET /api/khub/maps/{mapId}/topics/{contentId}/content returns raw HTML, converted to Markdown with markdownify.

Requests are throttled per host and retried once on 401/403/429/5xx.

Note: Only Fluidtopics-based portals are supported. Other platforms (MadCap Flare, Paligo, etc.) would need a different adapter.


Development

python -m pytest -q          # offline tests
python -m pytest -m live -q  # live smoke tests against the real hosts
ruff check .

Created with Claude AI

Available Tools

4 tools
get_tocA

Show the table of contents (1-6 levels via depth) of the publication, or of the subtree under the topic, in url: an indented outline, one line per node with title and reader URL. Reader URLs from any search result, TOC line or page link are valid input. Allowed hosts: docs.tia.siemens.cloud, docs.industrial-operations-x.siemens.cloud (default: docs.tia.siemens.cloud).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
depthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses the return shape ('an indented outline, one line per node with title and reader URL'), the usable depth range (1-6), and host restrictions/defaults. It omits permission requirements, rate limits, or failure behavior, but for a read-only outline tool this is solid coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences that lead with the core action and follow with input/constraint detail; nothing is wasted. The second sentence is somewhat packed with host rules, but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with an output schema, the description is nearly complete: purpose, input sources, depth semantics, and host constraints are all present. The main gap is the absence of any cross-reference to sibling tools for routing, which would fully close the loop.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it defines `depth` as 1-6 levels (a constraint absent from the schema) and explains `url` as any reader URL from search results, TOC lines, or page links, plus the allowed/default hosts. This meaningfully exceeds what the bare `string`/`integer` schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Show the table of contents') and pins down scope ('of the publication, or of the subtree under the topic, in `url`'), so the agent knows exactly what is returned. It does not, however, differentiate itself from the closely related sibling `list_publications`, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It usefully states where valid input comes from ('Reader URLs from any search result, TOC line or page link are valid input') and constrains the allowed hosts, which is real invocation guidance. But there is no explicit when-to-use-this vs. `read_page`/`list_publications` routing, so the selection logic is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_publicationsA

List documentation publications (up to 50 rows): title, product, version, locale and reader URL. Use it to discover valid product and version values; reader URLs are valid input to read_page and get_toc. product (e.g. 'STEP 7', 'WinCC Unified', 'SIMATIC AX') and version (e.g. 'V21', '6.1') are matched against the catalog; call list_publications to see valid values. Allowed hosts: docs.tia.siemens.cloud, docs.industrial-operations-x.siemens.cloud (default: docs.tia.siemens.cloud).

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNo
localeNo
productNo
versionNo
title_containsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose real constraints: a 50-row cap, an allowed-host list with a default host, and that product/version match against the catalog. It omits pagination behavior for results beyond 50 rows and any rate-limit or auth notes, which keeps it out of the top band.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The size is reasonable and the output fields are front-loaded, but the instruction to call list_publications to see valid product/version values is stated twice (once at the start, once mid-sentence), and the host allowlist detail is buried at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description still usefully previews row contents. It is complete enough to invoke correctly, with only the unpaginated 50-row cap and the two undocumented filter parameters (locale, title_contains) as residual gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage the description must compensate, and it does for three of five parameters: it gives host values/default, plus concrete product ('STEP 7', 'WinCC Unified') and version ('V21', '6.1') examples. `locale` and `title_contains` receive no explanation at all, leaving a measurable gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List documentation publications') and enumerates the exact fields returned (title, product, version, locale, reader URL). It also clarifies its role as a discovery catalog, which separates it from search_docs and the page-reading siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it to discover valid `product` and `version` values and notes reader URLs feed into read_page and get_toc. There is no explicit when-not-use guidance against sibling search_docs, but the primary scenario is unmistakable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pageA

Read one documentation page as Markdown, with a header (publication, version, breadcrumb, source URL). If the page is longer than max_chars, the end says how to continue with a larger offset. Reader URLs from any search result, TOC line or page link are valid input. Allowed hosts: docs.tia.siemens.cloud, docs.industrial-operations-x.siemens.cloud (default: docs.tia.siemens.cloud).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
offsetNo
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: truncation past `max_chars` with an in-band continuation hint, the header metadata prepended to output, and an allowed-host constraint with a default. It omits auth requirements, failure/error behavior, and whether offset is pagination-stable, but the operational traits it does cover are the ones an agent needs to avoid bad calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and output format, then behavior, then input constraints. The header parenthetical is informative rather than filler; no sentence is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, yet the description still names the header fields. For a read-only, 3-parameter tool it covers retrieval, continuation, and input validity adequately; only failure modes and any access prerequisites remain unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does explain `url` (valid sources, allowed hosts) and references `max_chars` and `offset` behavior (truncation and continuing with a larger offset), but the offset unit (characters vs. something else), the relationship between the two numbers, and the explicit max_chars threshold semantics are left implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one documentation page as Markdown') plus the return shape (header with publication, version, breadcrumb, source URL). It is clearly distinguishable from siblings that search, list, or return a TOC, since this retrieves a single page's content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent where valid input comes from ('Reader URLs from any search result, TOC line or page link'), which implies the intended search_docs/get_toc -> read_page workflow, and constrains allowed hosts with a stated default. It stops short of an explicit when-not-to-use or a direct pointer to the sibling that produces those URLs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_docsA

Full-text search of Siemens documentation. Returns a numbered Markdown list (1-50 hits via limit): title, publication + version, breadcrumb, excerpt and reader URL, with the total hit count. Pass a result's reader URL to read_page or get_toc. product (e.g. 'STEP 7', 'WinCC Unified', 'SIMATIC AX') and version (e.g. 'V21', '6.1') are matched against the catalog; call list_publications to see valid values. Allowed hosts: docs.tia.siemens.cloud, docs.industrial-operations-x.siemens.cloud (default: docs.tia.siemens.cloud).

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNo
limitNo
queryYes
localeNo
productNo
versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden; it discloses the return shape, the 1-50 hit range, the default host, and the allowed hosts. It omits auth/rate-limit or error behavior, which is a modest gap for a read-only search.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then output format, then chaining and parameter guidance. Dense but each sentence earns its place, with only mild verbosity in the host enumeration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, and the description still adds chaining and parameter guidance beyond it. With 5 of 6 parameters documented and follow-up routing explained, it is nearly complete, missing only locale and any auth caveats.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it documents limit (1-50), product with concrete examples and catalog matching, version with examples, and host with allowed values and default. Only locale is left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Full-text search of Siemens documentation') and scopes it to a domain. It also frames the result's role by pointing to read_page/get_toc for follow-up, letting an agent place it among its siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes follow-up actions clearly ('Pass a result's reader URL to read_page or get_toc') and points to list_publications for valid product/version values. It doesn't explicitly state when NOT to use it versus list_publications, but the intended context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedget_toc
    • First observedlist_publications
    • First observedread_page
    • First observedsearch_docs

TDQS

A4.3/5.0

Scored across 4 tools

Disambiguation5/5

Each tool serves a distinct purpose: search_docs for full-text discovery, read_page for page content, get_toc for hierarchical navigation, and list_publications for catalog browsing. Descriptions clearly state inputs/outputs and how reader URLs flow between tools, leaving no overlap.

Naming Consistency5/5

All four tools use a consistent snake_case verb_noun pattern: search_docs, read_page, get_toc, list_publications. The only minor variation is the plural 'docs' in search_docs, but the convention remains predictable.

Tool Count5/5

Four tools are well-scoped for a documentation access server, covering discovery, navigation, and reading without redundancy. Each tool earns its place and the set remains lightweight.

Completeness5/5

The read-only documentation workflow is fully covered: list_publications finds valid products/versions, search_docs locates content, get_toc reveals structure, and read_page retrieves pages with pagination. No obvious missing operation for this domain.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to interact with SIEMENS PLC S7-1500/1200 controllers through their JSON-RPC API, supporting authentication, tag browsing, variable read/write operations, alarm management, and diagnostic buffer access.
    11
    -
  • A
    license
    A
    quality
    A
    maintenance
    Provides AI assistants with direct access to Jamf official documentation, enabling them to answer Jamf-related questions by searching, retrieving articles, and browsing product documentation.
    6
    3,803 npm
    5
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to search, read, and traverse documentation bundles in Open Knowledge Format via MCP tools.
    506 npm
    73
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server for controlling Siemens TIA Portal, enabling AI agents to connect, read PLC tags, list blocks, inspect project structure, and compile projects. This community edition provides 5 read-only tools under the MIT license.
    MIT