nickol-knx-mcp
nickol-knx-mcp
설계 단계에서 사용할 수 있는 KNX / ETS6 어시스턴트로, MCP 서버로 제공됩니다.
이 도구로 할 수 있는 네 가지 기능 — 모두 운영 중인 KNX 버스에 전혀 접촉하지 않고 가능합니다:
사양서로부터 프로젝트 설계 — 장비 목록/프로젝트 사양을 완전한 그룹 주소 구조로 변환하고 전체 구현 문서 세트(ETS로 가져올 수 있는 XML/CSV, 사람이 읽을 수 있는 보고서, Home Assistant YAML, 인수 테스트 프로토콜, 시공 후 인계 패키지)를 생성합니다.
기존 프로젝트 감사, 수리 및 완성 — 명명 규칙 · DPT 및 하위 DPT · 명령↔상태 · KNX Secure · Matter 호환성을 검증하고, 구체적인 수정 제안(추론된 DPT, 합성된 상태 GA), 완성도 등급 부여, 두 프로젝트 버전 간 차이를 확인합니다.
스마트 홈 레이어 생성 — 실제 장치 상태를 읽는 조립된 Home Assistant 엔터티(컬러 조명, 기후, 커버, 센서)를 생성하며, 모호한 사항은 모두 사람의 검토로 넘깁니다.
매개변수화된 방 템플릿으로 새 프로젝트 구성 — 방 목록(각 슬롯에
basic/comfort프리셋 포함)으로부터 새로운 검증된 프로젝트를 조립 → 할당 명세서 + ETS GA XML/CSV + 장치 BOM 제안을 생성합니다. 드라이런, 신규 프로젝트 전용(R1).
내부적으로: 각 액추에이터를 실제 통신 객체로 확장하는 장치 라이브러리 — 일반 레시피부터 ETS 애플리케이션 프로그램에서 직접 파싱한 정확한 벤더 객체 모델까지 제공합니다.
🇷🇺 Русская версия: README.ru.md
새 소식 — 전체 데모 하우스.
examples/demo-home에는 합성된 239-GA / 47-Function 프로젝트, 도구가 생성한 보고서 + Home Assistant 구성 + ETS 내보내기, 그리고 전체 스마트 홈 "두뇌" — 일주기 조명, 8요소 기후 설정값, 재실/계절/시간 상태 머신 및 통계 — 가 포함된 5뷰 대시보드가 제공됩니다. 전체 내용은 **라이브 사이트 ↗**에서 확인하세요.
🖥️ 대시보드 — Home Assistant에서 실시간
라이브 Home Assistant에서 실행 중인 데모 하우스의 실제 스크린샷입니다. 도구가 조립한 엔터티가 작동하는 모습을 보여줍니다: RGBW / RGB / CCT 컬러 조명, 6개 바닥 난방 기후 구역(목표, 모드 및 밸브 %), 일주기 조명 곡선 및 계산된 기후 설정값 — 수동으로 설정하지 않았습니다.
기후 | 조명 |
에너지 및 통계 | 재실 |
▶ 라이브 사이트에서 대화형으로 살펴보기 → · 구성 파일은 examples/demo-home/ha-brain에 있습니다.
Related MCP server: PBIP Builder MCP Server
🧪 상태 및 테스터 모집
이것은 공개 베타입니다. 전체 파이프라인은 합성 프로젝트에서 종단 간 스모크 테스트를 통과했으며 실제 수천 개 GA 규모의 ETS5/ETS6 프로젝트(익명화)에 대해 검증되었습니다. 하지만 실제 ETS 프로젝트는 매우 다양하고 복잡하며, 더 많은 현장 보고가 도구를 개선합니다.
👉 ETS5/ETS6 프로젝트가 있다면 꼭 시도해보고 결과를 알려주세요. Real-project test report 이슈를 열어주세요. 도구는 읽기 전용이며 버스에 연결되지 않으므로 테스트는 안전합니다(참조: 안전 모델). 자세한 내용은 CONTRIBUTING.md를 참조하세요.
💬 토론에 참여하기 → — 인사말, 질문, 또는 도구가 프로젝트에서 발견한 내용을 공유하세요.
🗺️ 로드맵 — 실제 통합업체가 형성하는 방향
최근 실제 KNX 통합업체의 리뷰(토론)가 다음 개발 방향을 결정하고 있습니다:
장치 간 매개변수 일관성 (출시됨 —
check_device_parameters) — ETS 매개변수 설정이 동일한 N개 형제 장치와 다른 하나의 장치를 표시: 설정값/히스테리시스가 다른 온도 조절기, 감지 시간이 다른 재실 감지기..knxproj에서 직접 장치별 매개변수를 추출하여 실제 42–275개 장치 프로젝트에서 이상치를 찾아냅니다 — 읽기 전용, ETS 불필요, 버스 불필요 — 그리고 깨끗한 프로젝트에서는 아무것도 올바르게 보고하지 않습니다(다른 벤더/통합업체 스타일에서 오탐지 없음).프로젝트 정책 프로필 (출시됨 —
check_policy) — 하나의 보편적인 "전문 표준" 대신 자체 합의된 규칙(명명, GA 분류, 명령/상태 예외)에 대해 프로젝트를 검증합니다. 통합업체마다 관례가 다르기 때문입니다. 프로필이 없으면 프로젝트 자체에서 추론된 분류 체계에 대해 검증합니다.방 템플릿 라이브러리 — 매개변수화된 방 템플릿으로 새 프로젝트를 구성합니다. R1 출시됨 (
compose_rooms+validate_room_template: 새 프로젝트, 드라이런, 할당 명세서 + ETS XML/CSV + 장치 BOM). R2 계획: 기존 프로젝트 도킹 + 정확한 장치 선택.로직 머신 지원 (예정 — 연구 중) — 동일한 읽기 전용, 설계 단계 모델을 Logic Machine(Embedded Systems) 설치에 적용: LM 기반 KNX 프로젝트를 파싱하고 동일한 명명/DPT/상태/토폴로지 감사를 수행하며 동일한 인계 출력물을 생성하여, LM 통합업체도 원시
.knxproj에서 얻을 수 있는 것과 동일한 증거 기반 프로젝트 모델을 얻을 수 있도록 합니다. 현재 실제 Logic Machine 5 장치를 대상으로 범위를 설정 중입니다.ETS 내 그룹 주소 연결에 대해서는 의도적으로 바퀴를 재발명하지 않습니다. ETS 내에서 GA를 통신 객체에 연결하기 위해 이미 현재 ETS 앱 스토어 애드인이 있으며, ETS7에서 Smart Linking이 기본 제공될 예정입니다 — 이 도구는 해당 도구들을 안내하고, 읽기 전용 감사와 증거 기반 프로젝트 모델에 집중합니다.
테스트할 프로젝트, 문제가 있는 워크플로우, 또는 기능 제안이 있으신가요? → 토론.
이 도구가 필요한 이유
2026년 중반 기준으로 기성품 ETS6 ↔ Claude / MCP 도구는 없습니다. KNX 커뮤니티는 AI/CLI 워크플로우를 통해 프로젝트를 검사하고 수정(장치 및 그룹 주소 추가/이름 변경)할 수 있는 통합을 명시적으로 요청해 왔습니다. 이 패키지는 바로 그 설계 단계 레이어를 채웁니다 — 누락된 부분입니다.
권장되는 전체 구성은 4개 레이어이며, 하나만 처음부터 구축하면 됩니다:
레이어 | 목적 | 사용할 도구 | 직접 구축? |
1. 실시간 | 실행 중인 주택의 상태, 제어, 디버깅 | 공식 Home Assistant MCP 서버 + KNX (XKNX) 통합 | 아니요, 이미 존재함 |
2. 설계 단계 |
|
| 예 — 이것이 공백입니다 |
3. 파일 + Git | YAML/CSV/XML, 주소 스키마 버전 관리 | 표준 파일시스템 + git MCP 서버 | 아니요, 이미 존재함 |
4. 스킬 | 설계 규칙(GA 구조, 명명, DPT, 장면) + 운영 규율 |
| 아니요, 포함됨 |
안전한 설계: 레이어 2(이 서버)는 물리적으로 버스에 연결할 수 없습니다. 네트워크/버스 종속성이 전혀 없으며, 오직
.knxproj를 읽고 제한된 작업 공간에 파일을 씁니다. "운영 버스에 쓰지 않음" 요구 사항은 약속이 아닌 구조적으로 강제됩니다. 주택과의 실제 상호작용은 항상 레이어 1(Home Assistant)을 통해서만 이루어집니다.
이 도구로 할 수 있는 작업
📐 시나리오 1 — 사양서로부터 프로젝트 설계 (사양 → 구현 키트)
프로젝트 사양(장비 목록, 케이블 저널, 장치 목록)을 완전하고 검증된 그룹 주소 구조로 변환합니다. 그리고 이를 구현하기 위한 전체 문서 세트를 생성합니다:
장치 목록 → 객체 모델. 각 장치는 장치 라이브러리(
decompose_device)를 통해 실제 통신 객체로 확장됩니다. 조광 채널은 on/off + 상태 + 상대 조광(3.007) + 절댓값(5.001) + 밝기 상태로 구성되며, "하나의 GA"가 아닙니다. 바닥 난방 구역은 8개의 객체, 펄스 미터는 6개의 객체로 구성됩니다.전문 로직 계층. 단순한 명세는 프로젝트를 완성하는 요소(중앙 및 구역 매크로, 씬, 재실 로직, 기후 제어 스캐폴딩, 태양/바람 셔터 로직, 누수→차단 체인, 천문/기상 및 날짜-시간 소스, 모든 범위의 예약)를 언급하지 않습니다. 이 방법론은 이러한 완전성 패턴을 인코딩하며, KNX 협회 표준, 공개 제조업체 문서, 실제 전문 시공 ETS 프로젝트(익명화) 연구에서 추출되었습니다.
구조 및 규율. 3레벨 주소 지정, 구역+기능 명명, 명령↔상태 페어링, 모든 주소에 DPT 할당.
산출물(각각 하나의 명령): ETS 가져오기 가능 XML/CSV · Markdown 보고서 · Home Assistant YAML · 기능 승인 테스트 프로토콜 · 시공 인계 패키지(인벤토리, GA 맵, 커버리지 %, 보안 상태, QA 결과, 토폴로지 SVG).
전체 방법론: docs/spec-to-structure.md. 실제 시공 ETS 프로젝트(3,600개 이상의 그룹 주소)를 명세만으로 재구성하여 현장 검증 완료: ~92% 구조적 일치(분류 체계, 도메인, 자동화 로직, DPT 분포)에 검증 오류 0개 — 나머지 차이는 통합업체의 장치별 파라미터화로, 어떤 명세도 이를 인코딩하지 않습니다.
🔍 시나리오 2 — 기존 프로젝트 감사, 수리 및 완료
읽기 및 분류.
xknxproject를 통해 비밀번호로 보호된 ETS5/ETS6.knxproj를 파싱합니다. 모든 GA를 DPT + 다국어(EN/DE/RU) 이름 키워드로 범주(조명 / 셔터 / HVAC / 센서 / 씬 / 에너지 / 진단)와 종류(명령 / 상태 / 센서)별로 분류합니다. GA 목적 태깅(functional/reserve/logic/scratch)을 통해 의도된 자리 표시자를 오류 목록에서 제외하므로 보고서가 허위 경고를 하지 않습니다(실제 685-GA 프로젝트에서 허위 오류 29 → 6).검증(
analyze_all이 모든 것을 실행): 명명 및 구조 · 누락된 상태 객체(ETS-Function 역할 우선, 그 다음 이름-토큰 페어링, 위치 기반 페어링 — 1:1 이름으로 병렬 상태 중간 및 자체 보고 R+T 객체) · 누락/불일치 DPT + 하위 DPT 건전성(5.001을 가진 "온도" GA는 플래그 지정) · 상대 조광만 있는 디머 · KNX 보안 상태(보안 vs 일반 텍스트, 혼합 그룹, 키링 체크리스트 — 키 자료는 절대 읽지 않음) · Matter 준비 상태 · 에너지 도메인 커버리지.플래그 지정뿐만 아니라 수리(
suggest_repairs): 이름에서 DPT 추론, 의심스러운 하위 DPT 수정, 사용 가능한 주소 슬롯에 누락된 상태 GA 합성, 절대 밝기 GA 추가. 제안만 제공 — 사람이 검토하며, 수락된 GA는 ETS 내보내기에 반영됩니다. 실제 3,646-GA 프로젝트에서: 145개의 구체적 제안(32개 DPT 추론, 112개 합성 상태 GA).작업 완료:
grade_completeness(기본 골격 → 시공 점수),suggest_names,diff_projects(두.knxproj리비전 간 의미론적 차이: 추가/제거/DPT 변경/이름 변경/보안 변경), 이후 보고서, 인계 패키지 및 테스트 프로토콜 재생성.
🏠 시나리오 3 — 스마트 홈 레이어 생성(Home Assistant)
엔티티를 보수적으로 조합: 커버 → 색상/조광 가능 조명(on/off + 밝기 + RGBW/RGB/색온도 + 상태) → 스위치 → 기후(현재 온도, 목표 온도 상태, 작동/제어기 모드, 밸브 값) → 센서/바이너리. 모든 엔티티는 장치가 보고할 수 있는 경우
state_address를 가집니다. HA는 실제 상태를 읽으며, 가정하지 않습니다.검토 우선: 모호한 것(DPT 5.001 — 밝기인지 블라인드 위치인지)은 추측하지 않고 설명과 함께
review목록으로 이동합니다(액추에이터 종속 커버 플래그invert_position/ 이동 시간 포함, 이는 어떤.knxproj도 인코딩하지 않음).추가 기능: 날짜/시간 브로드캐스트(DPT 19.001)용
expose블록, Matter 준비 상태 린트, KNX IoT(Turtle/RDF) 의미론적 내보내기.집의 실시간 제어는 공식 Home Assistant 통합(레이어 1)에 남아 있으며, 이 서버는 해당 구성을 준비만 합니다.
운영 동반자:
skills/ha-git-backup— 배포 후 구성의 수명 주기:/config의 실제 git 히스토리(배포 키 + 사전 커밋 비밀 스캐너)와 GitHub Releases의 암호화된 오프사이트 백업, 월간 복구 훈련.
🧱 시나리오 4 — 방 템플릿으로 새 프로젝트 구성
빈 시트가 아닌 방에서 시작: 6개의 내장 파라미터화된 방 템플릿(침실, 어린이방, 거실, 주방, 욕실, 복도) 중 선택, 슬롯별로
basic/comfort프리셋 선택(집은 기본 조명과 편안한 기후를 혼합할 수 있음),compose_rooms가 새 프로젝트를 조합합니다.결과물: 할당
manifest(메인 = 도메인, 중간 = 역할, 서브 순차), 기존 생성기를 통한 ETS 가져오기 가능 GA XML/CSV, 장치 라이브러리의 장치 BOM 제안.실제 리더로 검증: 생성된
.knxproj는 표준load_project를 통해 다시 읽히며(타사 프로젝트와 동일한 경로), 4개의 린터(명명/누락 상태/DPT/정책)를 오류 0개 / 경고 0개로 통과합니다.기본적으로 드라이 런, 새 프로젝트만. 템플릿 형식은 공개 계약입니다(
room_templates/SCHEMA.md): ID는 로캘 중립적인slot_id이며, 사람 이름이 아닙니다.R2: 기존 프로젝트에 도킹 + 정확한 장치 선택 — 계획 중.
🧩 기반 — 성장하는 장치 라이브러리
parse_devices_from_project는.knxproj/.knxprod내부의 제조업체 애플리케이션 프로그램에서 정확한 벤더 객체 모델을 추출합니다(HDL/Ekinex와 같은 ref-level(ComObjectRef) 게시자 포함): 객체 번호, 이름, 크기, DPT, C/R/W/T/U 플래그, 채널당 블록 스트라이드 — 결정론적이며 PII 안전(벤더 카탈로그 데이터만; 클라이언트 프로젝트 파일 부분은 절대 읽지 않음).NICKOL_KNX_CATALOG을 카탈로그로 지정하면decompose_device가 일반 레시피 대신 정확한 모델(catalog-exact)로 응답합니다. 카탈로그는 사용자가 제공하는 프로젝트 및 제품 데이터베이스에서 요청에 따라 성장합니다.벤더가 DPT를 선언하지 않고 제공한 객체는 정직하게
unverified로 유지됩니다. 추측하지 않습니다.
모든 쓰기는 작업 공간 디렉터리(NICKOL_KNX_WORKSPACE, 기본값 ./knx-workspace)로만 이루어지며, 외부 쓰기는 거부됩니다.
설치
Python 3.10+ 필요.
git clone https://github.com/NickoScope/nickol-knx-mcp.git
cd nickol-knx-mcp
python3 -m venv .venv && source .venv/bin/activate
pip install -e .종속성: mcp>=1.10, xknxproject>=3.8, PyYAML>=6.0.
Debian/Ubuntu에서 pip가 외부 관리 환경에 대해 불평하는 경우, venv를 사용하거나(위와 같이)
pip install -e . --break-system-packages를 사용하세요.PyJWT충돌이 발생하면 먼저pip install mcp --ignore-installed PyJWT를 실행하세요.
확인:
python tests/test_pipeline.py # synthetic 16-GA project, end-to-end smoke test
nickol-knx-mcp # start the MCP server (stdio)Claude에 연결
Claude Desktop
examples/claude_desktop_config.json은 nickol-knx + filesystem + git + home-assistant를 연결합니다. 최소 조각(macOS 구성 경로: ~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"nickol-knx": {
"command": "nickol-knx-mcp",
"env": { "NICKOL_KNX_WORKSPACE": "/path/to/your/knx-workspace" }
}
}
}Claude Code
claude mcp add nickol-knx \
-e NICKOL_KNX_WORKSPACE="$HOME/knx-workspace" \
-- /absolute/path/to/.venv/bin/nickol-knx-mcp그런 다음 CLAUDE.md를 프로젝트 루트에 넣으면 ETS Assistant 스킬(설계 규칙, 안전 규칙, 3레벨 GA 구조, 명령/상태 페어링, DPT 규율, 명명, KNX 보안 키링 처리 및 권장 워크플로) 역할을 합니다.
MCP 도구 (31개)
읽기
도구 | 목적 |
|
|
| 분류 및 필터로 GA 나열 |
| 장치 및 해당 통신 객체 |
| 토폴로지(영역/라인/장치) |
| 하나의 GA에 대한 출처: 이렇게 분류된 이유 — 결정별 증거와 신뢰도 수준(권위적 ETS Function > 구조적 DPT > 휴리스틱 이름), 상태가 어떻게 페어링되었는지, 충돌(이름은 "AC"라고 하나 DPT는 조명 → |
검증
Tool | Purpose |
| 이름 규칙 / 3단계 구조 검증 |
| 상태 객체가 없는 액추에이터 |
| 누락/불일치 DPT + 하위 DPT 검증 (온도→9.001, 전력→14.056…) |
| 토폴로지 용량 + 개별 주소 유효성 (TP1 64/세그먼트, 256/라인, 유효하고 고유한 |
| KNX Data Secure 상태 + 키링 인계 체크리스트 |
| Matter 준비 상태 린트 (Matter 클러스터로 왕복 가능한 기능) |
| 계량/에너지 DPT 검사 + PV/배터리/EVSE 스캐폴드 |
| 모든 검사를 한 번에 실행 |
| 프로젝트 정책 프로필(메인 그룹 분류, 이름 지정, 페어링)에 대해 검증하거나, 프로필이 없으면 프로젝트 자체에서 추론된 분류에 대해 검증합니다. 보편적인 표준이 아닌 사용자의 규칙에서 벗어난 GA에 플래그를 지정합니다. |
수리 및 설계
Tool | Purpose |
| 플래그만 표시하지 않고 수정 제안 — DPT 추론, 상태/밝기 GA 합성 |
| 이름 규칙 제안 |
| 장치 → GA 분해: 로컬 카탈로그( |
| 내장 장치 라이브러리 (Zennio + ABB 제품군) |
|
|
| 장치 간 매개변수 QA: ETS 매개변수가 동일한 N개의 형제 장치와 다른 장치 찾기 (이상한 온도조절기/센서) — |
| 프로젝트 등급: 기본 골격 vs 시공 완료 |
| 두 |
생성
Tool | Purpose |
| HA KNX YAML (색상 + 기후 + 노출) + 검토 목록 |
| ETS 가져오기 가능 GA |
| 시공 완료 인계: 인벤토리, GA 맵, 커버리지, 보안, QA, topology.svg |
| 기능 승인 프로토콜 (명령 → 예상 상태) |
| KNX IoT 시맨틱 내보내기 (Turtle/RDF) |
| Markdown 보고서 |
| 작업 공간 경로 + 안전 보장 |
룸 라이브러리 (R1 — 룸 템플릿에서 새 프로젝트 구성)
Tool | Purpose |
| R1 스키마에 대해 룸 템플릿(내장 slot_id 또는 사용자 정의 YAML) 검증 |
| 룸 목록에서 새 프로젝트 구축 → 할당 |
일반적인 워크플로
load_project→.knxproj파일을 지정합니다 (보호된 경우 비밀번호 포함).analyze_all또는project_report→ 결과를 읽습니다. 먼저 사람이 검토합니다.ETS에서 이름/DPT/상태를 수정합니다 (생성된 GA를 가져오거나 수동으로).
generate_ets_group_addresses(fmt="xml")→ 누락된 GA를 ETS로 가져옵니다.generate_ha_package→ YAML을 Home Assistant에 배치합니다.review항목은 수동으로 해결합니다.모든 것 (
.knxproj내보내기, HA 설정, 주소 스키마)을 Git에 보관합니다.실제 집에는 Home Assistant MCP(레이어 1)를 통해서만 접근합니다.
제한 사항 (솔직하게)
명령/상태 및 카테고리 분류는 휴리스틱(DPT + 이름 + ETS Functions)입니다. Functions가 없고 비표준 이름이 있는 복잡한 프로젝트에서는 거짓 음성/양성이 발생할 수 있습니다. 따라서 보고서는 항상 사람이 검토하도록 되어 있으며, 모호한 항목은
review로 이동하고 설정에 포함되지 않습니다.DPT 5.001은 구조적으로 모호합니다 (밝기 vs 위치). 키워드로 구분되므로 비표준 이름의 경우 다시 확인하세요.
HA 생성기는 보수적입니다: 잘못된 엔티티를 내보내기보다는 항목을
review로 연기합니다.서버는 버스에 쓰지 않으며 ETS와 직접 통신하지 않습니다. ETS 교환은 GA의 파일 가져오기/내보내기만 가능합니다.
합성 데모 프로젝트와 실제 수천 개 GA의 ETS5/ETS6 프로젝트(익명화)에서 검증되었지만, 실제
.knxproj파일은 매우 다양하며 아직 베타 버전입니다. 따라서 테스터 모집 중입니다.
🔒 안전 모델
구조적으로 버스 접근 불가. 종속성 트리에 네트워킹 또는 버스 라이브러리가 없습니다.
workspace_info()는bus_access: false를 보고합니다.프로젝트에 대해 읽기 전용.
project.py는.knxproj를 다루는 유일한 모듈이며 읽기만 수행합니다.제한된 쓰기. 모든 출력은
NICKOL_KNX_WORKSPACE로 제한되며, 그 외부 경로는 거부됩니다.악의적인 프로젝트 파일에 대한 강화.
.knxproj는 신뢰할 수 없는 ZIP-of-XML이므로 구문 분석은safexml.py를 통해 실행됩니다: DTD/엔티티 XML은 거부되고(billion-laughs / XXE), 아카이브는 크기/항목/압축 해제 비율 제한에 대해 사전 점검되며 경로 탐색 이름은 거부됩니다(압축 폭탄 방어).사람이 개입.
project_report를 생성하고 ETS로 가져오거나 Home Assistant에 배포하기 전에 검토하세요.보안 문제를 발견하셨나요? SECURITY.md를 참조하세요.
패키지 레이아웃
nickol-knx-mcp/
├── nickol_knx_mcp/
│ ├── dpt_map.py # DPT → category / kind / HA platform / value_type
│ ├── project.py # the ONLY module that reads .knxproj (read-only)
│ ├── safexml.py # hardened ZIP/XML parsing of untrusted .knxproj (zip-bomb / XXE defense)
│ ├── pairing.py # command↔status pairing by name tokens
│ ├── analyze.py # naming / missing-status / DPT checks
│ ├── generate_ha.py # Home Assistant KNX YAML generation
│ ├── generate_ets.py # ETS XML + CSV generation
│ ├── report.py # Markdown report
│ ├── room_library.py # Room Library R1 — compose a new project from templates
│ ├── room_templates/ # built-in room YAML templates + SCHEMA.md (public contract)
│ └── server.py # FastMCP server, 31 tools, confined writes
├── tests/test_pipeline.py
├── examples/claude_desktop_config.json
├── skills/
│ └── ha-git-backup/ # ops companion: 2-circuit HA backup (git history + encrypted offsite)
├── CLAUDE.md # ETS Assistant skill / playbook
├── pyproject.toml
└── README.md기여
테스터와 기여자를 환영합니다. 특히 실제 프로젝트 테스트 보고서를 환영합니다. CONTRIBUTING.md 및 이슈 템플릿을 참조하세요.
라이선스
MIT © 2026 Nikolay Miroshnichenko
KNX Association과 제휴하거나 보증하지 않습니다. "KNX" 및 "ETS"는 KNX Association cc의 상표입니다. 이는 독립적인 커뮤니티 도구입니다.
Available Tools
32 toolsanalyze_allC
Run every check and return the report summary plus all findings.
| Name | Required | Description | Default |
|---|---|---|---|
| name_regex | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavioral traits. It does not state whether the tool is read-only, requires authentication, or has any side effects. The only behavior described is running checks and returning results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that conveys the core purpose without waste. However, it could be slightly more structured by stating the output format or mentioning the parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists but is not provided, the description is incomplete. It does not explain what 'every check' entails, the structure of the report summary, or how to interpret findings. The presence of 12 sibling tools suggests the need for clearer context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the description does not mention the only parameter, name_regex. The description adds no meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs every check and returns the report summary plus all findings. This distinguishes it from sibling tools like check_dpt, which are individual checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not explain scenarios where running all checks is appropriate, nor does it mention that individual sibling checks exist for specific needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_device_parametersA
Find the device whose ETS parameter settings differ from its N identical siblings — the odd thermostat/sensor out (e.g. one thermostat with a different setpoint/hysteresis, one presence detector with a different detection time).
Reads per-device parameter values straight from the `.knxproj` project part
(data xknxproject does not expose), groups identical devices by application
program, and returns `clear_outliers` (a strong majority with a small minority —
likely a mistake) and `split_configs` (balanced 2+ variants — review, often two
zones). Numeric config parameters are listed first; names are resolved from the
device application program. Read-only, no ETS/bus. Give a real `.knxproj` `path`
(a password-protected/encrypted project cannot be read).| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| password | No | ||
| min_group | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It clearly states the tool is read-only, does not touch ETS/bus, reads directly from the .knxproj part, and cannot handle password-protected projects. It also explains the grouping logic and output categories. It does not disclose potential performance characteristics or failure modes beyond the password limitation, but the core behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it opens with the core purpose, then explains the data source, grouping logic, output categories, and constraints. Every sentence adds value, though the final sentence about password protection could be integrated more tightly. It is longer than the ideal but earns its length by covering multiple behavioral aspects.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (grouping, outlier detection, two output categories) and the absence of annotations, the description is quite complete. It explains the input requirements, the read-only nature, the output structure, and a key limitation. It does not explicitly document the 'min_group' parameter or edge cases like what happens when no outliers are found, but the output schema likely covers return values. Overall, an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'path' parameter implicitly by requiring a real .knxproj path and noting password-protected projects cannot be read. It does not explicitly explain 'password' or 'min_group' semantics, but the description's mention of 'strong majority' and 'balanced 2+ variants' hints at grouping thresholds. The output schema likely clarifies the return structure, and the description adds meaningful context about what the tool does with the path.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Find'), a precise resource (device parameter settings in a .knxproj project), and a clear outcome (identifying the odd device among identical siblings). It also distinguishes itself from siblings by focusing on parameter settings rather than naming, topology, or security, and it names the two output categories (clear_outliers, split_configs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: to find devices whose ETS parameter settings differ from identical siblings. It also gives exclusions: a password-protected/encrypted project cannot be read, and it is read-only with no ETS/bus involvement. This is strong guidance for an agent deciding between this and sibling tools like check_naming or check_topology.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_dptB
Detect missing, inconsistent or mismatched DPTs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It only states the detection function but does not disclose whether the tool has side effects, requires authentication, or is read-only. The term 'detect' implies read-only, but this is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the verb 'Detect'. It is concise, though it could include more context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and an output schema, the description is minimal. It does not explain what DPTs are or in what context the tool operates (e.g., a loaded project). An agent may need to infer the scope from sibling tool names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the input schema is fully covered. Description adds no parameter info, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool detects missing, inconsistent, or mismatched DPTs, which is a specific verb and resource. It distinguishes from sibling tools like check_missing_status and check_naming, but 'DPTs' is not explicitly defined, mildly affecting clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as check_missing_status or check_naming. The agent has no context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_energyA
Check the metering/energy domain (energy DPTs 13.x / 14.056) and suggest a per-circuit / PV / battery / EVSE structure for the HA energy dashboard.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It mentions 'check' and 'suggest', implying a read-only operation, but does not explicitly state it does not modify data, require permissions, or have side effects. The description is adequate but lacks full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence of about 20 words. It front-loads the action and domain immediately, includes specific DPTs and components, and contains no redundant or unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, output schema exists), the description is fairly complete. It explains what the tool does and what it suggests. However, it does not mention prerequisites (e.g., a loaded project) or the output format, though the latter is presumably covered by the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain parameter behavior. With 100% schema coverage (none), the baseline is 4. The description adds no parameter info, which is appropriate given the lack of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'check the metering/energy domain' and 'suggest a per-circuit / PV / battery / EVSE structure for the HA energy dashboard.' It specifies the energy DPTs (13.x / 14.056) and distinguishes itself from sibling tools like check_dpt and check_matter by focusing on energy dashboard structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While the purpose implies the tool is for energy dashboard setup, there is no explicit guidance on when to use it versus alternatives or when not to use it. Users can infer its applicability from the domain focus, but clearer exclusions or prerequisites would improve this dimension.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_matterB
Matter-readiness lint: which controllable functions round-trip to a Matter cluster (have command + status + a decodable DPT) and which won't.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description should disclose behavioral traits. It describes the tool as a 'lint', implying a read-only analysis, but does not specify whether it modifies state, requires specific permissions, or how results are presented. The existence of an output schema helps but the description itself lacks behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-crafted sentence that conveys the core purpose efficiently. Every word earns its place; no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and an output schema present, the description is mostly complete, but it does not explain what the output represents or provide any examples of checks performed. For a lint tool with zero parameters, it could be more thorough in describing what 'Matter-readiness' entails.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the input schema trivially covers everything. The description does not need to add parameter information, and the baseline score of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a lint check for Matter-readiness, specifying it checks which controllable functions round-trip to a Matter cluster. This distinguishes it from sibling tools like check_dpt or check_energy, though it could be more explicit about what 'controllable functions' refers to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when assessing Matter compatibility but provides no explicit guidance on when to use this tool versus alternatives like check_dpt or check_naming. No exclusions or alternative references are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_missing_statusB
Detect controllable GAs lacking a status/feedback counterpart.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits, but it only states what the tool detects. It does not mention whether the tool is read-only, whether it has side effects, or what the output format looks like (though an output schema exists). The description is insufficiently transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that conveys the core purpose without unnecessary words. It is well front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and a simple input schema, the description provides a basic understanding. However, it does not explain domain-specific terms ('GA', 'controllable', 'status/feedback counterpart'), and with an output schema present, more detail on the return value would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, and schema description coverage is 100% (vacuous). The description adds meaning by specifying the detection criterion (controllable GAs missing status/feedback), which helps understand the tool's purpose without needing parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool detects controllable GAs lacking a status/feedback counterpart. The verb 'detect' and the resource 'controllable GAs' are specific. However, it does not explain what 'GA' stands for, which may reduce clarity for agents unfamiliar with the domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'analyze_all' or 'check_naming'. There are no preconditions, examples, or notes on when to use or avoid this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_namingC
Validate naming conventions and 3-level structure.
| Name | Required | Description | Default |
|---|---|---|---|
| name_regex | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden for behavioral disclosure. It indicates a read-only validation, but does not explicitly state that the tool does not modify state, nor does it disclose any side effects, permissions needed, or output format. The minimal description fails to provide transparency beyond the implied non-destructive nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise, but it lacks structure and does not front-load critical information. It earns its place but could be expanded with minimal additional sentences to improve clarity without sacrificing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema and one optional parameter, the description is insufficient. It does not explain return values, error conditions, or typical use cases. The sibling tools suggest a validation domain, but the description does not leverage that context to make the tool's role clearer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the sole parameter 'name_regex'. The agent has no clue what regex pattern to provide, its purpose, or format. This missing semantic information severely hampers correct usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a verb ('Validate') and objects ('naming conventions and 3-level structure'), making the tool's purpose clear. However, it lacks specificity about what constitutes naming conventions and 3-level structure, and does not distinguish it from sibling validation tools like check_dpt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. Sibling tools exist (e.g., check_dpt, check_missing_status) but the description does not mention any context, prerequisites, or exclusions, leaving the agent without decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_policyA
Validate the loaded project against a Project Policy Profile — your
agreed rules (main-group taxonomy, naming regex, command/status exemptions),
not one universal "standard". Flags GAs whose domain doesn't match the main
group your policy assigns, and names that don't match your pattern. Pass
profile_path to a YAML profile (omit to validate against the taxonomy inferred
from the project itself). Set write_example_to to drop a commented example
profile into the workspace — seeded from the loaded project's own main groups
(mains that do not exist in the project are not written). Report-only.
| Name | Required | Description | Default |
|---|---|---|---|
| profile_path | No | ||
| write_example_to | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It explicitly states 'Report-only', which is a critical behavioral trait. It also discloses what gets flagged (GAs with domain mismatches, naming mismatches) and how the example profile is seeded (only mains that exist in the project are written). This covers the key side effects and constraints well, though it doesn't detail return format (but output schema exists).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but every sentence carries meaningful information. It front-loads the core purpose and then details parameters and side effects. It could be slightly more concise, but the structure is logical and the emphasis on key constraints (report-only, seeding) is effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema and two optional parameters, the description covers the essential behavior: what it validates, how to customize via profile_path, and how to generate an example. It explains the seeding nuance and the report-only nature. It doesn't elaborate on error handling or performance, but these are minor given the output schema and the tool's relative simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must fully explain parameters. It does: profile_path is described as a YAML profile path with an omission behavior (infer from project), and write_example_to is described as dropping a commented example profile seeded from the project's main groups. This adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Validate') and resource ('loaded project against a Project Policy Profile'), and explicitly distinguishes it from a generic standard by emphasizing 'your agreed rules'. It also names the exact checks performed (main-group taxonomy, naming regex, exemptions), making it distinct from siblings like check_naming and check_secure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use the tool: when you have an agreed policy profile. It explains the optional profile_path behavior (omit to infer from project) and the write_example_to use case. It doesn't explicitly name alternative tools, but the purpose is specific enough that an agent can infer when not to use it. It lacks explicit 'use X instead' guidance, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_secureA
Summarise KNX Data Secure posture + the keyring handover checklist.
Reports how many group addresses are secured vs plaintext, flags middle groups that mix secure and plaintext addresses (a function is only as secure as its weakest GA), and emits the ETS/HA keyring workflow as a checklist. Report-only — this server never touches key material.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description clearly states the tool is report-only and never touches key material, providing key behavioral transparency. However, it does not disclose other traits like performance or limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no wasted words, front-loaded with the main purpose, and efficiently covers the tool's actions and safety.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an output schema (implied by 'report-only'), the description sufficiently describes the output: counts, flags, and checklist. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and 100% schema coverage, so the description need not add parameter info. Baseline is 4, and no additional detail is necessary.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('summarise') and resource ('KNX Data Secure posture') and clearly distinguishes from sibling tools like analyze_all or check_dpt by focusing on security posture and keyring workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but does not explicitly state when to use it versus alternatives or provide exclusions. Usage is implied for security assessments.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_topologyA
Check topology capacity and individual-address validity (KNX Handbook).
Flags lines over the TP1 segment (64) / line (256) limits, invalid or duplicate individual addresses, and multi-line projects missing a coupler.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes what it checks and flags (line limits, addresses, couplers), which is informative. However, it does not explicitly state whether the tool is read-only, whether it modifies the project, or what the output structure looks like. Since an output schema exists, the return format is covered, but the side-effect profile is not disclosed. The description adds value beyond the schema by detailing the specific validity checks, but stops short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first states the tool's purpose and reference to the KNX Handbook, the second lists the specific checks performed. It is front-loaded with the core action and resource, and every sentence adds essential information. There is no redundancy or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter check tool with an output schema, the description is complete. It clearly states what the tool does and what conditions it flags, which is all an agent needs to decide when to invoke it. The presence of an output schema covers return-value details. No prerequisites or additional context appear necessary, given the sibling tools like load_project imply a project is already loaded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is an empty object with zero parameters, so there is nothing to explain. The baseline for zero parameters is 4, and the description appropriately does not attempt to describe non-existent parameters. The tool takes no arguments, so parameter semantics are trivially satisfied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Check' with a clear resource: 'topology capacity and individual-address validity'. It enumerates exactly what it flags (line limits, invalid/duplicate addresses, missing coupler), which distinguishes it from sibling check tools like check_naming or check_secure. The purpose is unambiguous and action-oriented.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its use case: validating topology constraints on a project. It provides clear context on what it covers, but does not explicitly name alternatives or state when not to use it. Since the tool name and description align with the 'check' family, the usage context is clear enough for an agent to select it appropriately, but lacks explicit exclusionary guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compose_roomsA
Compose a new KNX project from a list of room templates (constructor).
``rooms`` is a list of specs, each: ``{template, preset?, slot_presets?,
params?, label?}`` — e.g. ``{"template": "bedroom", "preset": "comfort"}``.
``preset`` is basic|comfort (per-room); ``slot_presets`` overrides individual
slots (mix comfort climate with basic lighting); ``params`` overrides template
defaults (window/circuit counts); ``label`` sets a custom zone name.
Pipeline: resolve templates+params to a functional model, allocate group
addresses (main = domain, middle = role, sub sequential), write a real
``.knxproj`` and **re-read it with the standard loader**, then run our linters
on the re-read project. Output: a ``manifest`` (allocation), ETS GA XML/CSV,
and a device ``bom`` proposal from the device library.
R1 builds NEW projects only and is dry-run by default (nothing written). Set
``dry_run=false`` with ``output_dir`` (a folder inside the workspace) to write
the .knxproj, ETS exports, manifest.yaml and bom.yaml. Docking into an
existing project and exact device selection are R2. Never touches a bus.
| Name | Required | Description | Default |
|---|---|---|---|
| rooms | Yes | ||
| dry_run | No | ||
| language | No | ru | |
| output_dir | No | ||
| project_name | No | Room Library house |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the dry-run default, that setting dry_run=false with output_dir writes files, the re-read validation step, linter execution, and outputs (manifest, GA XML/CSV, BOM). It also states 'Never touches a bus' for safety. This is comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with the purpose up front, followed by parameter details and then the pipeline and constraints. It is longer than a one-liner but every sentence adds useful context. The front-loading of the core action aids quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the entire workflow: input structure, processing pipeline, outputs, and R1/R2 limitations. It includes the dry-run default and file writing conditions. Given the output schema exists, the lack of a detailed return format is not a gap. No essential information for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It thoroughly details the 'rooms' structure (template, preset, slot_presets, params, label) and explains dry_run and output_dir semantics. It does not explicitly describe 'language' or 'project_name', but these have sensible defaults and are self-explanatory; the description compensates for the critical parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb ('Compose') and resource ('new KNX project from a list of room templates'), clearly identifying the tool as a constructor. It distinguishes itself from siblings like check_naming or suggest_repairs by emphasizing 'new' projects, and the pipeline description further clarifies its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'R1 builds NEW projects only' and notes 'Docking into an existing project and exact device selection are R2', giving clear when-not-to-use guidance. It also explains the dry-run default and the conditions for writing files, so an agent knows exactly when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decompose_deviceA
Expand a device into its group-address decomposition recipe.
A KNX actuator channel is not one GA — it expands into command/status/dimming/ position/mode objects, each with its DPT. Given a device order number, type or alias (e.g. 'ZIO-MB24', 'dimmer', 'JRA/S', 'presence detector') and a channel count, returns the objects a professional wires per channel and the total GA count. Use when turning a spec/ТЗ device list into a group-address structure.
| Name | Required | Description | Default |
|---|---|---|---|
| channels | No | ||
| order_number | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It explains that the tool returns objects per channel and total GA count, but does not mention side effects, prerequisites, or whether it is read-only. This is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (4 sentences), front-loaded with the main purpose, and each sentence adds value. No redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the presence of an output schema, the description covers input, output, and usage context. It could be improved by mentioning error handling or behavior when the order_number is not found.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description adds significant meaning by explaining that 'order_number' can be a type or alias and that 'channels' defaults to 1. This compensates for the lack of schema-level documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool expands a device into its group-address decomposition recipe, with specific verbs and resource. It provides examples of inputs and output, and the purpose is distinct from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool (turning a spec device list into a group-address structure), but does not explicitly mention when not to use it or provide alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_projectsA
Semantic diff between two .knxproj files (path_a = base/old, path_b = new): added / removed GAs, DPT changes, renames, security-flag changes. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| path_a | Yes | ||
| path_b | Yes | ||
| password_a | No | ||
| password_b | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that the tool is read-only and performs a diff of specified aspects. However, it does not describe behavior like error handling, password usage, or performance implications. Still, it provides a good overview of what the tool does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that immediately conveys the tool's purpose. It is front-loaded with the key concept ('Semantic diff') and lists specifics concisely. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (not shown), the description does not need to detail return values. It adequately explains the tool's role within the set of sibling tools. A brief mention of output type or the meaning of 'diff' could enhance completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning for path_a and path_b by clarifying their roles as base/old and new, which the schema does not specify. However, password_a and password_b are not explained, though they are optional and likely self-explanatory. Given 0% schema coverage, the description partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs a semantic diff between two .knxproj files, specifying the exact aspects compared (added/removed GAs, DPT changes, renames, security-flag changes). This differentiates it from sibling tools like analyze_all or check_* tools, which have different scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (comparing two project files) but does not explicitly state when to use this tool versus alternatives. There is no mention of when it should not be used or any exclusions. The 'Read-only' hint is useful but not a usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_gaA
Provenance for one group address — why the tool classified it the way it did.
Replays the classification and shows, per decision (category / kind / status pairing),
the signals that fired with a confidence tier: authoritative (an ETS Function role) >
structural (the KNX DPT) > heuristic (a name keyword). Flags conflicts (e.g. a GA
the DPT calls lighting while its name says "AC") — the hotspot for silent
misclassification. Read-only; use before trusting a category or generating an entity.
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and succeeds. It discloses that the tool is read-only, replays classification, outputs per-decision signals with confidence tiers, and flags conflicts. It also explains the confidence ordering (authoritative > structural > heuristic). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact paragraph of about five sentences, front-loaded with the core purpose ('Provenance for one group address'). It then details the replay mechanism, confidence tiers, conflict flagging, and usage context. Every sentence adds value with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the tool's moderate complexity (replaying classification logic), the description covers the purpose, output structure (decisions, signals, confidence tiers), conflict behavior, and usage guidance. An output schema exists (though not shown), so return format is presumably defined there. The description is sufficient for an agent to decide when and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description is the only source. It clarifies that the 'address' parameter is a group address ('one group address'), which is essential context. While it doesn't specify format examples, the singular required parameter's meaning is well conveyed by the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Provenance for one group address — why the tool classified it the way it did.' It clearly explains the function (replaying classification and showing signals with confidence tiers) and distinguishes it from siblings like check_naming or suggest_repairs by focusing on explanation rather than checking or fixing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'use before trusting a category or generating an entity.' It also implies a single-address scope and mentions conflict detection as a hotspot for misclassification. However, it does not name alternative tools or state when NOT to use it, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_ets_group_addressesA
Generate an ETS-importable Group Address export.
Args:
fmt: 'xml' (ga-export/01, recommended) or 'csv' (native ETS layout).
output_path: optional file inside the workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| fmt | No | xml | |
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description fails to disclose critical behavioral traits such as whether the tool modifies any state, requires specific prerequisites (e.g., loaded project), or what happens if output_path already exists (overwrite? error?). It only mentions file creation obliquely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, front-loading the primary function in the first sentence, and then listing the arguments with brief explanations. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two simple parameters and an existing output schema, the description is largely complete. However, it omits prerequisites (e.g., project must be loaded) and does not clarify the workspace concept, which is necessary for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Given 0% schema description coverage, the description compensates well by explaining the two parameters: fmt (with values and recommendation), output_path (optional file). This adds meaning that the schema alone does not provide. However, it could further clarify the workspace context and file overwrite behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it generates an ETS-importable Group Address export, specifying the two output formats (xml and csv), which distinguishes it from sibling tools like list_group_addresses and generate_ha_package.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as list_group_addresses or generate_ha_package. The description only mentions fmt options without contextualizing the tool's role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_handover_packA
Generate a project handover pack (as-built deliverable for commissioning).
Assembles an equipment inventory, group-address map by domain, command/status
coverage, KNX Secure scope and QA state into ``handover.md``, plus a
``topology.svg`` diagram, the full ``group-addresses.csv`` and the
``ha-package.yaml``. When ``output_dir`` is given (a folder inside the
workspace) all files are written there and the paths returned; otherwise the
handover markdown + SVG are returned inline.
| Name | Required | Description | Default |
|---|---|---|---|
| output_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully shoulders behavioral disclosure. It explains conditional behavior based on 'output_dir' and lists generated files. However, it omits details like permission requirements or potential side effects, which would elevate it to a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph that front-loads the purpose and uses efficient phrasing. It is concise but not overly terse; it could be slightly shorter by removing redundant phrases while retaining clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one simple parameter and an output schema exists, the description covers the core behavior adequately. It explains both modes of output and lists all produced files, meeting completeness needs for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description thoroughly explains the lone parameter 'output_dir', including its effect on output routing. This compensates well for the schema gap, though it doesn't specify exact path format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it generates a project handover pack, listing specific deliverables (handover.md, topology.svg, etc.). It uses the verb 'Generate' and specifies the resource, distinguishing it from sibling tools like 'generate_ets_group_addresses' which focus on single artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for project completion via 'as-built deliverable for commissioning', but does not explicitly state when to use this tool over alternatives or when not to use it. No exclusions or comparisons to sibling tools are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_ha_packageA
Generate a Home Assistant KNX package YAML.
If output_path is given, the YAML is written into the workspace and the path
returned; otherwise the YAML text is returned inline.
| Name | Required | Description | Default |
|---|---|---|---|
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Describes two modes (file write vs inline return) but lacks details on file overwrite behavior or potential side effects. No annotations to contradict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two clear, front-loaded sentences with no wasted words. Efficiently communicates purpose and conditional behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers main functionality for a simple tool with one optional param and output schema. Lacks mention of error handling or file overwrite policy, but remains adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (output_path). Schema has no description, but the description adds crucial behavioral context: if provided, write to workspace; otherwise inline. This compensates for schema coverage of 0%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it generates a Home Assistant KNX package YAML, with two modes based on output_path. This distinguishes it from siblings like analyze_all or get_topology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use each mode (with/without output_path). Does not mention alternatives or when not to use, but the context is sufficient for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_knx_iotB
Export a KNX IoT semantic view (Turtle/RDF) of the project's functional datapoints — a pragmatic skeleton for the IP-native model, for review.
| Name | Required | Description | Default |
|---|---|---|---|
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral disclosure burden. It only states the output format and purpose (skeleton for review), but does not disclose whether the operation is read-only, requires authentication, or has other behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise and front-loaded with the action. However, it could be structured to better separate purpose, usage, and parameter details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema, return values need not be explained. However, the description lacks context about what 'functional datapoints' are and how the output relates to other tools, making it minimally adequate for a simple export tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description does not explain the only parameter (output_path) beyond what the schema provides. It neither describes its effect nor suggests typical values, leaving an agent with no semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports a KNX IoT semantic view in Turtle/RDF format for review. It specifies the verb 'Export', the resource 'functional datapoints', and the output format, distinguishing it from sibling export tools like generate_ets_group_addresses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for review purposes ('for review') but provides no explicit guidance on when to use this tool versus alternatives like generate_ets_group_addresses or generate_handover_pack. No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_test_protocolA
Draft a functional acceptance protocol (per function: command → expected status, pass/fail/sign-off) as Markdown. Execution is manual/on-site; this only drafts it.
| Name | Required | Description | Default |
|---|---|---|---|
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It correctly states the tool only drafts and doesn't execute, but fails to mention side effects like file creation or whether output_path saves the file. Behavior is partially transparent but lacks detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences efficiently convey the tool's purpose, format, and execution context without extraneous information. Every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 1 optional parameter and an output schema, the description covers the core function and boundary. The missing parameter documentation is the only gap, so it's mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole parameter output_path has 0% schema description coverage and no explanation in the description. The agent cannot infer what this parameter does (e.g., output file path) from the description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool drafts a functional acceptance protocol in Markdown format, specifying the structure (command → expected status, pass/fail/sign-off). It distinguishes itself from sibling tools like generate_handover_pack by focusing on test protocols.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes execution is manual/on-site and this tool only drafts the protocol, implying it's for planning not execution. However, it doesn't explicitly compare to alternatives or state when to use this over other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_devicesA
List devices (individual address, name, order number, manufacturer), sorted by
individual address (area/line/device numerically). Same paging contract as
list_group_addresses: next_cursor / total_matched / returned.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the sorting order (numerical by individual address), the paging contract (next_cursor/total_matched/returned), and the returned fields. This is valuable behavioral context beyond the schema, though it does not mention read-only nature explicitly or authentication/rate-limit details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler. The purpose is front-loaded, followed by sort order and paging contract. Every sentence carries substantive information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description adequately covers the listing behavior, returned fields, and paging. It is sufficient to call the tool correctly, though it assumes familiarity with list_group_addresses for full cursor semantics. Minor gaps around read-only confirmation and potential limits do not prevent correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that limit/cursor follow the same paging contract as list_group_addresses, giving context that cursor is a pagination token and limit controls page size. However, it does not explicitly define each parameter, relying on the sibling tool's contract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('devices'), enumerates the returned fields, and specifies the sort order. It also references a sibling tool for the paging contract, which helps an agent distinguish it from related tools like list_group_addresses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to use this tool (when a list of devices is needed), but there is no explicit guidance on when not to use it or which alternative to prefer. The reference to list_group_addresses is about paging behavior, not tool selection, so usage context is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_topologyB
Return the area/line/device topology tree.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It does not disclose any behavioral traits such as read-only nature, performance considerations, or side effects, only stating the return action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 6 words, highly concise. However, it could be slightly more informative while remaining concise, but it earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and an output schema exists, the description is minimally adequate. However, it does not clarify what 'topology tree' entails or how to interpret results, leaving gaps for a new user.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so baseline 4 applies. The description correctly implies no parameters are needed, and schema coverage is 100%, so no additional clarification is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Return' and identifies the resource as 'area/line/device topology tree', clearly stating what the tool does. It distinguishes from sibling tools like 'analyze_all' or 'get_devices' which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. The description lacks context for appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grade_completenessA
Grade the project: bare functional skeleton vs as-built grade — by the presence of the professional patterns (central macros, device tuning, astro/meteo, monitoring, deep metering, scenes, reserves, a debug main).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions grading logic and criteria but does not disclose if the tool is read-only, has side effects, or requires specific permissions. The output schema exists but the description does not hint at output structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that clearly states the purpose and criteria. It could be slightly more structured (e.g., bullet list of patterns) but it is efficient and front-loads the verb and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters and an output schema exists, so the description need not detail return values. It explains the grading criteria adequately, though it could mention the scale better. Given the context, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so baseline is 4 per rubric. The description adds no parameter info, which is acceptable since no parameters exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Grade') and resource ('project') with a clear scale ('bare functional skeleton vs as-built grade') and lists the professional patterns used as criteria. This distinguishes it from sibling tools like 'analyze_all' or 'check_*' which focus on different aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not specify when to use this tool versus the many siblings (e.g., when to use 'grade_completeness' vs 'analyze_all' or 'project_report'). No when-not-to-use or alternative guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_device_recipesA
List the device decomposition recipes in the built-in device library.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description clearly indicates a read-only listing operation with no side effects. However, since no annotations are provided, the description carries the full burden; it could additionally state idempotency or that it does not mutate state, but the current description is transparent enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no superfluous words. It is front-loaded and efficient, earning its place without any padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, an output schema exists (though not described), and the operation is simple listing, the description is complete. It provides all necessary information for an agent to understand the tool's purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema coverage is 100% trivially. The description does not need to add parameter details. Baseline score of 4 is appropriate as it fully covers the semantic gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list', the resource 'device decomposition recipes', and the scope 'built-in device library'. This specifically distinguishes from sibling tools like 'decompose_device' (which applies a recipe) and 'get_devices' (which lists devices).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage as a simple listing operation, but it does not provide any explicit guidance on when to use it, when not to use it, or mention alternatives. For a straightforward tool, this is adequate but not fully instructive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_group_addressesA
List parsed group addresses with classification, in a stable order.
Filters: category (lighting/shutter/hvac/sensor/scene/energy/diagnostics),
kind (command/status/sensor), missing_dpt_only.
Paging: results are always sorted by the group address itself (main/middle/sub
numerically, free-style addresses numerically, anything else lexically), so the
order does not depend on how the project happened to parse and a retry returns
the same page. Pass the returned `next_cursor` back as `cursor` for the next
page; `next_cursor` is null on the last page. `total_matched` reports how many
addresses match the filters, so a truncated answer is never silent.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No | ||
| cursor | No | ||
| category | No | ||
| missing_dpt_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full behavioral burden. It goes well beyond the minimum by disclosing stable ordering rules, exact sort behavior per address type, cursor handling, the meaning of next_cursor and total_matched, and the fact that truncation is never silent. This is rich, concrete, and directly useful for correct invocation and result interpretation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well organized: a one-line purpose, a compact filter list, and a focused paging paragraph. Every sentence carries useful information, and the most important differentiation (stable order) is front-loaded. The length is justified by the behavioral details provided.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a filtered-list tool with an output schema and no annotations, the description is complete: it explains filtering, ordering, paging, cursor lifecycle, and non-silent truncation. Nothing necessary for an agent to call or interpret the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It explicitly enumerates accepted values for category and kind, explains cursor semantics via next_cursor, and clarifies missing_dpt_only as a filter. The meaning of limit is not elaborated, but its default and the paging context make it adequately inferable. This is strong compensation, though not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'List parsed group addresses with classification'. It immediately conveys the main output and the stable-order guarantee, which clearly distinguishes this tool from siblings like generate_ets_group_addresses or get_devices. The purpose is unambiguous and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes when to use this tool: to list group addresses with filtering by category, kind, or missing DPT, and to page through results. It does not explicitly name alternatives or say when not to use it, but the filter and paging context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_ga_exportA
Load an ETS group-address export (ga-export/01 XML) instead of a full .knxproj.
For when you only have the GA list: an ETS "Export Group Addresses" file, or the
ETS import file a planning tool produces (TapPlan and similar). Names, addresses,
DPTs, descriptions, the security flag and the range tree are read; the result
replaces the loaded project for every other tool.
Works: check_naming, check_missing_status, check_dpt, check_policy, check_secure,
analyze_all, suggest_repairs, project_report, generate_ha_package,
generate_ets_group_addresses. Nothing to read (the export has no devices, ETS
Functions or topology): get_devices, get_topology, check_topology,
decompose_device, check_device_parameters, parse_devices_from_project. Pairing
relies on names only, since there are no ETS Function roles.
Args:
path: Path to the exported .xml file.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and meets it: it discloses the state-changing side effect ('the result replaces the loaded project for every other tool'), the data that is extracted, the absence of devices/functions/topology, and the name-only pairing limitation. These are exactly the behavioral traits an agent must know before calling a load tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but it is cleanly organized into purpose, usage trigger, data content, effect, tool-compatibility matrix, and caveat paragraphs, with the core purpose front-loaded. The tool-name lists are bulky but carry real routing value for an agent deciding which tool to pair with this loader.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter loader with an output schema, nothing essential is missing: the return contract is left to the output schema, while the description covers purpose, side effects, compatibility, and limitations. The only conceivable addition would be explicitly naming load_project as the alternative, but the contrast is already clear from 'instead of a full .knxproj.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter has zero schema description coverage, so the description must compensate. The Args section states 'path: Path to the exported .xml file,' which together with the main description's 'ga-export/01 XML' context conveys the file type and purpose. It is adequate but adds little beyond that basic meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Load an ETS group-address export (ga-export/01 XML) instead of a full .knxproj.' It signals that this is the GA-only counterpart to the full-project load, distinguishing it from the sibling load_project, and enumerates exactly what is read (names, addresses, DPTs, descriptions, security flag, range tree).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the explicit trigger condition ('For when you only have the GA list') and names the exact accepted file types (ETS 'Export Group Addresses' file, TapPlan-like planning imports). It goes further than most definitions by enumerating which downstream tools work after loading and which have 'nothing to read,' plus the pairing-relies-on-names caveat.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_projectA
Parse a .knxproj file (read-only) and cache it for the session.
Args:
path: Path to the .knxproj file.
password: Project password, if the .knxproj is protected.
language: Optional language code (e.g. 'de-DE', 'ru-RU').
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| language | No | ||
| password | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, but description discloses read-only nature, caching behavior, and optional password/language parameters. Does not detail error handling or file format expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single-line purpose followed by concise Args list. No redundant or missing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given output schema exists, description covers main function, caching, and parameter details. Could mention that tool should be called before others, but implied by sibling names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, description fully explains all three parameters: path, password (with conditional), and language (with examples). Adds value beyond schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states verb 'parse' and resource '.knxproj file', with read-only and session caching scope. Distinguishes from sibling tools which operate on loaded data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly foundational for other project tools, but no explicit when-to-use, when-not-to-use, or alternatives beyond sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parse_devices_from_projectA
Extract exact device object models from a .knxproj / .knxprod application programs.
Reads the manufacturer application programs (M-*) embedded in an ETS `.knxproj`
(devices actually used) or a `.knxprod` product database, and returns each device's
order number, app-program version, object counts and detected per-channel blocks —
the EXACT vendor comm-object model, not a generic recipe. Read-only and PII-safe: it
reads only vendor catalog data, never the client project (P-*/0.xml).
Use this to build/grow the local device catalog that `decompose_device` consumes
(set NICKOL_KNX_CATALOG to the catalog dir). If `output_path` is given, the full
catalog is written into the workspace as device-library YAML; the return value is
always a compact per-device summary + coverage manifest (the full object lists are
not inlined). DPT `unverified` = the vendor app-program declares none (never guessed).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| password | No | ||
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: read-only, PII-safe, reads only vendor data, never client project. It explains what 'unverified' DPT means and that the return value is a summary not full object lists. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single paragraph is efficient and front-loaded with purpose, but somewhat lengthy. Every sentence adds value, though could be structured with bullet points for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters, no annotations, and output schema exists, the description covers purpose, usage, return value summary, and limitations. Lacks details on error cases or prerequisites but is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so description must compensate. It indirectly explains path (file type) and output_path (catalog writing), but password parameter is not explained. Adds value beyond schema but could be more complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it extracts exact device object models from .knxproj/.knxprod files, specifying the resource (vendor application programs) and action (extract). It distinguishes from siblings by contrasting with 'generic recipe' and noting it reads only vendor catalog data, not client projects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: to build the local device catalog for decompose_device, and mentions setting NICKOL_KNX_CATALOG. Also explains behavior with output_path and the return value format, providing clear context for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_reportC
Produce the human-readable Markdown report (review before any import).
| Name | Required | Description | Default |
|---|---|---|---|
| name_regex | No | ||
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It only mentions 'produce' (implying read-only) but lacks details on side effects, required inputs, or response behavior. The output schema exists but is not referenced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is concise and front-loaded with purpose. However, it omits important information about parameters and usage, reducing its effectiveness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description fails to cover parameter behavior, usage context, or behavioral traits. Two parameters remain unexplained, and the tool's role among siblings is unclear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter (name_regex, output_path). No guidance on their purpose or usage, leaving the agent uninformed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it produces a human-readable Markdown report, with a specific hint about reviewing before import. However, it does not explicitly differentiate from sibling tools like analyze_all or check_*, which may also generate reports.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus siblings. The phrase 'review before any import' implies a specific context but does not outline alternatives or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_namesB
Naming hygiene suggestions (empty names, status GAs missing a status keyword).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description bears full burden. It does not disclose whether the tool is read-only, requires authorization, or has side effects. The description only states purpose, not behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with clear parenthetical example. No wasted words; front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no parameters and has an output schema, so description need not detail returns. However, more context about what kind of suggestions or output format would improve completeness. Adequate but minimal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with zero parameters, so baseline is 3. Description adds no parameter-specific info because there are none, but also no omission.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool provides naming hygiene suggestions for empty names and status GAs missing a status keyword. The verb 'suggest' and specific resources are identified, distinguishing it from sibling tools like check_naming or grade_completeness.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not mention prerequisites, typical use cases, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_repairsA
Propose concrete fixes for the project's findings — repair, don't just flag.
For each issue it suggests a reviewable fix: infer a DPT for a GA that has none,
correct a suspect sub-DPT, synthesise a status/feedback GA in a free address slot,
or add an absolute-brightness GA for a relative-only dimmer. Suggestions only —
a human reviews them; accepted new GAs feed generate_ets_group_addresses. The
server never writes to ETS or the bus.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behavioral traits: 'The server never writes to ETS or the bus,' ensuring agents know this is a non-destructive suggestion tool. It also clarifies that suggestions are reviewable, not automatic. Without annotations, this level of disclosure is commendable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured. The first sentence immediately conveys the core purpose. The second paragraph provides concrete examples and workflow context without unnecessary verbosity. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and an output schema exists (though not detailed here), the description sufficiently covers behavior, workflow integration with generate_ets_group_addresses, and non-destructive nature. It could mention output format briefly, but the presence of an output schema mitigates this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the description does not need to elaborate on parameters. It adds value by explaining the tool's operation and what it accomplishes, compensating for the lack of parameter details. The schema coverage is 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Propose concrete fixes for the project's findings — repair, don't just flag.' It provides specific examples of fixes (infer DPT, correct sub-DPT, etc.) and distinguishes itself from sibling tools by noting that suggestions feed into generate_ets_group_addresses and that the server never writes to ETS or the bus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates when to use the tool (for concrete fixes based on project findings) and provides workflow context: 'Suggestions only — a human reviews them; accepted new GAs feed into generate_ets_group_addresses.' It does not explicitly list alternatives, but the context implies appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_room_templateA
Validate a Room Library template against the R1 schema (report-only).
Pass ``template`` (a built-in template's semantic slot_id, e.g. 'bedroom',
'kitchen') or ``path`` to a custom template YAML. Checks the public contract:
a locale-neutral slot_id, ru/en labels, per-slot basic/comfort presets, known
function types, valid multiplicities, and that ``area_m2`` is a hint with
provenance (never a normative fact). Returns ok + findings; nothing is written.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| template | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states report-only, that nothing is written, and that results are ok + findings. It also reveals the nuanced policy that area_m2 is a hint with provenance, not a normative fact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then organizes parameter usage and validation details into a compact, readable list. Every sentence adds distinct value with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without annotations, the description covers purpose, parameter semantics, side-effect behavior, and return shape (ok + findings). Since an output schema exists, return details need not be spelled out further. The tool call can be understood and invoked correctly from this description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate fully. It does so by explaining both parameters: template is a built-in slot_id with examples like 'bedroom'/'kitchen', and path points to a custom template YAML. It also clarifies the or-relationship between them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: validate a Room Library template against the R1 schema. It also distinguishes itself from sibling check_* tools by declaring report-only behavior and focusing on template contract checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly communicates the context of use: validating either a built-in template by slot_id or a custom YAML by path. It does not explicitly name alternatives or state when not to use the tool, but the narrow template-validation scope makes the intended usage plain.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_infoB
Show the confined output workspace and the safety guarantees.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It states the tool shows information, implying a read-only operation, but provides no details on outcome, side effects, or necessary permissions. Minimal disclosure beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 9 words, with no redundant information. Every word serves a purpose, making it highly concise and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (presumably documenting return values), the description adequately covers the tool's function for a simple info tool. It could be slightly more explicit about what 'confined output workspace' means, but overall complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description does not need to add parameter meaning, and the schema coverage is 100% by default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Show' and identifies the resource ('confined output workspace' and 'safety guarantees'), which clearly indicates the tool's purpose. It is distinct from sibling tools that perform analysis or checks. However, the term 'confined' may be unclear to some users.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not provide context or suggest scenarios, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.8.2- Added
check_device_parameters - Added
check_policy - Added
check_topology - Added
compose_rooms - Added
explain_ga - Changed
get_devices6 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / limitAdded value: +{ + "default": 500, + "title": "Limit", + "type": "integer" +} - added
Output schema / additionalPropertiesAdded value: +true - removed
Output schema / propertiesRemoved value: -{ - "result": { - "items": { - "additionalProperties": true, - "type": "object" - }, - "title": "Result", - "type": "array" - } -} - removed
Output schema / requiredRemoved value: -[ - "result" -] - changed
Output schema / titlePrevious value: -"get_devicesOutput"New value: +"get_devicesDictOutput"
- Changed
list_group_addresses5 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Output schema / additionalPropertiesAdded value: +true - removed
Output schema / propertiesRemoved value: -{ - "result": { - "items": { - "additionalProperties": true, - "type": "object" - }, - "title": "Result", - "type": "array" - } -} - removed
Output schema / requiredRemoved value: -[ - "result" -] - changed
Output schema / titlePrevious value: -"list_group_addressesOutput"New value: +"list_group_addressesDictOutput"
- Added
load_ga_export - Added
validate_room_template
13 tool updates
v0.2.2- Added
check_energy - Added
check_matter - Added
check_secure - Added
decompose_device - Added
diff_projects - Added
generate_handover_pack - Added
generate_knx_iot - Added
generate_test_protocol - Added
grade_completeness - Added
list_device_recipes - Added
parse_devices_from_project - Added
suggest_names - Added
suggest_repairs
12 tool updates
v0.1.0- First observed
analyze_all - First observed
check_dpt - First observed
check_missing_status - First observed
check_naming - First observed
generate_ets_group_addresses - First observed
generate_ha_package - First observed
get_devices - First observed
get_topology - First observed
list_group_addresses - First observed
load_project - First observed
project_report - First observed
workspace_info
TDQS
Scored across 32 tools
Tools are largely distinct due to consistent prefixes (check_*, generate_*, list_*, get_*) and targeted domains. Minor overlap exists (e.g., check_naming vs. suggest_names, analyze_all vs. project_report), but descriptions clarify boundaries.
Most names follow a verb_noun pattern with strong category prefixes. A few irregular names (analyze_all, project_report, workspace_info) break the pattern, but the overall style is predictable and readable.
32 tools is a large surface, well above the typical 3-15 range. While each tool appears justified, the count feels heavy and could be consolidated (e.g., parameterized check/generate tools), increasing cognitive load for agents.
The toolset covers the full KNX workflow: project loading, multiple analysis/check dimensions, report generation, exports, room composition, and device decomposition. Minor gaps exist (no direct write/apply fixes, no update/delete operations), but these are intentional design constraints.
Maintenance
Related MCP Connectors
Read, analyze, and safely edit Microsoft Project MPP files.
Deterministic preflight for FactoryTalk View tag and alarm CSV imports before re-import.
Inspect XLSX/XLSM workbooks, validate email, generate SVG QR codes, and look up domain DNS and TLS.
Generate, validate and read EN 16931 and Peppol BIS 3.0 e-invoices, in UBL and CII.
Related MCP Servers
- -licenseNot gradedqualityNot gradedmaintenanceEnables automated project analysis and structured development specification generation. Supports multiple export formats and integrates with AI models for comprehensive project documentation and validation.-
- FlicenseCqualityDmaintenanceCreates, inspects, validates, and modifies Power BI Project (.pbip) folders, generating PBIR-style reports and TMDL semantic models from structured inputs.52-
- AlicenseAqualityAmaintenanceLocal electronics tools for MCP-capable assistants, enabling static analysis of CRUMB save files and Logisim-evolution projects, including net tracing, BOM building, electrical rule checks, and optional truth table generation.2279 npm1Apache 2.0
- FlicenseAqualityAmaintenanceProvides read-only analysis of Mitsubishi GX Works3 PLC projects via MCP, enabling device tracing, cross-referencing, ladder inspection, linting, and report generation without modifying source projects.4157-