firewalla-mcp-server
firewalla-mcp-server
Claude가 Firewalla MSP API를 통해 Firewalla 구성 및 네트워크 보안 상태를 감사할 수 있도록 지원하는 읽기 전용 Model Context Protocol (MCP) 서버입니다.
설계상 읽기 전용입니다. 이 서버는 장치 차단/차단 해제, 규칙 생성 또는 수정, 서비스 일시 중지, 또는 Firewalla에 대한 어떠한 변경도 수행할 수 없습니다. 오직 관찰만 가능합니다.
기능
Claude가 네트워크 내 장치, 활성 규칙, 보안 알람, 네트워크 흐름, 차단/허용 대상 목록 등 Firewalla를 검사하는 데 사용할 수 있는 8가지 도구를 제공합니다.
Related MCP server: mcp-infra-readonly
예시 프롬프트
보안 감사
이 프롬프트들은 Claude를 네트워크 보안 전문가로 설정하여 Firewalla 구성을 체계적으로 검토하도록 합니다. Firewalla MCP 도구를 사용할 수 있는 Claude Desktop 또는 Claude Code에서 가장 잘 작동합니다.
전체 네트워크 보안 감사:
당신은 내 홈 네트워크의 포괄적인 감사를 수행하는 수석 네트워크 보안 엔지니어입니다. Firewalla MCP 도구를 사용하여 다음 검토를 수행하고, 심각도 등급(치명적 / 높음 / 중간 / 낮음 / 정보 제공)이 포함된 체계적인 보고서로 결과를 제시하십시오:
장치 인벤토리 — 전체 장치 목록을 가져옵니다. 인식되지 않는 MAC 공급업체가 있는 장치, 모니터링되지 않는 장치, 또는 불법 액세스 포인트일 가능성이 있는 예상치 못한 라우터급 장치를 표시하십시오.
규칙 감사 — 모든 차단/허용 규칙을 검토합니다. 지나치게 허용적인(광범위한 범위, 인바운드 방향, 장치 제한 없음) 허용 규칙을 식별하십시오. 적중 횟수가 0인 오래된 규칙을 표시하십시오.
알람 검토 — 유형 및 심각도별로 그룹화된 최근 알람을 검색합니다. 패턴(동일한 장치에서 반복되는 알람, 예상치 못한 국가에서 발생하는 알람, 외부 노출이 없어야 하는 장치를 대상으로 하는 알람)을 식별하십시오.
대상 목록 범위 — 활성화된 차단 목록을 검토합니다. 현재 목록 구성이 일반적인 위협 범주(멀웨어, C2, 피싱, 암호화폐 채굴, 새로 등록된 도메인)에 대해 적절한 보호를 제공하는지 평가하십시오.
네트워크 보안 상태를 개선하기 위해 취해야 할 권장 조치 우선순위 목록으로 결론을 내리십시오.
방화벽 규칙 격차 분석:
방화벽 정책 분석가 역할을 수행하십시오. 내 모든 Firewalla 규칙과 전체 장치 목록을 가져와서 상호 참조하십시오. 다음을 식별해야 합니다: (1) 규칙이 전혀 적용되지 않은 장치 — 전역 규칙에만 의존하고 있는지, 의도적인 것인지 확인하십시오. (2) 인바운드 액세스를 허용하는 규칙 — 어떤 장치를 대상으로 하며 범위가 적절하게 좁은지 확인하십시오. (3) 한 번도 실행되지 않은(적중 횟수 = 0) 차단 규칙 — 오래된 규칙인지, 아니면 방어하려는 위협이 존재하지 않는 것인지 확인하십시오. 각 범주에 대한 평가와 권장 조치를 표 형식으로 제시하십시오.
의심스러운 트래픽 조사:
내 네트워크의 장치가 예상치 못한 외부 목적지와 통신하고 있는지 조사하고 싶습니다. 최근 네트워크 흐름에서 Firewalla에 의해 차단되지 않은 미국 이외의 지역으로의 트래픽을 검색하십시오. 결과를 장치 및 목적지 국가별로 그룹화하십시오. 차단되지 않은 트래픽이 비정상적인 지역으로 향하는 장치가 발견되면, 장치 목록과 상호 참조하여 해당 장치가 무엇인지 식별하고, 관련된 알람이 있는지 확인하십시오. 각 표시된 장치에 대한 위험 평가와 함께 결과를 요약하십시오.
빠른 쿼리
일상적인 모니터링 및 현장 점검을 위한 짧은 프롬프트입니다:
"내 네트워크의 모든 장치를 나열하고, 알 수 없는 MAC 공급업체가 있거나 Firewalla에서 모니터링되지 않는 장치를 표시해 줘."
"내 Firewalla의 모든 허용 규칙을 보여줘. 너무 광범위하게 설정된 규칙이 있어?"
"현재 내 네트워크에서 발생하는 상위 알람 유형은 뭐야? 유형별로 그룹화하고 개수를 알려줘."
"내가 활성화한 Firewalla 차단 목록과 각 목록의 항목 수를 확인해 줘. 중요한 범주가 빠져 있어?"
"지난 24시간 동안 차단된 흐름을 검색하고 목적지 국가별로 그룹화해 줘. 가장 많이 나타나는 국가는 어디야?"
"내 Firewalla 박스 정보를 가져와 줘 — 온라인 상태인지, 어떤 펌웨어 버전을 실행 중인지, 현재 활성 알람은 몇 개인지 알려줘."
도구
도구 | 설명 |
| MSP 계정의 Firewalla 박스 검색 (모델, 펌웨어, 온라인 상태, 장치/규칙/알람 수) |
| 네트워크의 모든 장치 인벤토리 (IP, MAC 공급업체, 장치 유형, 온라인 상태, 모니터링 플래그) |
| 쿼리 필터, 그룹화 및 커서 페이지 매김을 사용한 네트워크 흐름 검색 |
| 쿼리 필터, 그룹화 및 커서 페이지 매김을 사용한 활성 보안 알람 검색 |
| 박스 + 알람 ID별 단일 알람의 전체 세부 정보 가져오기 |
| 구성된 차단/허용 규칙 감사 (작업, 방향, 대상, 범위, 적중 횟수) |
| 차단/허용 대상 목록 나열 (Firewalla 관리 및 사용자 정의) |
| ID별 단일 대상 목록의 메타데이터 가져오기 |
모든 도구는 response_format: "json" | "markdown"을 지원하며 readOnlyHint: true로 주석 처리되어 있습니다.
필수 조건
Firewalla 박스가 MSP 계정에 연결되어 있어야 합니다. 독립형(비플릿) 박스도 MSP API를 사용합니다. 이는 유일하게 지원되는 공개 API입니다.
MSP 개인 액세스 토큰. 다음에서 생성하십시오:
https://<your-subdomain>.firewalla.net에서 MSP 포털에 로그인Account Settings → Personal Access Tokens로 이동
새 토큰을 생성하고 안전한 곳에 저장
자세한 설정 지침은 Getting Started with the Firewalla MSP API를 참조하십시오.
Node.js 18+
설치
git clone https://github.com/productengineered/firewalla-mcp.git
cd firewalla-mcp
npm install
npm run build구성
서버는 두 개의 환경 변수를 읽습니다:
변수 | 설명 | 예시 |
| MSP 하위 도메인 ( |
|
| MSP 계정 설정의 개인 액세스 토큰 |
|
로컬 개발을 위해 .env.example을 .env로 복사하고 값을 입력하십시오:
cp .env.example .env
# edit .env with your real valuesClaude Desktop에서 사용
claude_desktop_config.json(macOS의 경우 보통 ~/Library/Application Support/Claude/claude_desktop_config.json)에 추가하십시오:
{
"mcpServers": {
"firewalla": {
"command": "node",
"args": ["/absolute/path/to/firewalla-mcp/dist/index.js"],
"env": {
"FIREWALLA_MSP_DOMAIN": "yourname.firewalla.net",
"FIREWALLA_MSP_TOKEN": "your-token-here"
}
}
}
}참고: Claude Desktop은 최소한의
PATH로 실행됩니다.node를 찾을 수 없는 경우 Node.js 바이너리의 절대 경로(예:which node의 출력)를 사용하십시오.
구성을 편집한 후 Claude Desktop을 다시 시작하십시오.
Claude Code에서 사용
claude mcp add-json --scope user firewalla '{
"type": "stdio",
"command": "node",
"args": ["/absolute/path/to/firewalla-mcp/dist/index.js"],
"env": {
"FIREWALLA_MSP_DOMAIN": "yourname.firewalla.net",
"FIREWALLA_MSP_TOKEN": "your-token-here"
}
}'다음으로 확인하십시오:
claude mcp list
# firewalla: ... - ✓ Connected새로운 Claude Code 세션에서는 firewalla_* 도구를 자동으로 사용할 수 있습니다.
개발
# Source env for local dev
set -a; source .env; set +a
# Run with auto-reload
npm run dev
# Build
npm run build
# Test with MCP Inspector
npx @modelcontextprotocol/inspector --cli node dist/index.js --method tools/listFirewalla API 문서
라이선스
MIT
Available Tools
8 toolsfirewalla_get_alarmGet Firewalla AlarmARead-onlyIdempotent
Fetch the full detail of a single alarm by gid (box id) + aid (alarm id). Use this after firewalla_search_alarms to drill into one event.
Args:
gid (string, required): Box id (from firewalla_list_boxes).
aid (string, required): Alarm id (from firewalla_search_alarms).
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns the full alarm record, which may include device, remote endpoint, category, timestamps, and any alarm-type-specific detail fields the MSP API surfaces.
| Name | Required | Description | Default |
|---|---|---|---|
| gid | Yes | Box id (from firewalla_list_boxes). | |
| aid | Yes | Alarm id (from firewalla_search_alarms results). Accepts number or string; the API returns numeric ids. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already provide comprehensive behavioral hints (readOnlyHint: true, destructiveHint: false, idempotentHint: true, openWorldHint: true). The description adds valuable context beyond annotations by explaining the purpose of the response_format parameter ('markdown renders human-readable audit tables; json returns structured data suitable for chaining') and describing what the return contains ('full alarm record... may include device, remote endpoint, category, timestamps, and alarm-type-specific detail fields').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly structured with a clear purpose statement upfront, followed by a usage guideline, then parameter context in a formatted Args section, and finally return value information. Every sentence serves a distinct purpose with zero redundancy or wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with comprehensive annotations and full schema coverage, the description provides excellent contextual completeness. It explains the tool's role in the workflow, clarifies parameter sources, describes output format implications, and outlines what information the alarm record contains - all without needing to duplicate what's already in structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all parameters. The description adds minimal additional semantic context beyond the schema - it mentions that aid comes from firewalla_search_alarms results (already in schema) and explains the practical implications of response_format choices. This meets the baseline expectation when schema coverage is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Fetch the full detail of a single alarm') and identifies the required resources (gid and aid). It explicitly distinguishes this tool from its sibling firewalla_search_alarms by stating 'Use this after firewalla_search_alarms to drill into one event,' establishing a clear relationship and differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('Use this after firewalla_search_alarms to drill into one event') and references prerequisite tools for obtaining required parameters (firewalla_list_boxes for gid, firewalla_search_alarms for aid). This creates a clear workflow context and distinguishes it from other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_get_target_listGet Firewalla Target ListARead-onlyIdempotent
Fetch the metadata for a single target list by id.
MSP API limitation: For Firewalla-managed lists (owner="firewalla"), the MSP API does NOT return individual target entries — it returns the summary plus the aggregate count. User-created lists may include a targets array; if so, we surface it.
Use this to answer:
"What's the block mode / source / type of list X?"
"When was list X last updated?"
"How big is list X?" (use the
count/targetCountfield)
Do NOT use this to answer:
"Is domain example.com on list X?" — the entries aren't returned.
"Give me the first N entries of list X." — same reason.
Args:
id (string, required): Target-list id (from firewalla_list_target_lists).
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { id, name, owner, type?, source?, blockMode?, notes?, lastUpdated?, count?: number, // summary count reported by the API targetCount: number, // same as count, or actual targets.length when present targets?: string[] // only populated for user-created lists (rare) }
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Target-list id (from firewalla_list_target_lists). | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering safety and idempotency. The description adds valuable context beyond this: it discloses the MSP API limitation for Firewalla-managed lists (no individual entries returned), clarifies when targets array is populated (user-created lists), and explains the difference between count and targetCount fields. No contradictions with annotations exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with the core purpose. It uses bullet points for usage guidelines, separates arguments and returns clearly, and avoids redundant information. Every sentence adds value, such as explaining API limitations and field meanings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (fetching metadata with API limitations), the description is complete. It covers purpose, usage, behavioral nuances (like API constraints), parameters, and return structure in detail. Although there's no output schema, the description provides a comprehensive return object specification, compensating adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear descriptions for both parameters (id and response_format). The description adds minimal extra semantics: it reiterates that id comes from firewalla_list_target_lists (already in schema) and briefly explains response_format options (default and use cases). This meets the baseline for high schema coverage without significant added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the verb 'fetch' and resource 'metadata for a single target list by id', making the purpose specific. It distinguishes from sibling tools like firewalla_list_target_lists by focusing on a single list rather than listing all, and clarifies limitations compared to potential expectations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool (e.g., to answer questions about block mode, source, type, last updated, or size) and when not to use it (e.g., to check if a domain is on the list or get entries). It also references the sibling tool firewalla_list_target_lists for obtaining the id parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_list_boxesList Firewalla BoxesARead-onlyIdempotent
Discover the Firewalla boxes linked to this MSP account. This is the entry point for every audit — the returned gid is required by other tools.
Use this to answer:
"Is my box online and reporting in?"
"What firmware version is it running?"
"How many active devices, rules, alarms are there right now?"
Args:
group (string, optional): Filter to a specific group id.
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { count: number, boxes: Array<{ gid: string, // box id — save this, other tools need it name: string, model: string, // e.g. "gold_plus" mode: string, // routing mode version: string, // firmware online: boolean, publicIP?: string, lastSeen?: number, // epoch seconds — not always populated license?: string, location?: string, deviceCount: number, ruleCount: number, alarmCount: number, // currently-active alarms group?: { id, name } }> }
Audit framing:
Offline box → can't observe current state; surface it.
High alarmCount → follow up with firewalla_search_alarms.
publicIP exposed unexpectedly → investigate with firewalla_search_flows.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Filter to boxes in a specific group id. Omit to list all boxes on the account. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering the safety profile. The description adds valuable behavioral context beyond annotations: it explains the audit framing logic, clarifies that 'lastSeen' is 'not always populated', and provides guidance on interpreting results and next steps based on findings like offline boxes or high alarm counts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, usage questions, Args, Returns, audit framing) and efficiently conveys necessary information. While comprehensive, every section earns its place by adding value, though the Args section could be more concise given the schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity as an audit entry point with rich return data and sibling relationships, the description provides complete context. It explains the tool's role in the ecosystem, provides detailed return structure documentation (compensating for no output schema), and includes audit framing that guides interpretation and next steps with sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters well-documented in the schema. The description's Args section essentially repeats what's in the schema without adding significant semantic context beyond what's already structured. The baseline of 3 is appropriate when the schema does the heavy lifting for parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Discover'), resource ('Firewalla boxes linked to this MSP account'), and scope ('entry point for every audit'). It distinguishes from siblings by emphasizing this tool provides the essential 'gid' needed by other tools, unlike more specific tools like firewalla_search_alarms or firewalla_list_devices.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('entry point for every audit'), when to follow up with alternatives ('High alarmCount → follow up with firewalla_search_alarms', 'publicIP exposed unexpectedly → investigate with firewalla_search_flows'), and includes audit framing questions that guide appropriate usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_list_devicesList Firewalla DevicesARead-onlyIdempotent
Inventory every device Firewalla tracks — the "who's on my network right now" primitive.
Use this to answer:
"Are there any unknown/rogue devices on my network?"
"Which devices aren't being monitored?"
"What's the MAC vendor breakdown across my network?"
"Any router-class devices I didn't expect?"
Args:
box (string, optional): Filter to devices on a specific box gid.
online_only (boolean, optional): Drop offline devices client-side.
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { count: number, // devices after client-side filtering total: number, // devices returned by the API (pre-filter) devices: Array<{ id: string, // typically MAC gid: string, // box the device is attached to name: string, ip: string, mac?: string, macVendor?: string, ipReserved?: boolean, online: boolean, network?: { id, name }, deviceType?: string, // e.g. "phone", "computer", "iot" isRouter?: boolean, isFirewalla?: boolean, monitoring?: boolean, // false = device excluded from monitoring totalDownload?: number, // bytes (lifetime) totalUpload?: number }> }
Audit framing:
Unknown macVendor → possible squatter or spoofed MAC.
monitoring=false → device is excluded from Firewalla's visibility; review whether that's intentional.
Unexpected isRouter=true → shadow router on the LAN.
ipReserved=false on a server that should have a static lease → risk of address drift.
| Name | Required | Description | Default |
|---|---|---|---|
| box | No | Filter to devices attached to a specific box gid. | |
| online_only | No | If true, drop offline devices from the response. Client-side filter — the API returns all devices either way. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering safety and idempotency. The description adds valuable behavioral context about client-side filtering ('online_only' drops offline devices client-side), output format implications, and audit interpretations that help the agent understand how to process and interpret results beyond basic safety information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (purpose, usage questions, args, returns, audit framing) and every sentence adds value. While somewhat lengthy, it's efficiently organized with bullet points and structured returns documentation, making it easy to parse. Minor deduction for being slightly verbose in the returns section.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of an output schema, the description provides comprehensive context including detailed return structure documentation, audit interpretation guidance, and clear usage scenarios. With annotations covering safety aspects and the description filling in behavioral and interpretive gaps, this provides complete context for the agent to effectively use this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all three parameters. The description adds minimal additional context beyond what's in the schema (e.g., 'client-side filter' for online_only, output format implications), but doesn't provide significant semantic value beyond the structured documentation. Baseline 3 is appropriate given complete schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Inventory every device Firewalla tracks' and positions it as the 'who's on my network right now' primitive. It distinguishes from siblings by focusing on device inventory rather than alarms, rules, flows, or boxes, making the scope specific and differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage scenarios with bullet points answering specific questions like 'Are there any unknown/rogue devices on my network?' and 'Which devices aren't being monitored?'. It also includes an 'Audit framing' section that guides interpretation of results, effectively telling the agent when and how to use this tool for network auditing purposes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_list_rulesList Firewalla RulesARead-onlyIdempotent
Audit configured block / allow rules. Read-only — this tool does NOT pause, resume, create, or modify rules.
Use this to answer:
"Do I have any allow rules that bypass Firewalla's default blocks?"
"Which rules haven't fired in 90 days (candidates to remove)?"
"Are my block rules scoped to the right device/group?"
"Any rules with action=allow and broad scope?"
Args:
query (string, optional): Firewalla query-grammar filter (pass-through). Examples:
action:allow,status:paused,target.type:domain.response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { count: number, rules: Array<{ id: string, gid: string, action: string, // "block" | "allow" | "time_limit" | … direction?: string, // "outbound" | "inbound" | "bidirection" status?: string, // "active" | "paused" | "disabled" target: { type, value, dnsOnly?, port? }, scope?: { type?, value? }, notes?: string, hit?: { count?, lastHitTs? }, ts?: number, updateTs?: number }> }
Audit framing:
action=allow with scope=global → overly permissive, investigate.
status=paused with no notes → someone disabled a rule and didn't document why.
hit.count=0 & old updateTs → stale rule, candidate for removal.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Firewalla query string (pass-through). See Firewalla docs for the grammar — supports filters like `device.mac:AA:BB:CC:DD:EE:FF`, `blocked:true`, `region:CN`, `ts:>1700000000`, etc. Omit to match everything. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable behavioral context beyond the annotations. While annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, the description adds the 'audit framing' section that explains how to interpret the results for security analysis. This provides practical guidance on what patterns to look for in the returned data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exceptionally well-structured and front-loaded. The first sentence establishes the core purpose, followed immediately by usage examples, parameter details, return format, and audit guidance. Every section serves a distinct purpose with zero wasted text, making it easy for an AI agent to parse and understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the comprehensive annotations, detailed input schema with 100% coverage, and the rich description that includes usage examples, parameter context, return format explanation, and audit guidance, this description provides complete context for a read-only audit tool. The absence of an output schema is compensated by the detailed return structure documentation in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline would be 3. However, the description adds meaningful context by providing example queries in the 'Use this to answer' section that illustrate practical applications of the query parameter. The audit framing section also helps users understand how to interpret results based on parameter combinations.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('audit configured block/allow rules') and distinguishes it from siblings by explicitly stating what it does NOT do ('does NOT pause, resume, create, or modify rules'). This makes it immediately clear this is a read-only audit tool versus other Firewalla tools that might modify rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides excellent usage guidance with four specific example questions this tool can answer, giving concrete scenarios for when to use it. It also explicitly distinguishes from alternatives by stating what it doesn't do, helping users understand when NOT to use this tool versus modification tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_list_target_listsList Firewalla Target ListsARead-onlyIdempotent
List the block/allow target lists available on this MSP account — both Firewalla-managed ("global") and user-defined.
Use this to answer:
"Which block lists is Firewalla enforcing against?"
"Have I added any custom target lists, and what are their owners?"
"What categories (ad, tracker, malware, …) are covered?"
This endpoint returns summaries (including target count per list);
call firewalla_get_target_list for the actual targets array.
Args:
owner (string, optional): Filter by owner (e.g. 'global').
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { count: number, // number of target lists targetLists: Array<{ id: string, name: string, owner: string, // "global" | user id type?: string, // e.g. "ad", "tracker", "malware", "custom" source?: string, // upstream feed source (Firewalla-managed lists) count?: number, // number of entries in the list blockMode?: string, // e.g. "dns" | "ip" beta?: boolean, notes?: string, lastUpdated?: number }> }
Audit framing:
Custom lists (owner != global) without notes → undocumented intent.
blockMode=dns only, but target includes raw IPs → mismatch, investigate.
Zero-count list → may be stale / never populated.
| Name | Required | Description | Default |
|---|---|---|---|
| owner | No | Filter by owner. Common values: 'global' (Firewalla-managed), or a specific user id. Omit to list all. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true. The description adds valuable context beyond this: it explains the distinction between summaries vs. detailed targets, provides audit framing guidance (e.g., 'Custom lists without notes → undocumented intent'), and mentions output format implications. While it doesn't cover rate limits or authentication needs, it adds significant behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: it starts with the core purpose, provides usage examples in bullet points, explains the relationship with a sibling tool, documents parameters and returns, and ends with audit framing. Every sentence serves a clear purpose with zero waste, and information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, rich annotations (readOnly, idempotent, openWorld), and 100% schema coverage, the description is complete. It explains the tool's purpose, usage guidelines, relationship with siblings, parameter semantics (though schema covers this), return structure, and even includes audit framing for interpretation. No output schema exists, but the description thoroughly documents the return format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters. The description adds minimal value beyond the schema: it mentions the 'owner' filter can be used to list all (implied by omission) and provides example values, but doesn't add substantial semantic context. This meets the baseline of 3 when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and resource ('block/allow target lists available on this MSP account'), specifying both Firewalla-managed ('global') and user-defined lists. It distinguishes this tool from its sibling 'firewalla_get_target_list' by noting that this returns summaries while the sibling provides the actual targets array.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides three example questions this tool can answer, giving clear context for when to use it. It also distinguishes from the sibling 'firewalla_get_target_list' by stating this returns summaries while that tool provides the actual targets array, offering explicit guidance on alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_search_alarmsSearch Firewalla AlarmsARead-onlyIdempotent
Search active Firewalla alarms with the MSP query grammar. This is the primary tool for "what security events are happening right now?" audits.
Use this to answer:
"Any alarms from devices not in a known group?"
"How many alarms of type X in the last 24h, grouped by device?"
"Which remote countries are triggering the most alarms?"
"Any alarms relating to a specific device (by MAC)?"
Args:
query (string, optional): Firewalla query grammar. Examples:
type:1,device.mac:AA:BB:CC:DD:EE:FF,remote.country:CN,ts:>1700000000.group_by (string, optional): e.g.
device,type,remote.country.sort_by (string, optional): e.g.
ts:desc(default),ts:asc.limit (number, 1–500, default 200).
cursor (string, optional): pagination cursor from a prior response.
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { count: number, // items in this page next_cursor?: string, // echo back to fetch the next page alarms: Array<{ aid, gid, type, ts, message, status?, device?: { id?, name?, ip? }, remote?: { ip?, country?, name?, region?, category? } }> }
Audit framing:
Alarm from an unknown MAC (device.id not in firewalla_list_devices) → rogue device.
Repeated alarms to the same remote.country → likely a single piece of malware, check firewalla_list_rules.
When counts get big, use group_by=type first for a birds-eye view, then drill.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Firewalla query string (pass-through). See Firewalla docs for the grammar — supports filters like `device.mac:AA:BB:CC:DD:EE:FF`, `blocked:true`, `region:CN`, `ts:>1700000000`, etc. Omit to match everything. | |
| group_by | No | Group results by one or more fields (comma-separated). Examples: `device`, `device,domain`, `region`. When set, results are aggregated per group. | |
| sort_by | No | Sort expression. Format: `<field>:<asc|desc>`. Common: `ts:desc` (default, newest first), `ts:asc` (oldest first), `download:desc` (biggest flows first). | |
| limit | No | Maximum results per page (1–500, default 200). Smaller values are recommended when auditing — easier to review. | |
| cursor | No | Pagination cursor echoed from a prior response's `next_cursor`. Omit for the first page. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, the description adds valuable behavioral context beyond these annotations. It explains the tool's role in security audits, provides guidance on handling large result sets ('When counts get big, use group_by=type first'), and describes pagination behavior through the cursor parameter. The description doesn't contradict annotations and adds meaningful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with clear sections: purpose statement, usage examples, parameter details, return format, and audit guidance. Every sentence serves a specific purpose—no wasted words. The information is front-loaded with the core purpose, followed by progressively detailed guidance. The structure supports both quick understanding and deep reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, security audit focus) and the absence of an output schema, the description provides excellent contextual completeness. It fully documents the return structure in the 'Returns' section, explains pagination mechanics, provides audit-specific guidance, and references sibling tools for follow-up actions. The description compensates fully for the lack of output schema and provides comprehensive operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline would be 3, but the description adds significant value beyond the schema. The 'Args' section provides concrete query examples (`type:1`, `device.mac:AA:BB:CC:DD:EE:FF`, etc.) that illustrate the query grammar more vividly than the schema's description. It also explains the practical implications of parameters like 'group_by' for aggregation and 'response_format' for different use cases (human-readable vs. chaining).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool's purpose as 'Search active Firewalla alarms with the MSP query grammar' and positions it as 'the primary tool for "what security events are happening right now?" audits.' This clearly distinguishes it from sibling tools like firewalla_get_alarm (likely for single alarm retrieval) and firewalla_search_flows (for flow data rather than alarms), providing specific verb+resource+scope differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool through concrete example questions ('Any alarms from devices not in a known group?', 'How many alarms of type X in the last 24h, grouped by device?', etc.) and includes an 'Audit framing' section with specific scenarios (e.g., 'Alarm from an unknown MAC → rogue device'). It also implicitly suggests alternatives by referencing sibling tools like firewalla_list_devices and firewalla_list_rules for follow-up actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firewalla_search_flowsSearch Firewalla FlowsARead-onlyIdempotent
Search network flows observed by Firewalla with the MSP query grammar. Use this to inspect what's actually happening on the wire.
Use this to answer:
"Any outbound flows to region:CN that were NOT blocked?"
"Top talkers by download volume over the last 24h?"
"Which devices have made the most connections to blocklisted categories?"
"Are there any inbound flows from the public internet that shouldn't exist?"
"Flows from device X in the last hour?"
Args:
query (string, optional): Firewalla query grammar. Examples:
blocked:true,region:CN,direction:inbound,device.mac:AA:BB:CC:DD:EE:FF,category:malware,ts:>1700000000, combined with AND/OR.group_by (string, optional): e.g.
device,device,destination,region.sort_by (string, optional): e.g.
ts:desc(default),download:desc.limit (number, 1–500, default 200).
cursor (string, optional): pagination cursor from a prior response.
response_format ('markdown' | 'json'): Output format (default: markdown).
Returns: { count: number, // items in this page next_cursor?: string, flows: Array<{ ts, gid, protocol, direction, block?, blockType?, download?, upload?, total?, duration?, count?, device?: { id, ip?, name?, network? }, source?: { id?, ip?, name?, port? }, destination?: { id?, ip?, name?, port? }, // Flow-level classification fields (NOT nested under destination): country?, region?, domain?, category? }> }
Audit framing:
Start broad with
sort_by=download:descto find top bandwidth users.Narrow with
querywhen you've found a device/region of interest.block=falseflows to a category:malware destination = missed block, investigate rules.Use
group_byfor aggregates; use limit=50 or so for fine-grained review.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Firewalla query string (pass-through). See Firewalla docs for the grammar — supports filters like `device.mac:AA:BB:CC:DD:EE:FF`, `blocked:true`, `region:CN`, `ts:>1700000000`, etc. Omit to match everything. | |
| group_by | No | Group results by one or more fields (comma-separated). Examples: `device`, `device,domain`, `region`. When set, results are aggregated per group. | |
| sort_by | No | Sort expression. Format: `<field>:<asc|desc>`. Common: `ts:desc` (default, newest first), `ts:asc` (oldest first), `download:desc` (biggest flows first). | |
| limit | No | Maximum results per page (1–500, default 200). Smaller values are recommended when auditing — easier to review. | |
| cursor | No | Pagination cursor echoed from a prior response's `next_cursor`. Omit for the first page. | |
| response_format | No | Output format. 'markdown' (default) renders human-readable audit tables. 'json' returns structured data suitable for chaining into another tool call. | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering safety and idempotency. The description adds valuable behavioral context beyond annotations: it explains the tool's primary use for audit/inspection ('inspect what's actually happening on the wire'), provides strategic guidance in the 'Audit framing' section, and hints at typical workflows (e.g., 'Start broad... Narrow with query'). It doesn't mention rate limits or authentication needs, but adds meaningful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections: purpose statement, usage examples, parameter details, return format, and audit guidance. Every sentence adds value, though it's somewhat lengthy (which is justified given the tool's complexity). The information is front-loaded with the core purpose and usage examples immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex search tool with 6 parameters and no output schema, the description provides exceptional completeness. It includes: clear purpose, specific usage examples, detailed parameter explanations with examples, return format documentation, and strategic audit guidance. The combination of thorough parameter coverage in the schema and rich contextual information in the description makes this fully self-contained for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds value by providing concrete query examples in the 'Args' section (e.g., 'blocked:true', 'region:CN', 'device.mac:AA:BB:CC:DD:EE:FF') and explaining the purpose of each parameter in context. It also clarifies the relationship between parameters in the 'Audit framing' section (e.g., 'Use group_by for aggregates; use limit=50 or so for fine-grained review').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool's purpose: 'Search network flows observed by Firewalla with the MSP query grammar. Use this to inspect what's actually happening on the wire.' It clearly distinguishes this from sibling tools like firewalla_get_alarm or firewalla_list_devices by focusing on flow inspection rather than alarms, devices, or rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool through the 'Use this to answer' section with five concrete examples (e.g., 'Any outbound flows to region:CN that were NOT blocked?', 'Top talkers by download volume over the last 24h?'). The 'Audit framing' section offers strategic advice on starting broad and narrowing down, plus specific use cases like investigating missed blocks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
firewalla_get_alarm - First observed
firewalla_get_target_list - First observed
firewalla_list_boxes - First observed
firewalla_list_devices - First observed
firewalla_list_rules - First observed
firewalla_list_target_lists - First observed
firewalla_search_alarms - First observed
firewalla_search_flows
TDQS
Scored across 8 tools
Each tool has a distinct purpose targeting specific Firewalla resources: list_* tools fetch collections, get_* tools retrieve single items, and search_* tools query with filters. There is no overlap in functionality; for example, firewalla_get_alarm and firewalla_search_alarms serve complementary drill-down and overview roles without ambiguity.
All tools follow a consistent verb_noun pattern with the prefix 'firewalla_' and snake_case throughout. Verbs are clear and standardized: 'list' for collections, 'get' for single items, and 'search' for filtered queries. This predictability makes it easy to understand each tool's intent at a glance.
With 8 tools, the server is well-scoped for network security auditing. It covers essential resources (boxes, devices, rules, alarms, flows, target lists) without being overwhelming. Each tool earns its place by addressing a distinct aspect of Firewalla monitoring, fitting the domain's complexity appropriately.
The toolset provides comprehensive read-only coverage for auditing Firewalla MSP data, including inventory, rules, alarms, and network flows. Minor gaps exist, such as no tools for modifying rules or managing devices, but these are consistent with an audit-focused server, and agents can work around this by using the provided search and list tools effectively.
Maintenance
Related MCP Connectors
Read-only MCP access to a documented IT fleet: state, changes, posture. 15 tools.
Read-only local AI advice, shared reports and website audits. No PC scan or local actions.
Read-only MCP server for AIStatusDashboard status, incidents, metrics, and fallback recommendations.
Read-only MCP server for turva.dev's published service catalog, pricing and contact details. Five tools return JSON, including dated agent-readiness and security evidence with verification links. Connect over Streamable HTTP without an API key. The server answers questions about turva.dev and does not scan other websites or run audits.
Related MCP Servers
AlicenseAqualityDmaintenanceRead-only MCP server that allows AI assistants to query and monitor KVM Fleet devices, audit logs, and console sessions through the official REST API.59 npm1MIT- FlicenseNot gradedqualityBmaintenanceA read-only MCP server that gives Claude Code secure, non-invasive access to infrastructure logs, service status, metrics, Ansible facts, and Docker state via SSH, with a strict command allowlist and no write operations.-
- AlicenseAqualityCmaintenanceA read-only MCP server that allows Claude Code to securely access Zulip chat messages, streams, topics, and user information without modification capabilities.9MIT
- AlicenseNot gradedqualityBmaintenanceA read-only MCP server that gives Claude safe access to Kubernetes clusters, enabling listing, describing, and monitoring resources without mutation risks and with secret masking.1MIT