Skip to main content
Glama
amittell

firewalla-mcp-server

Firewalla MCP Server

npm version

A Model Context Protocol (MCP) server that provides real-time access to Firewalla firewall data through 24 read-only tools, plus 11 opt-in write tools, compatible with any MCP client.

Why Firewalla MCP Server?

Simple Network Security Integration

  • 24 read-only tools for network monitoring and analysis: 19 Direct API Endpoints + 5 Convenience Wrappers

  • 11 write tools, off unless FIREWALLA_ENABLE_WRITE_TOOLS=true, so by default nothing can change your box

  • Advanced Search with query syntax and filters

  • Clean, Verified Architecture with corrected API schemas

Related MCP server: Firewalla MCP Server

Features

  • Real-time Firewall Data: Query security alerts, network flows, and device status

  • Security Analysis: Get insights on threats, blocked attacks, and network anomalies

  • Bandwidth Monitoring: Track top bandwidth consumers and usage patterns

  • Rule Management: View firewall rules; with the write tools, create, pause, resume and delete them

  • Target Lists: View target lists; with the write tools, create, update and delete them

  • Search Tools: Query syntax with filters and logical operators

Client Setup Guides

Client

Quick Start

Full Guide

Claude Desktop

npm i -g firewalla-mcp-server → Configure MCP

Setup Guide

Claude Code

npm i -g firewalla-mcp-server → CLI integration

Setup Guide

VS Code

Install MCP extension → Configure server

Setup Guide

Cursor

Install Claude Code → VSIX method

Setup Guide

Roocode

Install MCP support → Configure server

Setup Guide

Cline

Configure in VS Code → Enable MCP

Setup Guide

Open WebUI

mcpo, or a native MCP connection over HTTP

Setup Guide

How It Works

Claude Desktop/Code ↔ MCP Server ↔ Firewalla API

The MCP server acts as a bridge between Claude and your Firewalla firewall, translating Claude's requests into Firewalla API calls and returning the results in a format Claude can understand.

Prerequisites

  • Node.js 18+ and npm. On Node 18-22, npm prints an EBADENGINE warning for geoip-lite 2.x, which declares Node 24 for its database update script; the lookups the server uses run on Node 18 and later, and CI tests 18, 20, 22 and 24.

  • Firewalla MSP account with API access

  • Your Firewalla device online and connected

Quick Start

1. Installation

# Install globally
npm install -g firewalla-mcp-server

# Or install locally in your project
npm install firewalla-mcp-server

Option B: Use Docker

Warning: Not for production use – secrets visible in process list

The examples below pass credentials directly in the command line, which exposes them to process listing and shell history. For production use, consider these secure alternatives:

  • Use --env-file with a .env file: docker run --env-file .env ...

  • Set environment variables in your shell before running Docker

  • Use Docker secrets for orchestration environments

Stdio Transport (Default - for Claude Desktop integration):

# Using Docker Hub image (minimal config)
docker run -it --rm \
  -e FIREWALLA_MSP_TOKEN=your_token \
  -e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
  amittell/firewalla-mcp-server

# Or with optional box filter
docker run -it --rm \
  -e FIREWALLA_MSP_TOKEN=your_token \
  -e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
  -e FIREWALLA_BOX_ID=your_box_gid \
  amittell/firewalla-mcp-server

# Or build locally
docker build -t firewalla-mcp-server .
docker run -it --rm \
  -e FIREWALLA_MSP_TOKEN=your_token \
  -e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
  firewalla-mcp-server

# Recommended: Using env file (more secure)
docker run -it --rm --env-file .env amittell/firewalla-mcp-server

HTTP Transport (for standalone Docker containers and external access):

# Run with HTTP transport on port 3000
docker run -d --name firewalla-mcp \
  -p 3000:3000 \
  -e MCP_TRANSPORT=http \
  -e MCP_HTTP_PORT=3000 \
  -e MCP_HTTP_BEARER_TOKEN=a_long_random_secret \
  -e FIREWALLA_MSP_TOKEN=your_token \
  -e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
  amittell/firewalla-mcp-server

# Add FIREWALLA_BOX_ID if you want to filter to a specific box
# -e FIREWALLA_BOX_ID=your_box_gid \

# The server will be accessible at http://localhost:3000/mcp, and clients
# send the header: Authorization: Bearer a_long_random_secret

# Using env file (recommended)
docker run -d --name firewalla-mcp \
  -p 3000:3000 \
  --env-file .env \
  amittell/firewalla-mcp-server

# For docker-compose
cat > docker-compose.yml << EOF
version: '3.8'
services:
  firewalla-mcp:
    image: amittell/firewalla-mcp-server
    ports:
      - "3000:3000"
    environment:
      - MCP_TRANSPORT=http
      - MCP_HTTP_PORT=3000
      # The image already listens on every interface of the container
      - MCP_HTTP_HOST=0.0.0.0
      - MCP_HTTP_BEARER_TOKEN=\${MCP_HTTP_BEARER_TOKEN}
      # Other containers reach it as http://firewalla-mcp:3000/mcp
      - MCP_HTTP_ALLOWED_HOSTS=firewalla-mcp
      - FIREWALLA_MSP_TOKEN=\${FIREWALLA_MSP_TOKEN}
      - FIREWALLA_MSP_ID=\${FIREWALLA_MSP_ID}
      # Optional: filter to specific box
      # - FIREWALLA_BOX_ID=\${FIREWALLA_BOX_ID}
    restart: unless-stopped
EOF

docker-compose up -d

The image sets MCP_HTTP_HOST=0.0.0.0 so that a published port reaches the server. -p 3000:3000 publishes it on every interface of the Docker host, so anyone on your network can reach it: set MCP_HTTP_BEARER_TOKEN (for example openssl rand -hex 32) whenever you publish the port, or publish it on this machine only with -p 127.0.0.1:3000:3000. The server answers only requests whose Host header is localhost, 127.0.0.1 or [::1], so a client that connects by another name, such as the host's LAN address or a compose service name, needs that name in MCP_HTTP_ALLOWED_HOSTS. See HTTP transport security.

Option C: Install from source

git clone https://github.com/amittell/firewalla-mcp-server.git
cd firewalla-mcp-server
npm install
npm run build

2. Configuration

Create a .env file with your Firewalla credentials:

# Required
FIREWALLA_MSP_TOKEN=your_msp_access_token_here
FIREWALLA_MSP_ID=yourdomain.firewalla.net

# Optional - filters all queries to a specific box
# FIREWALLA_BOX_ID=your_box_gid_here

# Optional - default box for single-box operations, without filtering queries
# FIREWALLA_DEFAULT_BOX_ID=your_box_gid_here

# Optional - register the 11 write tools (default: off, read-only)
# FIREWALLA_ENABLE_WRITE_TOOLS=true

Getting Your Credentials:

  1. Log into your Firewalla MSP portal at https://yourdomain.firewalla.net

  2. Your MSP ID is the full domain (e.g., company123.firewalla.net)

  3. Generate an access token in API settings

  4. (Optional) Find your Box GID in device settings to filter queries to a specific box, or retrieve available boxes using the get_boxes tool

Box ID is optional. Without FIREWALLA_BOX_ID, queries cover every box on the account. The few operations that act on one box (get_specific_alarm, archive_alarm, mute_alarm, delete_alarm, create_rule, rename_device) take a gid argument, and without one they use FIREWALLA_BOX_ID, then FIREWALLA_DEFAULT_BOX_ID, then the account's only box. On an account with several boxes and none of those set, get_specific_alarm, archive_alarm, mute_alarm and delete_alarm check each box, and create_rule and rename_device refuse and list the boxes.

Test mode. MCP_TEST_MODE=true starts the server without credentials, to check that it starts: it uses a dummy token, API (https://test.firewalla.net) and box instead of your settings, always runs on stdio, and cannot read your Firewalla data. With NODE_ENV=production the server refuses test mode and exits with code 1. The Docker image sets NODE_ENV=production, so pass -e NODE_ENV=development along with -e MCP_TEST_MODE=true.

Transport Configuration

The MCP server supports two transport modes:

Stdio Transport (Default): Standard input/output communication for Claude Desktop and similar MCP clients

MCP_TRANSPORT=stdio

HTTP Transport: HTTP server mode for Docker containers, MCP orchestrators, and external access

MCP_TRANSPORT=http
MCP_HTTP_PORT=3000          # Default: 3000
MCP_HTTP_PATH=/mcp          # Default: /mcp. Other paths get 404
MCP_HTTP_HOST=127.0.0.1     # Address to listen on, no port (::1 or [::1] for IPv6). Default: 127.0.0.1 (0.0.0.0 in the Docker image)
MCP_HTTP_BEARER_TOKEN=      # When set, clients must send Authorization: Bearer <token>
MCP_HTTP_ALLOWED_HOSTS=     # More Host header names to accept, comma-separated
MCP_HTTP_ALLOWED_ORIGINS=   # Browser origins to accept, comma-separated, e.g. http://localhost:6274

HTTP transport security: every request can spend your MSP token, so the HTTP server follows the security rules of the MCP transport specification:

  • It listens on 127.0.0.1, this machine only. Set MCP_HTTP_HOST=0.0.0.0 (or one address) to accept other machines, and set MCP_HTTP_BEARER_TOKEN with it: the server logs a warning when it listens beyond loopback without a token.

  • With MCP_HTTP_BEARER_TOKEN set, a request without Authorization: Bearer <token> gets 401.

  • A request whose Host header is not localhost, 127.0.0.1, [::1], the MCP_HTTP_HOST address or a name in MCP_HTTP_ALLOWED_HOSTS gets 403. This stops DNS rebinding, where a web page points its own domain name at your machine.

  • A request with an Origin header, which browsers send, gets 403 unless the origin is in MCP_HTTP_ALLOWED_ORIGINS. Non-browser MCP clients send no Origin and are not affected. An allowed origin gets CORS headers on every answer, refusals such as a 401 for a missing token included, so a web page on it can call the server and read why a request was refused.

  • The MCP endpoint is the MCP_HTTP_PATH path exactly: /mcp, /mcp/ and either with a query string. Any other path, such as /mcpx or /mcp/x, gets 404, CORS preflight requests included.

  • A request body may be at most 1 MB, and a client has 10 seconds to send the headers and 30 seconds for the whole request. An answer given without reading the body (401, 403, 404, 405, a malformed session ID, a CORS preflight) closes the connection, so a client cannot hold one open by sending the body slowly.

  • A session ends when the client sends DELETE, after MCP_SESSION_IDLE_TIMEOUT_MS without a request (default 30 minutes), or when the server restarts. A request with the ID of a session the server does not hold gets 404, and the transport specification has the client start a new session with a new initialize. A request with no session ID, other than initialize, gets 400.

When to use HTTP transport:

  • Running in Docker containers independently

  • Accessing from MCP orchestrators (e.g., open-webui)

  • Multiple clients need to connect to the same server instance

  • Network-based access to the MCP server

When to use stdio transport:

  • Claude Desktop integration (default)

  • Claude Code CLI integration

  • Single-process MCP client setups

  • Standard MCP client configurations

3. Build and Start

npm run build
npm run mcp:start

4. Connect Claude Desktop

Add this configuration to your Claude Desktop claude_desktop_config.json:

If installed via npm

{
  "mcpServers": {
    "firewalla": {
      "command": "npx",
      "args": ["firewalla-mcp-server"],
      "env": {
        "FIREWALLA_MSP_TOKEN": "your_msp_access_token_here",
        "FIREWALLA_MSP_ID": "yourdomain.firewalla.net",
        "FIREWALLA_BOX_ID": "your_box_gid_here"
      }
    }
  }
}

If using Docker

{
  "mcpServers": {
    "firewalla": {
      "command": "docker",
      "args": ["run", "-i", "--rm", 
        "-e", "FIREWALLA_MSP_TOKEN=your_token",
        "-e", "FIREWALLA_MSP_ID=yourdomain.firewalla.net",
        "-e", "FIREWALLA_BOX_ID=your_box_gid",
        "amittell/firewalla-mcp-server"
      ]
    }
  }
}

If installed from source

{
  "mcpServers": {
    "firewalla": {
      "command": "node",
      "args": ["/full/path/to/firewalla-mcp-server/dist/server.js"],
      "env": {
        "FIREWALLA_MSP_TOKEN": "your_msp_access_token_here",
        "FIREWALLA_MSP_ID": "yourdomain.firewalla.net",
        "FIREWALLA_BOX_ID": "your_box_gid_here"
      }
    }
  }
}

Config file locations:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

5. Next Steps

Usage Examples

Step-by-Step First Use

1. Verify Connection After completing the setup, verify the MCP server is working:

# Start the server
npm run mcp:start

# You should see output like:
# MCP Server starting...
# Firewalla client initialized
# Server ready on stdio transport

2. Test with Claude Open Claude Desktop and try these starter queries:

Basic Health Check:

"Can you check my Firewalla status and show me a summary?"

This uses: firewall_summary resource + get_simple_statistics tool

Security Overview:

"What security alerts do I have? Show me the 5 most recent ones."

This uses: get_active_alarms tool with limit parameter

Practical Workflows

Daily Security Review:

"Give me today's security report. Include:
1. Any new security alerts
2. Top 3 devices using bandwidth
3. Any devices that went offline
4. Status of critical firewall rules"

Investigating Suspicious Activity:

"I noticed unusual traffic. Can you:
1. Show me all security and abnormal upload alarms from the last 4 hours
2. Find any blocked connections to external IPs
3. Check which devices had the most network activity"

Network Troubleshooting:

"A device seems to have connectivity issues. Can you:
1. Check if device 192.168.1.100 is online
2. Show its recent network flows
3. See if any rules are blocking its traffic"

Bandwidth Investigation:

"Our internet is slow. Help me find the cause:
1. Show top 10 bandwidth users in the last hour
2. Look for any devices with unusual upload/download patterns
3. Check for any streaming or video traffic"

Advanced Search Examples

Find Specific Threats:

search for: security activity alarms from IP range 10.0.0.* in the last 24 hours

Uses: search_alarms with query: "type:1 AND device.ip:10.0.0. AND ts:>=<unix time 24 hours ago>"*

Analyze Rule Effectiveness:

"Show me firewall rules that blocked the most connections this week"

Uses: get_network_rules + search_flows for blocked traffic analysis

Device Behavior Analysis:

"Find all devices that were online yesterday but are offline now"

Uses: search_devices with temporal queries + get_offline_devices

Troubleshooting Common Issues

Connection Problems: If you get authentication errors:

  1. Verify your .env file has correct credentials

  2. Check your MSP token hasn't expired

  3. Confirm your Box ID is the full GID format

Empty Results: If queries return no data:

  1. Check your Firewalla is online and reporting

  2. Verify the time range isn't too narrow

  3. Try broader search terms first

Performance Issues: If responses are slow:

  1. Reduce the limit parameter in queries

  2. Use more specific time ranges

  3. Check your network connection to the MSP API

Available Tools (24 read-only, 11 opt-in write tools)

Core Tools

  • Security: Get alarms, analyze threats

  • Network: Monitor traffic flows, track bandwidth usage

  • Devices: Check device status, find offline devices

  • Rules: View firewall rules and their summary

  • Search: Advanced search across all data types

  • Analytics: Statistics, trends, and geographic analysis

  • Target Lists: View security target lists

Quick Reference

Security: get_active_alarms, get_specific_alarm
Network: get_flow_data, get_recent_flow_activity, get_bandwidth_usage
Devices: get_device_status, get_offline_devices, get_boxes
Rules: get_network_rules, get_network_rules_summary
Target lists: get_target_lists, get_specific_target_list
Search: search_flows, search_alarms, search_rules, search_devices, search_target_lists
Analytics: get_simple_statistics, get_statistics_by_region, get_statistics_by_box, get_flow_insights, get_flow_trends, get_alarm_trends, get_rule_trends
Write (opt-in): create_rule, delete_rule, pause_rule, resume_rule, create_target_list, update_target_list, delete_target_list, rename_device, archive_alarm, mute_alarm, delete_alarm

Response format

Every read-only tool (readOnlyHint: true) takes an optional response_format. json, the default, returns the compact JSON response. markdown returns the same response as markdown for reading: a heading with the tool name and the number of records in each list, the fields as bullet lists, and each list of records as a table of at most 8 columns and DEFAULT_PAGE_SIZE rows (default 100). Fields with one value in every record are stated once above the table, fields that are not columns are named below it, and a list of one record is shown as a list of its fields. The last line says what the view leaves out (the response's meta block, rows past the cap, cells cut at 120 characters); response_format: json returns all of it. Values and field names are escaped, since device names, domains and alarm messages come from the network: a character that could start markdown or HTML ([, <, *, `, & before an entity, ://, @, ...) gets a backslash, so a renderer shows [a](https://...) or <img ...> as text. Addresses, MACs, gids and timestamps read as they are. Omitting response_format or sending null means json. Errors are JSON in either format, and the tools that change state do not take response_format.

get_device_status with {"limit": 2, "response_format": "markdown"}, shortened:

## get_device_status (devices: 2)

- **total_devices:** 2
- **online_devices:** 1
- **devices:** 2 records, under devices below

### devices (2 records)

Same in every record: gid = 00000000-0000-0000-0000-000000000000; ip_reserved = false.

| name | id | last_seen | ip | online |
| --- | --- | --- | --- | --- |
| nas | AA:BB:CC:DD:EE:01 | 2026-09-26T10:00:00.000Z | 192.168.1.10 | true |
| printer | AA:BB:CC:DD:EE:02 | 2026-09-25T18:30:00.000Z | 192.168.1.11 | false |

_This view leaves out the meta block (request_id req_1790000000000_abc123). Call again with `response_format: json` for the full JSON response._

Write tools (opt-in)

Every tool that changes something is off by default, so the server is read-only unless you turn them on. Set FIREWALLA_ENABLE_WRITE_TOOLS=true to register the 11 write tools: create_rule, delete_rule, pause_rule and resume_rule (rules), create_target_list, update_target_list and delete_target_list (target lists), rename_device (devices), and archive_alarm, mute_alarm and delete_alarm (alarms). Without it, calling one answers "Unknown tool" and sends nothing. MCP clients that honor tool annotations can ask before calling them; see Tool annotations.

Up to 1.5.0, pause_rule, resume_rule and the three target-list tools were always registered. If you use them, set FIREWALLA_ENABLE_WRITE_TOOLS=true.

The write tools need an MSP API token with write access. Firewalla said on 2026-09-08 that MSP 2.12 adds read-only API tokens, which cannot make changes. When a write gets HTTP 403, the error says the token may be read-only and that the write tools need a token with write access. A 403 can also mean the request named a box the token cannot access, and the error says that too; get_boxes lists the boxes the token can access.

IDs that go into a request path (id, rule_id, alarm_id, gid, device_id) are checked before anything is sent. One holding /, a backslash, ?, #, %, whitespace or a control character, or that is . or .., is refused as a validation error naming the argument; leading or trailing whitespace is refused too, not trimmed. Rule IDs such as <box gid>:<n>, MAC device IDs and ovpn: device IDs are accepted.

create_rule and rename_device act on one box: gid, else FIREWALLA_BOX_ID or FIREWALLA_DEFAULT_BOX_ID, else the account's only box. On a multi-box account with none of those, they refuse without writing anything, because the MSP API applies a rule with no gid to every box in the account, including boxes added later. delete_rule needs MSP 2.11.0 or later.

archive_alarm and mute_alarm need MSP 2.11.0 or later. archive_alarm takes an alarm out of the active alarms and does nothing else: future matching traffic can still raise new alarms. mute_alarm archives the alarm and has the box create a lasting silence exception. Its target_type says what is silenced: alarmType every future alarm of that alarm's type, whatever the destination; domain a domain and its subdomains; ip one address. Its scope_type says for which devices: all of them, or the one device, group, user or network named by scope_value. The mute is checked against the documented request model before anything is sent, and this server has no tool to remove the exception afterwards. delete_alarm deletes the alarm for good (measured on 2026-09-26: the alarm was gone afterwards); use archive_alarm to keep it among the archived alarms. Alarm IDs are per box, and the same ID can name different alarms on different boxes: all three tools use the alarm's gid, else FIREWALLA_BOX_ID, else check each box, and they refuse when several boxes have that alarm ID and none of them is FIREWALLA_DEFAULT_BOX_ID. All three read the alarm before writing, so a wrong ID fails without changing anything.

Tool annotations

Every tool carries MCP tool annotations: a title, readOnlyHint, and openWorldHint: true, since each one calls the Firewalla MSP API. All get_* and search_* tools are read-only. The tools that change state have readOnlyHint: false, and all of them need FIREWALLA_ENABLE_WRITE_TOOLS=true:

Tool

destructiveHint

idempotentHint

pause_rule, resume_rule

false

true

create_target_list

false

false

update_target_list, delete_target_list

true

true

create_rule

true

false

delete_rule

true

true

rename_device

false

true

archive_alarm

false

true

mute_alarm

true

false

delete_alarm

true

true

pause_rule and resume_rule check the rule's status first and change nothing when it is already paused or active. Clients that honor annotations can ask before calling any tool that is not read-only.

Development

Scripts

npm run dev          # Start development server with hot reload
npm run build        # Build TypeScript to JavaScript
npm run test         # Run all tests
npm run test:watch   # Run tests in watch mode
npm run lint         # Run ESLint
npm run lint:fix     # Fix ESLint issues

MCP Execution Methods

Why npx for MCP servers?

  • Version Management: Always uses the correct/latest version

  • Dependency Resolution: Handles package dependencies automatically

  • No global installation required: Works without global installation

  • MCP Standard: Follows Model Context Protocol conventions

  • Reliable: Works consistently across different environments

Alternative execution methods:

# Development (from source)
npm run mcp:start

# Production (npm installed)
npx firewalla-mcp-server

# Direct execution (from source after build)
node dist/server.js

Launch checks in CI

scripts/launch-smoke.mjs starts the server the way a user would and requires an answer to an MCP initialize over stdio. It uses dummy credentials, and initialize is answered locally, so nothing is sent to Firewalla.

  • The CI workflow's launch job packs the package and starts it through the global bin, npx and node dist/server.js on Linux, macOS and Windows.

  • The Docker Build workflow runs on pull requests that touch the Dockerfile, the package files or the Docker workflows. It builds the image for amd64, arm64 and arm/v7, then runs the amd64 image with docker run -i --rm and requires serverInfo.version to equal package.json's version. Run the workflow by hand with the image input (for example amittell/firewalla-mcp-server:1.4.1) to pull and check a published image instead.

To run the Docker check locally:

docker build -t firewalla-mcp-server:local .
node scripts/launch-smoke.mjs --docker firewalla-mcp-server:local

# A published image; the expected version comes from the tag
node scripts/launch-smoke.mjs --docker amittell/firewalla-mcp-server:1.4.1 --pull

Project Structure

firewalla-mcp-server/
├── src/
│   ├── server.ts           # Main MCP server
│   ├── firewalla/          # Firewalla API client
│   ├── tools/              # MCP tool implementations
│   ├── resources/          # MCP resource implementations
│   └── prompts/            # MCP prompt implementations
├── tests/                  # Test files
├── docs/
│   └── firewalla-api-reference.md  # API documentation
├── CLAUDE.md              # Comprehensive development guide
├── SPEC.md                # Technical specifications
└── README.md              # This file

Documentation

  • README.md (this file) - Setup and basic usage

  • USAGE.md - Simple usage guide with examples

  • TROUBLESHOOTING.md - Common issues and solutions

  • docs/clients/ - Client-specific setup guides

  • CLAUDE.md - Development guide and commands

Security

See SECURITY.md to report a vulnerability. For the HTTP transport, see HTTP transport security.

  • MSP tokens are stored securely in environment variables

  • No credentials are logged or stored in code

  • Requests are paced to API_RATE_LIMIT per 5 minutes (default 100, the MSP API's quota per token). A request that cannot start within 20 s fails at once, saying when capacity returns

  • Input validation prevents injection attacks

  • Device names, domains and alarm messages are set by the devices and sites on your network, so the server treats them as untrusted: characters that do not display are shown as markers such as <U+E0041>, the prompts quote API data only inside a marked data block, and the initialize instructions tell the client to treat results as data. See Untrusted data

  • All API communications use HTTPS

Known Behaviors and Limitations

Category Classification

  • Flow Categories: Many network flows may show as empty category ("") in the Firewalla API response. This is expected behavior - Firewalla categorizes traffic when it recognizes the domain/service (e.g., "av" for audio/video, "social" for social media).

  • Target List Categories: Some target lists may show category as "unknown". This is normal for user-created or certain system lists.

  • Timeline: Category classification happens at the Firewalla device level and may take time to build up meaningful categorization data.

Data Characteristics

  • Response Sizes: The get_recent_flow_activity tool returns up to 150 recent flows to stay within token limits. For larger datasets or historical analysis, use search_flows with time filters for more targeted queries.

  • Geographic Data: IP geolocation is enriched by the MCP server and includes country, city, and risk scores when available.

API Limitations

  • Alarm Deletion: In July 2025 the MSP API answered DELETE /v2/alarms/{gid}/{aid} with {"message": "success", "success": true} and kept the alarm, so delete_alarm was withdrawn. Measured again on 2026-09-26, the same request deleted the alarm (a GET of it then answered 404, and the archived alarm count fell by one), and delete_alarm is back as an opt-in write tool.

Troubleshooting

Quick Fixes

Server won't start:

# Clean and rebuild
npm run clean
npm run build

# If build fails, try:
npm install
npm run build

Authentication errors:

  • Check your MSP token is valid

  • Verify Box ID format (long UUID)

  • Confirm MSP domain is correct

No data returned:

  • Try broader queries: "last week" vs "last hour"

  • Check if Firewalla is online

  • Test with: "show me basic statistics"

Slow responses:

  • Add limits: "top 10 devices"

  • Use shorter time ranges

  • Restart the server

Debug Mode

Enable detailed logging:

DEBUG=mcp:* npm run mcp:start

For more detailed troubleshooting, see TROUBLESHOOTING.md

Contributing

  1. Fork the repository

  2. Create a feature branch

  3. Make your changes

  4. Add tests for new functionality

  5. Run the test suite

  6. Submit a pull request

What's New

Release notes for every version, newest first, are in CHANGELOG.md. Changes that are merged but not yet released are listed there under Unreleased.

License

MIT License

Support

For issues and questions:


GitHub Repository

Repository: https://github.com/amittell/firewalla-mcp-server

Repository Stats

GitHub issues GitHub stars GitHub license

Available Tools

28 tools
create_target_listA

Create a new target list (POST /v2/target-lists); each call creates another list. owner global makes it shareable across all boxes, a box GID ties it to that box.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesTarget list name (required, max 24 chars)
notesNoAdditional description (optional)
ownerYesOwner: "global" or box GID (required)
targetsYesArray of domains, IPs, or CIDR ranges (required)
categoryNoContent category (optional)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnlyHint=false, idempotentHint=false), the description spells out 'each call creates another list,' making non-idempotency concrete for an agent. It also discloses the scope consequence of owner selection: 'global' shares the list across all boxes, while a box GID ties it to that box.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences deliver the core operation, endpoint, non-idempotency, and owner semantics without unnecessary information. The most decision-relevant details are front-loaded and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with a fully documented input schema, the description sufficiently covers the operation, endpoint, non-idempotent behavior, and owner-scope options. Minor omissions such as return format are acceptable because no output schema is declared and the schema already describes parameter constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all five parameters, giving a baseline of 3. The description adds semantic value by explaining the owner field's effect ('shareable across all boxes' vs 'ties it to that box') rather than only restating accepted values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Create a new target list' and identifies the exact REST endpoint (POST /v2/target-lists), naming the action and resource unambiguously. The phrase 'each call creates another list' also distinguishes it from update/delete target-list siblings by emphasizing non-idempotent creation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this tool is for creating a fresh target list, but it does not explicitly state when not to use it or mention update_target_list/delete_target_list as alternatives for modifying or removing lists. Usage context is inferable rather than directly instructed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_target_listA
DestructiveIdempotent

Permanently delete a target list (DELETE /v2/target-lists/{id}); cannot be undone. The tool does not check whether a rule still targets the list.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesTarget list ID to delete (required)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation destructive and non-read-only; the description adds the permanence warning and the important caveat that no rule-reference check is performed. This is meaningful behavioral context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load the action and endpoint, then add the two essential caveats. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive delete with no output schema, the description covers the core action, permanence, and the rule-reference caveat. It does not describe response/error behavior, but that is a minor gap given the tool's simplicity and annotation coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter id is fully described in the schema (100% coverage), so the description adds no additional parameter-level detail. Baseline 3 is appropriate because the schema carries the semantic weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Permanently delete'), the resource ('target list'), the HTTP endpoint, and a key consequence ('cannot be undone'). This clearly distinguishes it from sibling create/update/get target-list tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: deleting a target list by ID, and warns that it is permanent and does not check for rule references. It does not explicitly name alternatives or conditions when to avoid use, so guidance is implied rather than fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_active_alarmsA
Read-only

Retrieve active security alarms from the Firewalla MSP API (GET /v2/alarms): status:1 is added unless the query names a status (status:2 for archived alarms). Without a ts: qualifier the API covers the last 30 days. Returns up to limit alarms and a cursor for the next page, or groups with groupBy. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoResults per page (optional, default: 200, API maximum: 500)
queryNoSearch query for filtering alarms. Active alarms only (status:1) unless the query names a status. Use type:N where N is: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Examples: type:8 (video), type:10 (porn), region:US, device.ip:192.168.*
cursorNoPagination cursor from previous response
sortByNoSort alarms (default: ts:desc)
groupByNoFields to group by, comma-separated, e.g. "type", "status", "device" or "type,box". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true and openWorldHint, but the description adds extensive behavioral detail: the default status:1 filter unless overridden, the default time range, pagination via cursor, grouping behavior and response format, and scope handling. This goes far beyond annotations and clearly discloses operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet information-dense. It front-loads the main purpose and then covers default status, time range, pagination, grouping, and scope in a few well-structured sentences. Every sentence contributes essential information with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a retrieval tool with 5 optional parameters and no output schema, the description covers all critical aspects: default filters, pagination mechanism, grouping output shape, and scope. It does not detail alarm object fields, but that is not necessary for invoking the tool correctly; the agent knows what to expect in the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already documented. However, the description adds meaning by explaining the default status injection, the query syntax with examples (type codes, region, device.ip), and the grouping output structure. This adds value beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the verb 'Retrieve' and the resource 'active security alarms from the Firewalla MSP API' with the endpoint (GET /v2/alarms). It distinguishes itself by detailing default status, time window, pagination, grouping, and scope, making it clear what this tool does and how it differs from a generic search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool: it covers active alarms by default, explains the default 30-day window, and notes scoping to a specific box ID. It does not explicitly name alternative tools or when not to use it, but the context is strong enough for an agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_bandwidth_usageA
Read-only

Top devices by upload plus download over the period, summed client-side from up to 10 times limit (1,000 at most) of the period's most recent flows (GET /v2/flows, 500 per request), so on a busy network the totals cover a sample. Scoped to box, else FIREWALLA_BOX_ID, else every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNoOnly flows of this box (box gid). Defaults to FIREWALLA_BOX_ID; without either, every box.
limitNoNumber of top devices to return
periodYesTime period for bandwidth calculation

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark read-only and open-world, so the description carries the behavioral burden. It adds a substantial behavioral profile: client-side summation over at most 10x limit flows, cap of 1,000, sample-based incompleteness, and box scoping precedence. This is exactly the information an agent needs to interpret results appropriately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loads the outcome before the algorithm details. It is dense but every clause contributes; only the fairly long single sentence and the embedded API endpoint reference add slight complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only aggregate tool with no output schema, the description covers the key operational details: period, scope, sample limitation, and result ordering. It could still mention return shape or what happens with no flows, but the definition is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all three params, so baseline is 3. Description adds meaning by relating `limit` to the client-side sampling cap (up to 10x limit, max 1,000) and by giving the `box` resolution order (box → FIREWALLA_BOX_ID → every box), which the schema doesn't provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific result: top devices by combined upload+download over a period. It also differentiates itself from sibling tools by describing a client-side summation algorithm and scope resolution, though it never names a specific sibling tool or explicitly contrasts with alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: use this to get top bandwidth consumers over a period, scoped to a box or all boxes, and warns that on busy networks totals are sampled. However, it doesn't state when not to use it or explicitly mention alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_boxesA
Read-only

List the Firewalla boxes this MSP token can see (GET /v2/boxes, optionally one box group), with online status, model and version. Not limited by FIREWALLA_BOX_ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNoGet boxes within a specific group (requires group ID)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds useful behavioral context: the endpoint, the optional group filter, the returned fields, and the scope being all boxes visible to the MSP token rather than a single box.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, information-dense sentence front-loads the action and resource, then adds the endpoint, optional filter, output fields, and scope limitation. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with zero required parameters and one optional parameter, the description covers what the tool returns, its scope, and the relevant endpoint. No output schema exists, but the returned fields are stated directly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already documents the optional group parameter and the requirement for a group ID. The description only restates that a group is optional, adding minimal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a clear resource ('Firewalla boxes this MSP token can see'), names the exact endpoint, and states what information is returned. It is distinct from siblings like get_device_status or search_devices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the access scope explicit ('MSP token can see') and clarifies that it is not limited by FIREWALLA_BOX_ID, which helps agents decide when this is the right tool. It does not explicitly name alternatives or exclusions, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_device_statusA
Read-only

Check online/offline status of devices on the Firewalla network. Reads the device list from GET /v2/devices (box, else FIREWALLA_BOX_ID, else every box; group limits it to a box group) and returns up to limit devices, sorted by name.

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNoGet devices under a specific Firewalla box (requires box ID)
groupNoGet devices under a specific box group (requires group ID)
limitYesMaximum number of devices to return (required)

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description reveals useful behavior: it reads from GET /v2/devices, explains the box/group selection precedence, and states the result is limited and sorted by name. This gives the agent meaningful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary purpose, and packs the key behavioral details into a compact parenthetical. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool, the description is quite complete: it covers purpose, data source, scoping rules, limit behavior, and sort order. With no output schema, it could say slightly more about the exact shape of the returned device objects, but the core behavior is sufficiently clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning by explaining parameter precedence: box takes priority, then FIREWALLA_BOX_ID, then every box, while group further restricts. It also clarifies that limit caps the returned device count.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: checking online/offline status of Firewalla network devices. It does not explicitly differentiate from related siblings like get_offline_devices or search_devices, so it misses the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Check online/offline status of devices' implies a clear use case, and the selection logic for box/group provides scoping context. However, it does not state when to prefer this tool over siblings or mention any exclusions, so guidance is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_flow_dataA
Read-only

Query network traffic flows from the Firewalla MSP API (GET /v2/flows). Without a ts: qualifier the API covers the last 24 hours. Returns up to limit flows and a cursor for the next page, or groups with groupBy. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum results (optional, default: 200, API maximum: 500)
queryNoSearch query for flows. Supports region:US for geographic filtering, protocol:tcp, status:blocked, domain:*, category:social, etc.
cursorNoPagination cursor from previous response
sortByNoSort flows (default: "ts:desc")
groupByNoFields to group by, comma-separated, e.g. "category", "domain", "device", "box" or "device,category". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for "device,category") and the rest are the group's summed connection count and bytes.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation read-only and open-world; the description adds materially: default 24h lookback unless a ts: qualifier is used, paginated cursor behavior, grouping semantics, and environment scoping. No contradictions with annotations. It doesn't cover error/rate-limit behavior but that is not required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences with no filler; the endpoint and core behavior are front-loaded, and each sentence adds a distinct operational fact. Could be slightly shorter, but all content earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 optional params, no output schema, and read-only annotations, the description covers the main invocation concerns: time window, pagination, grouping, scoping. It still omits an explicit differentiation from search_flows and does not describe the shape of a returned flow record for non-grouped results, which leaves minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% description coverage, so the baseline is 3. The description goes beyond the schema by explaining the ts: qualifier's effect on the time window and by detailing groupBy's shape (key, count, download, upload, total) and peculiar key behavior for device vs device,category. That compensation justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb ('Query'), resource ('network traffic flows'), and the underlying API endpoint (GET /v2/flows), making the operation unambiguous. It does not explicitly distinguish itself from sibling search_flows, which also appears to target flow data, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to prefer this tool over search_flows, get_recent_flow_activity, or get_flow_insights. It states behavioral conditions (24h window, FIREWALLA_BOX_ID scoping) but not selection criteria or exclusions. This is the main clarity gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_flow_insightsA
Read-only

Get category-based flow analysis for a period: top content categories and their domains, top devices by bandwidth, and optionally blocked traffic. Ideal for answering questions like "what porn sites were accessed" or "what social media was used". Computed client-side from the period's largest flows (GET /v2/flows by total bytes: up to 500 for categories, 200 for devices) and, with include_blocked, the 50 most frequent blocked flows, so on a busy network it covers the largest flows, not all of them. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNoTime period for analysis (default: 24h)24h
categoriesNoFilter to specific content categories (optional)
include_blockedNoInclude blocked traffic analysis (default: false)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint, and the description adds substantial behavioral detail: it is computed client-side from the largest flows, with explicit caps (500 categories, 200 devices, 50 blocked flows), and it explicitly warns that on a busy network it covers the largest flows, not all of them. It also discloses FIREWALLA_BOX_ID scoping. This goes well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler. The main purpose is front-loaded, followed by usage examples and then necessary computation/scoping caveats. Every sentence contributes to correct tool selection and invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining what the tool returns, and it does: categories, domains, devices, and optional blocked traffic. It also covers limits, scoping, and the optional parameter behavior, making it complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explaining that include_blocked yields the 50 most frequent blocked flows and that categories map to top content categories and domains. It does not add much about the period parameter, but the schema already covers it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Get category-based flow analysis for a period,' and enumerates concrete outputs (top content categories and their domains, top devices by bandwidth, optionally blocked traffic). It also gives example questions ('what porn sites were accessed') that make the tool's niche unmistakable among flow-related siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context by saying it is 'Ideal for answering questions like...' and by explaining the client-side computation scope. It does not explicitly name alternative tools or state when not to use it, but the context is strong enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_network_rulesA
Read-only

Retrieve firewall rules and conditions (GET /v2/rules). The API returns every matching rule; the tool returns the first limit. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitYesMaximum number of rules to return (required)
queryNoSearch conditions for filtering rules

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description discloses two important behaviors: the tool returns only the first 'limit' results even though the API returns all matching rules, and it is scoped to FIREWALLA_BOX_ID when set. This adds meaningful context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three compact sentences with no wasted words. It front-loads the primary action, then explains the truncation behavior and scoping, making every sentence valuable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only tool with 100% schema coverage and clear annotations, the description is nearly complete. It covers the endpoint, result truncation, and scoping, though it could be stronger by distinguishing itself from search_rules.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies. The description clarifies the limit parameter's effect by noting the tool returns only the first 'limit' results, but it does not add substantive meaning beyond the schema's descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Retrieve') and identifies the resource ('firewall rules and conditions') with an endpoint reference (GET /v2/rules). It is clear, but it does not explicitly differentiate from the sibling search_rules tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus search_rules or get_network_rules_summary. The description provides behavioral context such as limiting results and scoping by FIREWALLA_BOX_ID, but it does not state exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_network_rules_summaryA
Read-only

Get overview counts of network rules by action, direction, status and target type (convenience wrapper): reads the rules from GET /v2/rules and counts them locally. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_typeNoFilter by rule type
active_onlyNoOnly include active rules in summary (default: true)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, and the description is consistent with them. It adds valuable behavior beyond annotations: the tool reads from GET /v2/rules and counts locally, and it reveals global scope when FIREWALLA_BOX_ID is unset. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, each earning its place: the first communicates purpose and mechanism, the second covers scoping behavior. There is no filler, repetition of schema details, or irrelevant context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only summary tool with zero required parameters, the description adequately states what is counted, the data source, and the box-scoping behavior. Since there is no output schema, describing the exact return shape in more detail would make it fully complete, but 'counts by action, direction, status and target type' conveys the essential result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both rule_type and active_only are already documented in the input schema. The description does not add extra meaning about accepted values, filtering behavior, or edge cases, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description leads with a specific verb-resource combination: 'Get overview counts of network rules by action, direction, status and target type.' It further clarifies the aggregate nature by calling itself a 'convenience wrapper' that reads GET /v2/rules and counts locally, making it clearly distinguishable from sibling tools like get_network_rules or search_rules.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'convenience wrapper' label and 'overview counts' wording imply this is for summary-level insights rather than raw rule listings, but the description never explicitly names alternatives or states when not to use it. The only explicit guidance is scoping behavior (FIREWALLA_BOX_ID vs. every box), which is execution context rather than tool-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_offline_devicesA
Read-only

List offline devices from the full device list (GET /v2/devices), most recently seen first by default, up to limit; total_offline_devices counts all of them. Scoped to box, else FIREWALLA_BOX_ID, else every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNoFilter devices under a specific Firewalla box
limitNoMaximum number of offline devices to return
sort_by_last_seenNoSort devices by last seen time (default: true)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true; the description adds scoping precedence, default sorting, and total_offline_devices count behavior. These details go beyond annotations and help an agent anticipate response semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with zero filler; the action is front-loaded and every clause conveys scoping, ordering, limit, and count behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with 3 optional params and no output schema, it covers return semantics (total count, sort, limit) and scoping. Does not define 'offline' or pagination, but those are minor given the annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds box scoping precedence ('Scoped to box, else FIREWALLA_BOX_ID, else every box') beyond the schema's 'filter' wording, and confirms limit/sort defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List offline devices from the full device list'), names the exact endpoint (GET /v2/devices), and clarifies the filtering semantics. The scoping rules distinguish it from generic device listing tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implied usage context is present through sort order, limit, and scoping fallback, but there is no explicit when-to-use guidance or comparison with the sibling search_devices. No exclusions or alternatives are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_recent_flow_activityA
Read-only

Get a snapshot of the 50 most recent network flows (one GET /v2/flows request) with protocol, region and blocked/allowed counts; the minutes they span depend on how busy the network is. Use this for: "what's happening right now?", current security threats, immediate network issues. DO NOT use for: historical analysis, more than 50 flows, or daily/weekly patterns; use search_flows with time queries like "ts:>24h" for those. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, and the description adds meaningful behavioral context: exactly one API request, variable time span based on network busyness, and FIREWALLA_BOX_ID scoping with fallback to every box. There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three well-organized sentences: definition, use/don't-use guidance, and scope. Every sentence earns its place, and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameters, the description adequately covers return content, limits, request behavior, and scope. An agent has all the information needed to select and invoke this tool accurately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so there is no parameter burden on the description. It still adds relevant invocation context by explaining the implicit FIREWALLA_BOX_ID environment scoping, which is directly useful for calling the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get a snapshot of the 50 most recent network flows (one GET /v2/flows request)'. It names the output dimensions (protocol, region, blocked/allowed counts), and its contrast with search_flows distinguishes it from the closest sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool ('what's happening right now?', current security threats, immediate network issues) and when not to use it (historical analysis, >50 flows, daily/weekly patterns), pointing to search_flows with a concrete query example 'ts:>24h'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_simple_statisticsA
Read-only

Get account-wide counts from GET /v2/stats/simple: online boxes, offline boxes, alarms and rules, optionally for one box group. Not limited by FIREWALLA_BOX_ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNoGet statistics for specific box group

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=true. The description reinforces and extends these by explaining the account-wide scope, the optional group parameter, and the fact that it is not constrained by FIREWALLA_BOX_ID. It also implies the return shape by listing the counted entity types.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler. The endpoint and data scope are front-loaded, the optional parameter is mentioned, and the scope caveat is placed last for emphasis.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only statistics tool with one optional parameter and no output schema, the description is complete: it explains what is counted, where the data comes from, and how scope is determined. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single optional 'group' parameter at 100%, so the baseline is 3. The description adds only 'optionally for one box group', which slightly clarifies usage but does not significantly extend beyond the schema's own description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Get account-wide counts from GET /v2/stats/simple' and enumerates the exact data categories (online boxes, offline boxes, alarms and rules). It also distinguishes this from box-scoped siblings by adding 'Not limited by FIREWALLA_BOX_ID'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly communicates when the tool is appropriate: account-wide counts rather than per-box statistics, with an optional box-group filter. It does not explicitly name alternative tools like get_statistics_by_box or get_statistics_by_region, so it stops short of full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_specific_alarmA
Read-only

Get detailed information for one Firewalla alarm (GET /v2/alarms/{gid}/{aid}). Alarm IDs are per box: pass gid on a multi-box account, or each box is checked, one request per box, until one has the alarm.

ParametersJSON Schema
NameRequiredDescriptionDefault
gidNoBox the alarm belongs to (the gid field of get_active_alarms or search_alarms results). Defaults to FIREWALLA_BOX_ID; without either, each box on the account is checked.
alarm_idYesAlarm ID (required for API call): the aid from get_active_alarms or search_alarms, as a number or a string

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses the multi-box scan behavior: without gid, 'each box is checked, one request per box, until one has the alarm.' This is a meaningful behavioral trait involving multiple HTTP requests and stop-at-first-match logic that annotations do not convey. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and endpoint, followed by the essential per-box ID caveat. No filler or redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only getter with full schema coverage and readOnlyHint, the description covers the endpoint and the only non-obvious behavior (multi-box lookup). It does not describe the return shape or not-found behavior, but there is no output schema and the operation is simple enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces the per-box semantics of alarm IDs, but the input schema already documents the gid fallback and the aid source, so it adds little beyond what structured data already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get detailed information for one Firewalla alarm' plus the REST endpoint (GET /v2/alarms/{gid}/{aid}). This clearly differentiates it from sibling list tools like get_active_alarms and search_alarms by targeting a single alarm.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when the tool is relevant and explains the gid/fallback behavior ('pass gid on a multi-box account, or each box is checked...'). It does not explicitly name alternatives or exclusions, but the schema notes that alarm_id comes from get_active_alarms or search_alarms, which is enough to route an agent from a list result to this details call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_specific_target_listA
Read-only

Retrieve one target list by ID, including its targets (GET /v2/target-lists/{id}).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesTarget list ID (required)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds that the tool returns the list 'including its targets', which is useful but not a behavioral disclosure beyond what the annotations imply. It does not mention any side effects, error conditions, or pagination, but for a read-only retrieval this is acceptable. It adds some value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is concise and front-loaded with the verb and resource. It includes the API endpoint for reference, which is helpful without being verbose. Every word earns its place, and there is no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple retrieval tool with one parameter, no output schema, and read-only annotations, the description is complete. It states what it does, what it returns, and the parameter is fully documented. There are no missing pieces an agent would need to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents the only parameter (id) with 100% coverage, so the baseline is 3. The description does not add any extra semantics about the parameter format (e.g., UUID vs. string) or constraints beyond what the schema states. It simply reiterates the purpose. No additional value is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Retrieve'), a precise resource ('one target list by ID'), and clarifies the returned content ('including its targets'). It clearly differentiates from sibling tools like get_target_lists (plural) and search_target_lists by implying single-ID retrieval. This is unambiguous and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it: when you have a specific target list ID and want its details. It does not explicitly mention alternatives or when not to use it, but the context is clear enough that an agent can infer it is for single-list retrieval rather than listing or searching. No explicit exclusions are given, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statistics_by_boxA
Read-only

Top boxes by blocked flows (the default) or by Security Activity alarms, from GET /v2/stats/{type}, with each box's details from GET /v2/boxes; each box's value is the statistic, over about the last 30 days when measured. Not limited by FIREWALLA_BOX_ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoStatistics type to retrievetopBoxesByBlockedFlows
groupNoGet statistics for specific box group
limitNoMaximum number of results (optional, default: 5)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, lowering the burden on the description. The description usefully adds the approximate 30-day measurement window, the fact that results are not limited by FIREWALLA_BOX_ID, and the semantics that each box's value is the statistic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence where every clause earns its place: default vs. alternative statistic, source endpoints, value semantics, time window, and scope limitation. No filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with three optional parameters and no output schema, the description gives enough to invoke it correctly and interpret the result concept. It could be slightly more explicit about the exact response shape, but the endpoint references and value semantics largely compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents type, group, and limit. The description reinforces the default type and the two allowed enum values, but does not add meaningful parameter-level details beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: top boxes ranked by blocked flows or Security Activity alarms, sourced from a specific endpoint and enriched with box details. It also distinguishes this from other statistics tools by emphasizing per-box scope and the 'Not limited by FIREWALLA_BOX_ID' qualifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clarifies the default statistic type and that the tool is not limited by FIREWALLA_BOX_ID, giving some usage context. However, it does not explicitly say when to prefer this tool over siblings like get_statistics_by_region or get_simple_statistics, nor does it state any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statistics_by_regionA
Read-only

Top regions by blocked flows, from GET /v2/stats/topRegionsByBlockedFlows, optionally for one box group; the API returned no more than 5 regions. Not limited by FIREWALLA_BOX_ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNoGet statistics for specific box group
limitNoMaximum number of regions (optional, default: 5; the API returned no more than 5 when a larger limit was tried)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already signaling a safe read operation, the description adds useful behavioral detail: the API caps results at 5 regions even when a larger limit is requested, and the query is not constrained by FIREWALLA_BOX_ID. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence front-loads the result semantics and packs in endpoint, optional filter, API cap, and scope limitation without wasted words. Every clause adds useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only statistics tool with two optional parameters and no output schema, the description plus schema fully covers what an agent needs: what is returned, the optional group filter, the limit behavior, and the scope. The readOnly and openWorld annotations complement the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and both parameters (group, limit) are already described with meaningful detail including default and minimum. The description's mentions of optional box-group filtering and the 5-region cap reinforce the schema but do not add new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly identifies the result as top regions ranked by blocked flows, cites the exact endpoint, and disambiguates scope with optional box group and the absence of a box-ID restriction. The noun-phrase style is still unambiguous because the endpoint and resource are explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this tool for regional blocked-flow statistics, optionally filtered by one box group, and it explicitly notes the query is not limited by FIREWALLA_BOX_ID. It does not name sibling alternatives like get_statistics_by_box, but the scope guidance is strong enough to route an agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_target_listsA
Read-only

Retrieve target lists (GET /v2/target-lists). Without owner the API returns the MSP's global lists and the Firewalla-managed lists; owner selects global lists, a box's lists, or several. entry_count is the number of entries in each list; the API does not return the entries of Firewalla-managed lists, so their targets is null. Returns up to limit lists.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitYesMaximum number of target lists to return (required)
ownerNoOnly lists with this owner: "global" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. "global,<box_gid>". Default: global and Firewalla-managed lists.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint and openWorldHint. The description adds valuable context about the response: entry_count is provided, and Firewalla-managed lists have null targets because entries are not returned. It also notes the limit behavior. This complements the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the endpoint and core behavior. It avoids unnecessary filler and includes important details about response fields and limits. It is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list-retrieval tool with no output schema, the description covers the key aspects: owner selection, default scope, entry_count, null targets for Firewalla-managed lists, and the limit parameter. Nothing essential is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are well-documented in the schema itself. The description adds minimal extra value: it repeats the default owner behavior and mentions the limit behavior. Since the schema already explains owner options and default, the description does not significantly enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves target lists, specifies the endpoint, and explains the behavior with and without the owner parameter. It distinguishes the scope (global, box, or multiple owners) and the default behavior, making it unambiguous what this tool does relative to siblings like get_specific_target_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use this tool: to fetch target lists with optional owner filtering, and describes default behavior. It does not explicitly mention alternatives or when not to use it, but the context is clear enough for an agent to select it for list retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pause_ruleA
Idempotent

Pause an active firewall rule on the box until resume_rule reactivates it (POST /v2/rules/{id}/pause, no body). The MSP API takes no duration, so the pause does not expire on its own. Checks the rule's status first and changes nothing if it is already paused.

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_idYesRule ID to pause

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already include idempotentHint=true and destructiveHint=false, but the description adds significant context: the API takes no duration and the pause does not expire on its own, and it checks the rule's status first, changing nothing if already paused. This explains the idempotent behavior in detail and provides endpoint information, exceeding what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all substantive: action, endpoint, no-duration behavior, and idempotency check. No filler or redundancy, and the core action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation tool with annotations covering safety and idempotency, the description includes the endpoint, no-duration constraint, and status check. It provides everything an agent needs to call it correctly, including when it is safe to invoke.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the description does not add any meaning to the rule_id parameter beyond the schema's 'Rule ID to pause'. No additional constraints, format hints, or examples are provided, so it relies on the schema as expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Pause') and resource ('active firewall rule'), and explicitly references 'resume_rule' as the counterpart, distinguishing it from sibling tools like search_rules. It is unambiguous about the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions that the pause lasts until resume_rule reactivates it, which implies when to use the reverse tool, and notes that it checks status and does nothing if already paused, indicating safe calling conditions. However, it does not explicitly list alternatives or exclusions beyond resume_rule, so it's clear but not fully exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_ruleA
Idempotent

Resume a paused firewall rule on the box, restoring it to active (POST /v2/rules/{id}/resume, no body). Checks the rule's status first and changes nothing if it is already active.

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_idYesRule ID to resume

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the HTTP method and body requirement, and explains the status-checking behavior that makes the operation effectively idempotent. This adds meaningful behavioral context that the annotations alone do not provide, such as 'changes nothing if it is already active.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences deliver the action, resource, endpoint, request body requirement, and idempotency behavior with no filler. The key purpose is front-loaded, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with no output schema, the description is complete: it states what happens, how it happens, and the edge case of an already-active rule. Annotations cover the mutation and idempotency profile, so nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter rule_id is already described as 'Rule ID to resume.' The description adds the path template showing where the ID goes, but does not add new semantic detail about the parameter's format, constraints, or behavior. Baseline 3 is appropriate given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Resume'), a specific resource ('paused firewall rule'), and the intended outcome ('restoring it to active'). It also names the exact endpoint (POST /v2/rules/{id}/resume), which unambiguously distinguishes it from sibling tools like pause_rule.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool: when a paused firewall rule should be restored to active. It also provides useful guidance that the tool checks status first and is a no-op if already active, so callers need not pre-check state. It does not explicitly name alternatives, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_alarmsB
Read-only

Search alarms using full-text or field filters. Alarm types: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Reads GET /v2/alarms, 500 per request, following the cursor up to limit. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum results (optional, default: 200, API maximum: 500)
queryYesSearch query using Firewalla syntax. Supported fields: type:1-16 (see alarm types above), status:1/2 (active/archived), device.ip:192.168.*, region:US (country code), box.id:box_gid, device.name:*. Examples: "type:8 AND region:US" (video from US), "type:10 AND status:1" (active porn alerts), "device.ip:192.168.* AND status:1" (active alarms from the LAN), "porn" (free text: a term without a qualifier searches alarm text)
cursorNoPagination cursor from previous response
sortByNoSort alarms (default: ts:desc)
groupByNoFields to group by, comma-separated, e.g. "type", "status", "device" or "type,box". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group.

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint. The description adds valuable behavior: it mentions the HTTP endpoint (GET /v2/alarms), pagination limit of 500 per request, cursor following up to the limit, and scoping to a box ID. This goes beyond the annotation flags and helps the agent anticipate API behavior. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that lists all 16 alarm types, making it somewhat long but logically structured. It front-loads the core purpose, then provides the type mapping and behavioral details. While it could be more concise (e.g., moving the type list to a reference section), it is organized and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema, the description covers input semantics, pagination, and scoping but does not describe the response format (e.g., what fields are returned, how groups look). This is a gap because agents need to know how to parse results. Given the complexity of 5 parameters, this omission prevents full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all parameters have descriptions. The description adds the alarm type mapping (1-16) referenced by the query parameter, which is essential for constructing valid queries. It also explains pagination behavior with cursor and limit, adding meaning beyond the schema's generic descriptions. This is a meaningful supplement to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Search alarms using full-text or field filters.' It specifies the resource (alarms) and the action (search), and even lists alarm types, which helps understanding. However, it does not explicitly differentiate from sibling tools like get_active_alarms or get_specific_alarm, so it falls short of the 5-level criterion that requires distinguishing from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention conditions for choosing this over get_active_alarms or get_specific_alarm, nor does it state any exclusions or prerequisites. The only context given is scoping to FIREWALLA_BOX_ID, which is a parameter detail, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_devicesA
Read-only

Search devices by name, IP, MAC or status (convenience wrapper with client-side filtering): reads the device list from GET /v2/devices (box, else FIREWALLA_BOX_ID, else every box) and filters it locally.

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNoFilter devices under a specific Firewalla box
limitNoMaximum number of devices to return
queryYesSearch query using Firewalla syntax. Supported fields: mac:AA:BB:CC:DD:EE:FF, ip:192.168.1.*, name:*iPhone*, online:true/false, mac_vendor:Apple, gid:box_gid, network.name:*, group.name:*. Examples: "online:false AND mac_vendor:Apple", "ip:192.168.1.* AND name:*laptop*", "mac:AA:* OR name:*phone*"

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it read-only and open-world. The description adds valuable behavioral context beyond annotations: it fetches the complete device list and filters client-side, and describes the box/FIREWALLA_BOX_ID/every-box precedence. This helps the agent understand the tool's scope and performance implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the purpose and then packs in the essential implementation detail. Every phrase earns its place, and the key behavior (local filtering) is highlighted early.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with no output schema, the description covers the data source, filtering approach, and scope selection logic. It doesn't describe the return shape, but that is partially inferred from the query schema and the tool's search nature. The main minor gap is the lack of an explicit return-value description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining the box parameter's fallback behavior (box, else FIREWALLA_BOX_ID, else every box), which is not present in the schema. Limit and query are well-documented in the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search devices by name, IP, MAC or status'. It clearly identifies the tool as a convenience wrapper and differentiates from sibling search tools by naming the device-list endpoint it wraps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: it reads the full device list from GET /v2/devices and filters locally, with explicit fallback logic for box selection. It doesn't explicitly name alternatives, but no sibling tool searches devices, so the usage context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_flowsA
Read-only

Search network flows with advanced query filters. Use this for: historical analysis, specific time ranges, complex filtering, or when you need more than 50 flows. Supports pagination, time-based queries (e.g., "ts:>1h" for the last hour, or Unix seconds such as "ts:1735689600-1735693200"), and all flow fields including geographic filtering. For quick "what's happening now" snapshots, use get_recent_flow_activity instead. Reads GET /v2/flows, 500 per request, following the cursor up to limit. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum results (optional, default: 200, API maximum: 500)
queryYesSearch query using Firewalla syntax. Supported fields: protocol:tcp/udp, direction:inbound/outbound/local, status:blocked/ok, total:>1MB (download + upload in B/KB/MB/GB/TB), download:>10MB, upload:>10MB, domain:*.example.com, region:US (country code), category:social/games/porn/etc, box.id:box_gid, device.ip:192.168.*, source.ip:*, destination.ip:*, ts:>1h. Examples: "region:US AND protocol:tcp", "status:blocked AND region:CN", "category:social OR category:games"
cursorNoPagination cursor from previous response
sortByNoSort flows (default: "ts:desc")
groupByNoFields to group by, comma-separated, e.g. "category", "domain", "device", "box" or "device,category". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for "device,category") and the rest are the group's summed connection count and bytes.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint, but the description adds substantive behavioral details: it reads GET /v2/flows, caps at 500 per request, follows the cursor up to limit, and scopes results to FIREWALLA_BOX_ID when set. These clarify pagination, rate limits, and scoping without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: purpose first, then usage conditions, then capability highlights, then a pointer to an alternative, then technical specifics. Every sentence serves a clear function with no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex search tool with no output schema, the description covers everything needed to invoke it correctly: purpose, when to use, pagination behavior, scoping, and even the endpoint. Grouping behavior is conveyed via the schema's groupBy description, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds marginal value by illustrating time-based query syntax ('ts:>1h', Unix seconds) and explaining how cursor and limit interact ('following the cursor up to limit'). These details are not fully captured in the schema's parameter descriptions, so a slight premium is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Search network flows with advanced query filters', giving a specific verb+resource. It then enumerates concrete uses (historical analysis, specific time ranges, complex filtering, >50 flows) that sharply differentiate it from the sibling get_recent_flow_activity, making the tool's role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('Use this for: historical analysis, specific time ranges, complex filtering, or when you need more than 50 flows') and when not to, naming the alternative ('For quick "what's happening now" snapshots, use get_recent_flow_activity instead'). This direct routing leaves no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_rulesA
Read-only

Search firewall rules by target, action or status; the MSP API applies the query (GET /v2/rules). Supports all rule fields. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of rules to return
queryYesSearch query using Firewalla syntax. Supported fields: action:allow/block/timelimit, target.type:domain/ip/device, target.value:*.facebook.com, status:active/paused, direction:bidirection/inbound/outbound, protocol:tcp/udp, box.id:box_gid, scope.type:device/network, notes:"description text". Examples: "action:block AND target.value:*.social.com", "status:paused", "target.type:domain AND action:block"

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is known. The description adds behavioral value beyond that: it identifies the API endpoint (GET /v2/rules), notes that all rule fields are supported, and explains the FIREWALLA_BOX_ID scoping behavior that affects which rules are returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core purpose appears in the first clause, followed by only necessary scope and capability details. There is no redundant filler, and each sentence contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with a rich schema, the description adequately covers the query scope and the box-scoping behavior. It does not describe the response payload, and there is no output schema to fill that gap, but the return of firewall rules is strongly implied; a brief note on result format would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% parameter coverage, including a detailed query syntax and examples for both query and limit. The description's reference to 'target, action or status' is only a light restatement of the schema, not additional semantic meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Search firewall rules by target, action or status', which names a specific verb, resource, and query dimensions. Adding 'Supports all rule fields' clarifies the scope and distinguishes it from simple list/get siblings like get_network_rules.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful context, especially 'Scoped to FIREWALLA_BOX_ID when set, otherwise every box', and implies query-based use. However, it does not explicitly state when to prefer search_rules over alternatives such as get_network_rules or search_flows, nor does it give when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_target_listsA
Read-only

Search target lists (convenience wrapper with client-side filtering): reads GET /v2/target-lists, sending owner if given (without it, the global and Firewalla-managed lists), and applies the query locally.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of target lists to return
ownerNoOnly lists with this owner, sent to the API: "global" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. "global,<box_gid>". Default: global and Firewalla-managed lists.
queryYesSearch query for target lists. Supported fields: name:*Social*, owner:global/box_gid, category:social/games/ad/porn/etc, targets:*.facebook.com, notes:"description text", target_count:>100 (entries: n, >n, >=n, <n, <=n or a range 10-50), last_updated:>2026-09-01 (a date or Unix seconds, with the same comparisons). Examples: "category:social", "owner:global AND name:*Block*", "targets:*.gaming.com", "target_count:>1000", "last_updated:<2026-01-01"

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds meaningful behavior: the query is not sent to the server; filtering happens locally after reading the endpoint, and owner defaults to global and Firewalla-managed lists when omitted. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that is front-loaded with the tool's purpose and contains no filler. Every clause contributes behavioral or scoping information, making it easy to extract the key facts quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with full schema coverage, the description covers the endpoint, owner default, and local-filtering behavior. It does not describe the return shape, but no output schema exists and the resource name makes the return type predictable; the absence of explicit sibling routing is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds one useful semantic beyond the schema: the query is applied client-side while owner is passed through to the API. It does not add detail about limit or query syntax, but the schema already documents those thoroughly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly identifies the action ('Search') and resource ('target lists') and specifies that it is a convenience wrapper that reads GET /v2/target-lists and applies the query locally. This differentiates it from a direct list call or single-list retrieval without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: this is a wrapper over GET /v2/target-lists for client-side search/filtering, and the owner parameter controls scope. It does not explicitly name sibling alternatives such as get_target_lists or get_specific_target_list, nor state when not to use it, so it stops short of full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_target_listA
DestructiveIdempotent

Update an existing target list (PATCH /v2/target-lists/{id}). Only the fields given are sent; targets, when given, is the complete new list and is not merged with the current targets.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesTarget list ID (required)
nameNoUpdated target list name (max 24 chars)
notesNoUpdated description
targetsNoUpdated array of domains, IPs, or CIDR ranges
categoryNoUpdated content category

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as destructive and idempotent, and the description adds precise behavioral context: only supplied fields are sent, and targets is a complete replacement, not a merge. This meaningfully clarifies what gets overwritten and prevents an incorrect mental model of additive target updates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The endpoint and operation are front-loaded, and the second sentence delivers the most important behavioral nuance without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with destructive potential and no output schema, the description covers the critical invocation semantics: PATCH behavior, partial field sending, and full target-list replacement. It does not mention the response body or failure cases, but these are not essential for correctly calling the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying the targets parameter's full-replacement semantics and the general partial-update behavior, which are not visible from parameter descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Update an existing target list (PATCH /v2/target-lists/{id})'. It conveys this is a modification operation on a single existing list, but it does not explicitly contrast it with create_target_list, delete_target_list, or get_specific_target_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong usage guidance by explaining PATCH semantics: 'Only the fields given are sent' and that targets, when provided, replaces the full list rather than merging. It does not explicitly state when to prefer this over create/delete/search alternatives, but it provides clear context for correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 15 tool updatesv1.5.0
    • Changedget_active_alarms2 fields changed
      • changedInput schema / properties / groupBy / description
        Previous value: -"Group alarms by field (e.g., type, box)"New value: +"Fields to group by, comma-separated, e.g. \"type\", \"status\", \"device\" or \"type,box\". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group."
      • changedInput schema / properties / query / description
        Previous value: -"Search query for filtering alarms (default: status:1 for active). Use type:N where N is: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Examples: type:8 (video), type:10 (porn), region:US, source_ip:*"New value: +"Search query for filtering alarms. Active alarms only (status:1) unless the query names a status. Use type:N where N is: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Examples: type:8 (video), type:10 (porn), region:US, device.ip:192.168.*"
    • Changedget_alarm_trends1 field changed
      • addedInput schema / properties / period
        Added value: +{
        +  "default": "30d",
        +  "description": "Return the days that overlap this period (default: 30d). The API has no finer resolution than a day, so 1h returns today so far and 24h returns yesterday and today",
        +  "enum": [
        +    "1h",
        +    "24h",
        +    "7d",
        +    "30d"
        +  ],
        +  "type": "string"
        +}
    • Changedget_bandwidth_usage1 field changed
      • changedInput schema / properties / box / description
        Previous value: -"Filter devices under a specific Firewalla box"New value: +"Only flows of this box (box gid). Defaults to FIREWALLA_BOX_ID; without either, every box."
    • Changedget_flow_data2 fields changed
      • changedInput schema / properties / groupBy / description
        Previous value: -"Group flows by specified values (e.g., \"domain,box\")"New value: +"Fields to group by, comma-separated, e.g. \"category\", \"domain\", \"device\", \"box\" or \"device,category\". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for \"device,category\") and the rest are the group's summed connection count and bytes."
      • changedInput schema / properties / query / description
        Previous value: -"Search query for flows. Supports region:US for geographic filtering, protocol:tcp, blocked:true, domain:*, category:social, etc."New value: +"Search query for flows. Supports region:US for geographic filtering, protocol:tcp, status:blocked, domain:*, category:social, etc."
    • Changedget_rule_trends1 field changed
      • addedInput schema / properties / period
        Added value: +{
        +  "default": "30d",
        +  "description": "Return the days that overlap this period (default: 30d). The API has no finer resolution than a day",
        +  "enum": [
        +    "1h",
        +    "24h",
        +    "7d",
        +    "30d"
        +  ],
        +  "type": "string"
        +}
    • Changedget_specific_alarm3 fields changed
      • changedInput schema / properties / alarm_id / description
        Previous value: -"Alarm ID (required for API call)"New value: +"Alarm ID (required for API call): the aid from get_active_alarms or search_alarms, as a number or a string"
      • changedInput schema / properties / alarm_id / type
        Previous value: -"string"New value: +[
        +  "string",
        +  "number"
        +]
      • addedInput schema / properties / gid
        Added value: +{
        +  "description": "Box the alarm belongs to (the gid field of get_active_alarms or search_alarms results). Defaults to FIREWALLA_BOX_ID; without either, each box on the account is checked.",
        +  "type": "string"
        +}
    • Changedget_statistics_by_region1 field changed
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum number of results (optional, default: 5)"New value: +"Maximum number of regions (optional, default: 5; the API returned no more than 5 when a larger limit was tried)"
    • Changedget_target_lists1 field changed
      • addedInput schema / properties / owner
        Added value: +{
        +  "description": "Only lists with this owner: \"global\" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. \"global,<box_gid>\". Default: global and Firewalla-managed lists.",
        +  "type": "string"
        +}
    • Changedpause_rule3 fields changed
      • removedInput schema / properties / box
        Removed value: -{
        -  "description": "Box GID for context (required by API)",
        -  "type": "string"
        -}
      • removedInput schema / properties / duration
        Removed value: -{
        -  "default": 60,
        -  "description": "Duration in minutes to pause the rule (optional, default: 60, range: 1-1440)",
        -  "maximum": 1440,
        -  "minimum": 1,
        -  "type": "number"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "rule_id",
        -  "box"
        -]New value: +[
        +  "rule_id"
        +]
    • Changedresume_rule2 fields changed
      • removedInput schema / properties / box
        Removed value: -{
        -  "description": "Box GID for context (required by API)",
        -  "type": "string"
        -}
      • changedInput schema / required
        Previous value: -[
        -  "rule_id",
        -  "box"
        -]New value: +[
        +  "rule_id"
        +]
    • Changedsearch_alarms2 fields changed
      • changedInput schema / properties / groupBy / description
        Previous value: -"Group alarms by specified fields (comma-separated)"New value: +"Fields to group by, comma-separated, e.g. \"type\", \"status\", \"device\" or \"type,box\". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group."
      • changedInput schema / properties / query / description
        Previous value: -"Search query using Firewalla syntax. Supported fields: type:1-16 (see alarm types above), resolved:true/false, status:1/2 (active/archived), source_ip:192.168.*, region:US (country code), gid:box_id, device.name:*, message:\"text search\". Examples: \"type:8 AND region:US\" (video from US), \"type:10 AND status:1\" (active porn alerts), \"source_ip:192.168.* AND NOT resolved:true\""New value: +"Search query using Firewalla syntax. Supported fields: type:1-16 (see alarm types above), status:1/2 (active/archived), device.ip:192.168.*, region:US (country code), box.id:box_gid, device.name:*. Examples: \"type:8 AND region:US\" (video from US), \"type:10 AND status:1\" (active porn alerts), \"device.ip:192.168.* AND status:1\" (active alarms from the LAN), \"porn\" (free text: a term without a qualifier searches alarm text)"
    • Changedsearch_devices2 fields changed
      • changedInput schema / properties / query / description
        Previous value: -"Search query using Firewalla syntax. Supported fields: mac:AA:BB:CC:DD:EE:FF, ip:192.168.1.*, name:*iPhone*, online:true/false, vendor:Apple, gid:box_id, network.name:*, group.name:*. Examples: \"online:false AND vendor:Apple\", \"ip:192.168.1.* AND name:*laptop*\", \"mac:AA:* OR name:*phone*\""New value: +"Search query using Firewalla syntax. Supported fields: mac:AA:BB:CC:DD:EE:FF, ip:192.168.1.*, name:*iPhone*, online:true/false, mac_vendor:Apple, gid:box_gid, network.name:*, group.name:*. Examples: \"online:false AND mac_vendor:Apple\", \"ip:192.168.1.* AND name:*laptop*\", \"mac:AA:* OR name:*phone*\""
      • removedInput schema / properties / status
        Removed value: -{
        -  "default": "any",
        -  "description": "Filter by online status",
        -  "enum": [
        -    "online",
        -    "offline",
        -    "any"
        -  ],
        -  "type": "string"
        -}
    • Changedsearch_flows2 fields changed
      • changedInput schema / properties / groupBy / description
        Previous value: -"Group flows by specified values (e.g., \"domain,box\")"New value: +"Fields to group by, comma-separated, e.g. \"category\", \"domain\", \"device\", \"box\" or \"device,category\". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for \"device,category\") and the rest are the group's summed connection count and bytes."
      • changedInput schema / properties / query / description
        Previous value: -"Search query using Firewalla syntax. Supported fields: protocol:tcp/udp, direction:inbound/outbound/local, blocked:true/false, bytes:>1MB, domain:*.example.com, region:US (country code), category:social/games/porn/etc, gid:box_id, device.ip:192.168.*, source_ip:*, destination_ip:*. Examples: \"region:US AND protocol:tcp\", \"blocked:true AND bytes:>1MB\", \"category:social OR category:games\""New value: +"Search query using Firewalla syntax. Supported fields: protocol:tcp/udp, direction:inbound/outbound/local, status:blocked/ok, total:>1MB (download + upload in B/KB/MB/GB/TB), download:>10MB, upload:>10MB, domain:*.example.com, region:US (country code), category:social/games/porn/etc, box.id:box_gid, device.ip:192.168.*, source.ip:*, destination.ip:*, ts:>1h. Examples: \"region:US AND protocol:tcp\", \"status:blocked AND region:CN\", \"category:social OR category:games\""
    • Changedsearch_rules1 field changed
      • changedInput schema / properties / query / description
        Previous value: -"Search query using Firewalla syntax. Supported fields: action:allow/block/timelimit, target.type:domain/ip/device, target.value:*.facebook.com, status:active/paused, direction:bidirection/inbound/outbound, protocol:tcp/udp, gid:box_id, scope.type:device/network, notes:\"description text\". Examples: \"action:block AND target.value:*.social.com\", \"status:paused\", \"target.type:domain AND action:block\""New value: +"Search query using Firewalla syntax. Supported fields: action:allow/block/timelimit, target.type:domain/ip/device, target.value:*.facebook.com, status:active/paused, direction:bidirection/inbound/outbound, protocol:tcp/udp, box.id:box_gid, scope.type:device/network, notes:\"description text\". Examples: \"action:block AND target.value:*.social.com\", \"status:paused\", \"target.type:domain AND action:block\""
    • Changedsearch_target_lists3 fields changed
      • removedInput schema / properties / category
        Removed value: -{
        -  "description": "Filter by category",
        -  "type": "string"
        -}
      • changedInput schema / properties / owner / description
        Previous value: -"Filter by owner (global or box gid)"New value: +"Only lists with this owner, sent to the API: \"global\" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. \"global,<box_gid>\". Default: global and Firewalla-managed lists."
      • changedInput schema / properties / query / description
        Previous value: -"Search query for target lists. Supported fields: name:*Social*, owner:global/box_gid, category:social/games/ad/porn/etc, targets:*.facebook.com, notes:\"description text\". Examples: \"category:social\", \"owner:global AND name:*Block*\", \"targets:*.gaming.com\""New value: +"Search query for target lists. Supported fields: name:*Social*, owner:global/box_gid, category:social/games/ad/porn/etc, targets:*.facebook.com, notes:\"description text\", target_count:>100 (entries: n, >n, >=n, <n, <=n or a range 10-50), last_updated:>2026-09-01 (a date or Unix seconds, with the same comparisons). Examples: \"category:social\", \"owner:global AND name:*Block*\", \"targets:*.gaming.com\", \"target_count:>1000\", \"last_updated:<2026-01-01\""
  2. 5 tool updatesv1.3.0
    • Changedsearch_alarms1 field changed
      • changedInput schema / required
        Previous value: -[]New value: +[
        +  "query"
        +]
    • Changedsearch_devices1 field changed
      • changedInput schema / required
        Previous value: -[]New value: +[
        +  "query"
        +]
    • Changedsearch_flows1 field changed
      • changedInput schema / required
        Previous value: -[]New value: +[
        +  "query"
        +]
    • Changedsearch_rules2 fields changed
      • addedInput schema / properties / limit
        Added value: +{
        +  "description": "Maximum number of rules to return",
        +  "type": "number"
        +}
      • changedInput schema / required
        Previous value: -[]New value: +[
        +  "query"
        +]
    • Changedsearch_target_lists1 field changed
      • changedInput schema / required
        Previous value: -[]New value: +[
        +  "query"
        +]
  3. 28 tool updates
    • First observedcreate_target_list
    • First observeddelete_target_list
    • First observedget_active_alarms
    • First observedget_alarm_trends
    • First observedget_bandwidth_usage
    • First observedget_boxes
    • First observedget_device_status
    • First observedget_flow_data
    • First observedget_flow_insights
    • First observedget_network_rules
    • First observedget_network_rules_summary
    • First observedget_offline_devices
    • First observedget_recent_flow_activity
    • First observedget_rule_trends
    • First observedget_simple_statistics
    • First observedget_specific_alarm
    • First observedget_specific_target_list
    • First observedget_statistics_by_box
    • First observedget_statistics_by_region
    • First observedget_target_lists
    • First observedpause_rule
    • First observedresume_rule
    • First observedsearch_alarms
    • First observedsearch_devices
    • First observedsearch_flows
    • First observedsearch_rules
    • First observedsearch_target_lists
    • First observedupdate_target_list

TDQS

A3.8/5.0

Scored across 28 tools

Disambiguation3/5

Several tools overlap in purpose: get_offline_devices and get_device_status both read device lists and report online/offline status; get_flow_data, search_flows, and get_recent_flow_activity all query flows with different scopes; get_network_rules and search_rules both retrieve rules. Descriptions help distinguish them, but the boundaries are not always obvious.

Naming Consistency4/5

Most tools follow a consistent get_/search_/create_/update_/delete_/pause_/resume_ verb pattern with clear noun objects. Minor deviations like get_network_rules_summary and get_specific_target_list vs get_specific_alarm are acceptable, but the mix of get_ and search_ for similar resources (flows, alarms, rules) creates slight inconsistency.

Tool Count3/5

28 tools is on the heavy side for a single server, though the domain (Firewalla MSP API) is broad with devices, alarms, flows, rules, target lists, and statistics. The count is borderline: many tools are convenience wrappers or variations on the same resource, which could be consolidated.

Completeness4/5

The tool set covers the main Firewalla MSP resources well: devices, alarms, flows, rules, target lists, and statistics. Minor gaps exist (e.g., no create/update/delete for rules or alarms, no device management actions), but the core read and search operations are comprehensive.

Maintenance

ActivityMaintained
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Provides real-time access to Firewalla firewall data through 28 specialized tools for network monitoring, security analysis, bandwidth tracking, and firewall rule management. Enables users to query security alerts, analyze network flows, monitor device status, and manage firewall configurations through natural language.
    516 npm
    1
    -
  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables interaction with Firewalla network security devices for network monitoring, device management, traffic analysis, and security rule configuration through MCP tools.
    -
  • A
    license
    A
    quality
    B
    maintenance
    A secure MCP server for managing OPNsense firewalls through AI assistants. Provides 81 tools across system, firewall, network, DNS, DHCP, VPN, HAProxy, services, diagnostics, and security domains.
    81
    232 PyPI
    21
    MIT