firewalla-mcp-server
This server provides real-time access to Firewalla firewall data through 24 read-only tools and 11 opt-in write tools, enabling network monitoring, security analysis, and management via MCP clients.
Monitor security alarms: Retrieve active alarms, search by type/region/device, get details, view alarm trends.
Analyze network traffic: Query flows, search with filters (protocol, domain, region, status), get recent activity, flow insights by category, bandwidth usage by device.
Manage devices: Check online/offline status, search by name/IP/MAC, list offline devices, rename devices (write).
Inspect firewall rules: Get rules, search by action/target/status, view rule summaries, trends.
Manage rules (opt-in): Create, pause, resume, delete firewall rules.
Manage target lists (opt-in): Create, update, delete, search target lists; view global and box-specific lists.
Handle alarms (opt-in): Archive, mute, delete alarms.
Get statistics: Account-wide counts, top boxes/regions by blocked flows, simple statistics, alarm/rule trends.
Advanced search: Query syntax with filters and logical operators across flows, alarms, rules, devices, target lists.
Geographic insights: Filter by region, view top regions by blocked traffic, enrich IP geolocation.
Write tools are opt-in: Off by default (read-only mode); enable with FIREWALLA_ENABLE_WRITE_TOOLS=true.
Supports multiple boxes: Scope queries to specific box or aggregate across all; list boxes the token can access.
Multiple transports: Stdio for local MCP clients, HTTP for Docker/remote access with bearer token and host/origin validation.
Response format options: JSON (default) or markdown for readable summaries.
Supports containerized deployment with Docker, allowing the MCP server to run in isolated containers with appropriate environment variables.
Provides real-time access to Firewalla firewall data through 28 specialized tools for network monitoring, security analysis, bandwidth tracking, rule management, and target list management.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@firewalla-mcp-servershow me recent security alerts from the last hour"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Firewalla MCP Server
A Model Context Protocol (MCP) server that provides real-time access to Firewalla firewall data through 24 read-only tools, plus 11 opt-in write tools, compatible with any MCP client.
Why Firewalla MCP Server?
Simple Network Security Integration
24 read-only tools for network monitoring and analysis: 19 Direct API Endpoints + 5 Convenience Wrappers
11 write tools, off unless
FIREWALLA_ENABLE_WRITE_TOOLS=true, so by default nothing can change your boxAdvanced Search with query syntax and filters
Clean, Verified Architecture with corrected API schemas
Related MCP server: Firewalla MCP Server
Features
Real-time Firewall Data: Query security alerts, network flows, and device status
Security Analysis: Get insights on threats, blocked attacks, and network anomalies
Bandwidth Monitoring: Track top bandwidth consumers and usage patterns
Rule Management: View firewall rules; with the write tools, create, pause, resume and delete them
Target Lists: View target lists; with the write tools, create, update and delete them
Search Tools: Query syntax with filters and logical operators
Client Setup Guides
Client | Quick Start | Full Guide |
Claude Desktop |
| |
Claude Code |
| |
VS Code | Install MCP extension → Configure server | |
Cursor | Install Claude Code → VSIX method | |
Roocode | Install MCP support → Configure server | |
Cline | Configure in VS Code → Enable MCP | |
Open WebUI | mcpo, or a native MCP connection over HTTP |
How It Works
Claude Desktop/Code ↔ MCP Server ↔ Firewalla APIThe MCP server acts as a bridge between Claude and your Firewalla firewall, translating Claude's requests into Firewalla API calls and returning the results in a format Claude can understand.
Prerequisites
Node.js 18+ and npm. On Node 18-22, npm prints an
EBADENGINEwarning for geoip-lite 2.x, which declares Node 24 for its database update script; the lookups the server uses run on Node 18 and later, and CI tests 18, 20, 22 and 24.Firewalla MSP account with API access
Your Firewalla device online and connected
Quick Start
1. Installation
Option A: Install from npm (Recommended)
# Install globally
npm install -g firewalla-mcp-server
# Or install locally in your project
npm install firewalla-mcp-serverOption B: Use Docker
Warning: Not for production use – secrets visible in process list
The examples below pass credentials directly in the command line, which exposes them to process listing and shell history. For production use, consider these secure alternatives:
Use
--env-filewith a.envfile:docker run --env-file .env ...Set environment variables in your shell before running Docker
Use Docker secrets for orchestration environments
Stdio Transport (Default - for Claude Desktop integration):
# Using Docker Hub image (minimal config)
docker run -it --rm \
-e FIREWALLA_MSP_TOKEN=your_token \
-e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
amittell/firewalla-mcp-server
# Or with optional box filter
docker run -it --rm \
-e FIREWALLA_MSP_TOKEN=your_token \
-e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
-e FIREWALLA_BOX_ID=your_box_gid \
amittell/firewalla-mcp-server
# Or build locally
docker build -t firewalla-mcp-server .
docker run -it --rm \
-e FIREWALLA_MSP_TOKEN=your_token \
-e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
firewalla-mcp-server
# Recommended: Using env file (more secure)
docker run -it --rm --env-file .env amittell/firewalla-mcp-serverHTTP Transport (for standalone Docker containers and external access):
# Run with HTTP transport on port 3000
docker run -d --name firewalla-mcp \
-p 3000:3000 \
-e MCP_TRANSPORT=http \
-e MCP_HTTP_PORT=3000 \
-e MCP_HTTP_BEARER_TOKEN=a_long_random_secret \
-e FIREWALLA_MSP_TOKEN=your_token \
-e FIREWALLA_MSP_ID=yourdomain.firewalla.net \
amittell/firewalla-mcp-server
# Add FIREWALLA_BOX_ID if you want to filter to a specific box
# -e FIREWALLA_BOX_ID=your_box_gid \
# The server will be accessible at http://localhost:3000/mcp, and clients
# send the header: Authorization: Bearer a_long_random_secret
# Using env file (recommended)
docker run -d --name firewalla-mcp \
-p 3000:3000 \
--env-file .env \
amittell/firewalla-mcp-server
# For docker-compose
cat > docker-compose.yml << EOF
version: '3.8'
services:
firewalla-mcp:
image: amittell/firewalla-mcp-server
ports:
- "3000:3000"
environment:
- MCP_TRANSPORT=http
- MCP_HTTP_PORT=3000
# The image already listens on every interface of the container
- MCP_HTTP_HOST=0.0.0.0
- MCP_HTTP_BEARER_TOKEN=\${MCP_HTTP_BEARER_TOKEN}
# Other containers reach it as http://firewalla-mcp:3000/mcp
- MCP_HTTP_ALLOWED_HOSTS=firewalla-mcp
- FIREWALLA_MSP_TOKEN=\${FIREWALLA_MSP_TOKEN}
- FIREWALLA_MSP_ID=\${FIREWALLA_MSP_ID}
# Optional: filter to specific box
# - FIREWALLA_BOX_ID=\${FIREWALLA_BOX_ID}
restart: unless-stopped
EOF
docker-compose up -dThe image sets MCP_HTTP_HOST=0.0.0.0 so that a published port reaches the server. -p 3000:3000 publishes it on every interface of the Docker host, so anyone on your network can reach it: set MCP_HTTP_BEARER_TOKEN (for example openssl rand -hex 32) whenever you publish the port, or publish it on this machine only with -p 127.0.0.1:3000:3000. The server answers only requests whose Host header is localhost, 127.0.0.1 or [::1], so a client that connects by another name, such as the host's LAN address or a compose service name, needs that name in MCP_HTTP_ALLOWED_HOSTS. See HTTP transport security.
Option C: Install from source
git clone https://github.com/amittell/firewalla-mcp-server.git
cd firewalla-mcp-server
npm install
npm run build2. Configuration
Create a .env file with your Firewalla credentials:
# Required
FIREWALLA_MSP_TOKEN=your_msp_access_token_here
FIREWALLA_MSP_ID=yourdomain.firewalla.net
# Optional - filters all queries to a specific box
# FIREWALLA_BOX_ID=your_box_gid_here
# Optional - default box for single-box operations, without filtering queries
# FIREWALLA_DEFAULT_BOX_ID=your_box_gid_here
# Optional - register the 11 write tools (default: off, read-only)
# FIREWALLA_ENABLE_WRITE_TOOLS=trueGetting Your Credentials:
Log into your Firewalla MSP portal at
https://yourdomain.firewalla.netYour MSP ID is the full domain (e.g.,
company123.firewalla.net)Generate an access token in API settings
(Optional) Find your Box GID in device settings to filter queries to a specific box, or retrieve available boxes using the
get_boxestool
Box ID is optional. Without FIREWALLA_BOX_ID, queries cover every box on the account. The few operations that act on one box (get_specific_alarm, archive_alarm, mute_alarm, delete_alarm, create_rule, rename_device) take a gid argument, and without one they use FIREWALLA_BOX_ID, then FIREWALLA_DEFAULT_BOX_ID, then the account's only box. On an account with several boxes and none of those set, get_specific_alarm, archive_alarm, mute_alarm and delete_alarm check each box, and create_rule and rename_device refuse and list the boxes.
Test mode. MCP_TEST_MODE=true starts the server without credentials, to check that it starts: it uses a dummy token, API (https://test.firewalla.net) and box instead of your settings, always runs on stdio, and cannot read your Firewalla data. With NODE_ENV=production the server refuses test mode and exits with code 1. The Docker image sets NODE_ENV=production, so pass -e NODE_ENV=development along with -e MCP_TEST_MODE=true.
Transport Configuration
The MCP server supports two transport modes:
Stdio Transport (Default): Standard input/output communication for Claude Desktop and similar MCP clients
MCP_TRANSPORT=stdioHTTP Transport: HTTP server mode for Docker containers, MCP orchestrators, and external access
MCP_TRANSPORT=http
MCP_HTTP_PORT=3000 # Default: 3000
MCP_HTTP_PATH=/mcp # Default: /mcp. Other paths get 404
MCP_HTTP_HOST=127.0.0.1 # Address to listen on, no port (::1 or [::1] for IPv6). Default: 127.0.0.1 (0.0.0.0 in the Docker image)
MCP_HTTP_BEARER_TOKEN= # When set, clients must send Authorization: Bearer <token>
MCP_HTTP_ALLOWED_HOSTS= # More Host header names to accept, comma-separated
MCP_HTTP_ALLOWED_ORIGINS= # Browser origins to accept, comma-separated, e.g. http://localhost:6274HTTP transport security: every request can spend your MSP token, so the HTTP server follows the security rules of the MCP transport specification:
It listens on
127.0.0.1, this machine only. SetMCP_HTTP_HOST=0.0.0.0(or one address) to accept other machines, and setMCP_HTTP_BEARER_TOKENwith it: the server logs a warning when it listens beyond loopback without a token.With
MCP_HTTP_BEARER_TOKENset, a request withoutAuthorization: Bearer <token>gets 401.A request whose
Hostheader is notlocalhost,127.0.0.1,[::1], theMCP_HTTP_HOSTaddress or a name inMCP_HTTP_ALLOWED_HOSTSgets 403. This stops DNS rebinding, where a web page points its own domain name at your machine.A request with an
Originheader, which browsers send, gets 403 unless the origin is inMCP_HTTP_ALLOWED_ORIGINS. Non-browser MCP clients send noOriginand are not affected. An allowed origin gets CORS headers on every answer, refusals such as a 401 for a missing token included, so a web page on it can call the server and read why a request was refused.The MCP endpoint is the
MCP_HTTP_PATHpath exactly:/mcp,/mcp/and either with a query string. Any other path, such as/mcpxor/mcp/x, gets 404, CORS preflight requests included.A request body may be at most 1 MB, and a client has 10 seconds to send the headers and 30 seconds for the whole request. An answer given without reading the body (401, 403, 404, 405, a malformed session ID, a CORS preflight) closes the connection, so a client cannot hold one open by sending the body slowly.
A session ends when the client sends DELETE, after
MCP_SESSION_IDLE_TIMEOUT_MSwithout a request (default 30 minutes), or when the server restarts. A request with the ID of a session the server does not hold gets 404, and the transport specification has the client start a new session with a newinitialize. A request with no session ID, other thaninitialize, gets 400.
When to use HTTP transport:
Running in Docker containers independently
Accessing from MCP orchestrators (e.g., open-webui)
Multiple clients need to connect to the same server instance
Network-based access to the MCP server
When to use stdio transport:
Claude Desktop integration (default)
Claude Code CLI integration
Single-process MCP client setups
Standard MCP client configurations
3. Build and Start
npm run build
npm run mcp:start4. Connect Claude Desktop
Add this configuration to your Claude Desktop claude_desktop_config.json:
If installed via npm
{
"mcpServers": {
"firewalla": {
"command": "npx",
"args": ["firewalla-mcp-server"],
"env": {
"FIREWALLA_MSP_TOKEN": "your_msp_access_token_here",
"FIREWALLA_MSP_ID": "yourdomain.firewalla.net",
"FIREWALLA_BOX_ID": "your_box_gid_here"
}
}
}
}If using Docker
{
"mcpServers": {
"firewalla": {
"command": "docker",
"args": ["run", "-i", "--rm",
"-e", "FIREWALLA_MSP_TOKEN=your_token",
"-e", "FIREWALLA_MSP_ID=yourdomain.firewalla.net",
"-e", "FIREWALLA_BOX_ID=your_box_gid",
"amittell/firewalla-mcp-server"
]
}
}
}If installed from source
{
"mcpServers": {
"firewalla": {
"command": "node",
"args": ["/full/path/to/firewalla-mcp-server/dist/server.js"],
"env": {
"FIREWALLA_MSP_TOKEN": "your_msp_access_token_here",
"FIREWALLA_MSP_ID": "yourdomain.firewalla.net",
"FIREWALLA_BOX_ID": "your_box_gid_here"
}
}
}
}Config file locations:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
5. Next Steps
See USAGE.md for practical examples and common queries
Check TROUBLESHOOTING.md if you encounter issues
Review client-specific setup guides in docs/clients/
Usage Examples
Step-by-Step First Use
1. Verify Connection After completing the setup, verify the MCP server is working:
# Start the server
npm run mcp:start
# You should see output like:
# MCP Server starting...
# Firewalla client initialized
# Server ready on stdio transport2. Test with Claude Open Claude Desktop and try these starter queries:
Basic Health Check:
"Can you check my Firewalla status and show me a summary?"This uses: firewall_summary resource + get_simple_statistics tool
Security Overview:
"What security alerts do I have? Show me the 5 most recent ones."This uses: get_active_alarms tool with limit parameter
Practical Workflows
Daily Security Review:
"Give me today's security report. Include:
1. Any new security alerts
2. Top 3 devices using bandwidth
3. Any devices that went offline
4. Status of critical firewall rules"Investigating Suspicious Activity:
"I noticed unusual traffic. Can you:
1. Show me all security and abnormal upload alarms from the last 4 hours
2. Find any blocked connections to external IPs
3. Check which devices had the most network activity"Network Troubleshooting:
"A device seems to have connectivity issues. Can you:
1. Check if device 192.168.1.100 is online
2. Show its recent network flows
3. See if any rules are blocking its traffic"Bandwidth Investigation:
"Our internet is slow. Help me find the cause:
1. Show top 10 bandwidth users in the last hour
2. Look for any devices with unusual upload/download patterns
3. Check for any streaming or video traffic"Advanced Search Examples
Find Specific Threats:
search for: security activity alarms from IP range 10.0.0.* in the last 24 hoursUses: search_alarms with query: "type:1 AND device.ip:10.0.0. AND ts:>=<unix time 24 hours ago>"*
Analyze Rule Effectiveness:
"Show me firewall rules that blocked the most connections this week"Uses: get_network_rules + search_flows for blocked traffic analysis
Device Behavior Analysis:
"Find all devices that were online yesterday but are offline now"Uses: search_devices with temporal queries + get_offline_devices
Troubleshooting Common Issues
Connection Problems: If you get authentication errors:
Verify your
.envfile has correct credentialsCheck your MSP token hasn't expired
Confirm your Box ID is the full GID format
Empty Results: If queries return no data:
Check your Firewalla is online and reporting
Verify the time range isn't too narrow
Try broader search terms first
Performance Issues: If responses are slow:
Reduce the limit parameter in queries
Use more specific time ranges
Check your network connection to the MSP API
Available Tools (24 read-only, 11 opt-in write tools)
Core Tools
Security: Get alarms, analyze threats
Network: Monitor traffic flows, track bandwidth usage
Devices: Check device status, find offline devices
Rules: View firewall rules and their summary
Search: Advanced search across all data types
Analytics: Statistics, trends, and geographic analysis
Target Lists: View security target lists
Quick Reference
Security: get_active_alarms, get_specific_alarm
Network: get_flow_data, get_recent_flow_activity, get_bandwidth_usage
Devices: get_device_status, get_offline_devices, get_boxes
Rules: get_network_rules, get_network_rules_summary
Target lists: get_target_lists, get_specific_target_list
Search: search_flows, search_alarms, search_rules, search_devices, search_target_lists
Analytics: get_simple_statistics, get_statistics_by_region, get_statistics_by_box, get_flow_insights, get_flow_trends, get_alarm_trends, get_rule_trends
Write (opt-in): create_rule, delete_rule, pause_rule, resume_rule, create_target_list, update_target_list, delete_target_list, rename_device, archive_alarm, mute_alarm, delete_alarmResponse format
Every read-only tool (readOnlyHint: true) takes an optional response_format. json, the default, returns the compact JSON response. markdown returns the same response as markdown for reading: a heading with the tool name and the number of records in each list, the fields as bullet lists, and each list of records as a table of at most 8 columns and DEFAULT_PAGE_SIZE rows (default 100). Fields with one value in every record are stated once above the table, fields that are not columns are named below it, and a list of one record is shown as a list of its fields. The last line says what the view leaves out (the response's meta block, rows past the cap, cells cut at 120 characters); response_format: json returns all of it. Values and field names are escaped, since device names, domains and alarm messages come from the network: a character that could start markdown or HTML ([, <, *, `, & before an entity, ://, @, ...) gets a backslash, so a renderer shows [a](https://...) or <img ...> as text. Addresses, MACs, gids and timestamps read as they are. Omitting response_format or sending null means json. Errors are JSON in either format, and the tools that change state do not take response_format.
get_device_status with {"limit": 2, "response_format": "markdown"}, shortened:
## get_device_status (devices: 2)
- **total_devices:** 2
- **online_devices:** 1
- **devices:** 2 records, under devices below
### devices (2 records)
Same in every record: gid = 00000000-0000-0000-0000-000000000000; ip_reserved = false.
| name | id | last_seen | ip | online |
| --- | --- | --- | --- | --- |
| nas | AA:BB:CC:DD:EE:01 | 2026-09-26T10:00:00.000Z | 192.168.1.10 | true |
| printer | AA:BB:CC:DD:EE:02 | 2026-09-25T18:30:00.000Z | 192.168.1.11 | false |
_This view leaves out the meta block (request_id req_1790000000000_abc123). Call again with `response_format: json` for the full JSON response._Write tools (opt-in)
Every tool that changes something is off by default, so the server is read-only unless you turn them on. Set FIREWALLA_ENABLE_WRITE_TOOLS=true to register the 11 write tools: create_rule, delete_rule, pause_rule and resume_rule (rules), create_target_list, update_target_list and delete_target_list (target lists), rename_device (devices), and archive_alarm, mute_alarm and delete_alarm (alarms). Without it, calling one answers "Unknown tool" and sends nothing. MCP clients that honor tool annotations can ask before calling them; see Tool annotations.
Up to 1.5.0, pause_rule, resume_rule and the three target-list tools were always registered. If you use them, set FIREWALLA_ENABLE_WRITE_TOOLS=true.
The write tools need an MSP API token with write access. Firewalla said on 2026-09-08 that MSP 2.12 adds read-only API tokens, which cannot make changes. When a write gets HTTP 403, the error says the token may be read-only and that the write tools need a token with write access. A 403 can also mean the request named a box the token cannot access, and the error says that too; get_boxes lists the boxes the token can access.
IDs that go into a request path (id, rule_id, alarm_id, gid, device_id) are checked before anything is sent. One holding /, a backslash, ?, #, %, whitespace or a control character, or that is . or .., is refused as a validation error naming the argument; leading or trailing whitespace is refused too, not trimmed. Rule IDs such as <box gid>:<n>, MAC device IDs and ovpn: device IDs are accepted.
create_rule and rename_device act on one box: gid, else FIREWALLA_BOX_ID or FIREWALLA_DEFAULT_BOX_ID, else the account's only box. On a multi-box account with none of those, they refuse without writing anything, because the MSP API applies a rule with no gid to every box in the account, including boxes added later. delete_rule needs MSP 2.11.0 or later.
archive_alarm and mute_alarm need MSP 2.11.0 or later. archive_alarm takes an alarm out of the active alarms and does nothing else: future matching traffic can still raise new alarms. mute_alarm archives the alarm and has the box create a lasting silence exception. Its target_type says what is silenced: alarmType every future alarm of that alarm's type, whatever the destination; domain a domain and its subdomains; ip one address. Its scope_type says for which devices: all of them, or the one device, group, user or network named by scope_value. The mute is checked against the documented request model before anything is sent, and this server has no tool to remove the exception afterwards. delete_alarm deletes the alarm for good (measured on 2026-09-26: the alarm was gone afterwards); use archive_alarm to keep it among the archived alarms. Alarm IDs are per box, and the same ID can name different alarms on different boxes: all three tools use the alarm's gid, else FIREWALLA_BOX_ID, else check each box, and they refuse when several boxes have that alarm ID and none of them is FIREWALLA_DEFAULT_BOX_ID. All three read the alarm before writing, so a wrong ID fails without changing anything.
Tool annotations
Every tool carries MCP tool annotations: a title, readOnlyHint, and openWorldHint: true, since each one calls the Firewalla MSP API. All get_* and search_* tools are read-only. The tools that change state have readOnlyHint: false, and all of them need FIREWALLA_ENABLE_WRITE_TOOLS=true:
Tool |
|
|
| false | true |
| false | false |
| true | true |
| true | false |
| true | true |
| false | true |
| false | true |
| true | false |
| true | true |
pause_rule and resume_rule check the rule's status first and change nothing when it is already paused or active. Clients that honor annotations can ask before calling any tool that is not read-only.
Development
Scripts
npm run dev # Start development server with hot reload
npm run build # Build TypeScript to JavaScript
npm run test # Run all tests
npm run test:watch # Run tests in watch mode
npm run lint # Run ESLint
npm run lint:fix # Fix ESLint issuesMCP Execution Methods
Why npx for MCP servers?
Version Management: Always uses the correct/latest version
Dependency Resolution: Handles package dependencies automatically
No global installation required: Works without global installation
MCP Standard: Follows Model Context Protocol conventions
Reliable: Works consistently across different environments
Alternative execution methods:
# Development (from source)
npm run mcp:start
# Production (npm installed)
npx firewalla-mcp-server
# Direct execution (from source after build)
node dist/server.jsLaunch checks in CI
scripts/launch-smoke.mjs starts the server the way a user would and requires an answer to an MCP initialize over stdio. It uses dummy credentials, and initialize is answered locally, so nothing is sent to Firewalla.
The CI workflow's
launchjob packs the package and starts it through the global bin,npxandnode dist/server.json Linux, macOS and Windows.The Docker Build workflow runs on pull requests that touch the Dockerfile, the package files or the Docker workflows. It builds the image for amd64, arm64 and arm/v7, then runs the amd64 image with
docker run -i --rmand requiresserverInfo.versionto equalpackage.json's version. Run the workflow by hand with theimageinput (for exampleamittell/firewalla-mcp-server:1.4.1) to pull and check a published image instead.
To run the Docker check locally:
docker build -t firewalla-mcp-server:local .
node scripts/launch-smoke.mjs --docker firewalla-mcp-server:local
# A published image; the expected version comes from the tag
node scripts/launch-smoke.mjs --docker amittell/firewalla-mcp-server:1.4.1 --pullProject Structure
firewalla-mcp-server/
├── src/
│ ├── server.ts # Main MCP server
│ ├── firewalla/ # Firewalla API client
│ ├── tools/ # MCP tool implementations
│ ├── resources/ # MCP resource implementations
│ └── prompts/ # MCP prompt implementations
├── tests/ # Test files
├── docs/
│ └── firewalla-api-reference.md # API documentation
├── CLAUDE.md # Comprehensive development guide
├── SPEC.md # Technical specifications
└── README.md # This fileDocumentation
README.md (this file) - Setup and basic usage
USAGE.md - Simple usage guide with examples
TROUBLESHOOTING.md - Common issues and solutions
docs/clients/ - Client-specific setup guides
CLAUDE.md - Development guide and commands
Security
See SECURITY.md to report a vulnerability. For the HTTP transport, see HTTP transport security.
MSP tokens are stored securely in environment variables
No credentials are logged or stored in code
Requests are paced to
API_RATE_LIMITper 5 minutes (default 100, the MSP API's quota per token). A request that cannot start within 20 s fails at once, saying when capacity returnsInput validation prevents injection attacks
Device names, domains and alarm messages are set by the devices and sites on your network, so the server treats them as untrusted: characters that do not display are shown as markers such as
<U+E0041>, the prompts quote API data only inside a marked data block, and theinitializeinstructions tell the client to treat results as data. See Untrusted dataAll API communications use HTTPS
Known Behaviors and Limitations
Category Classification
Flow Categories: Many network flows may show as empty category ("") in the Firewalla API response. This is expected behavior - Firewalla categorizes traffic when it recognizes the domain/service (e.g., "av" for audio/video, "social" for social media).
Target List Categories: Some target lists may show category as "unknown". This is normal for user-created or certain system lists.
Timeline: Category classification happens at the Firewalla device level and may take time to build up meaningful categorization data.
Data Characteristics
Response Sizes: The
get_recent_flow_activitytool returns up to 150 recent flows to stay within token limits. For larger datasets or historical analysis, usesearch_flowswith time filters for more targeted queries.Geographic Data: IP geolocation is enriched by the MCP server and includes country, city, and risk scores when available.
API Limitations
Alarm Deletion: In July 2025 the MSP API answered
DELETE /v2/alarms/{gid}/{aid}with{"message": "success", "success": true}and kept the alarm, sodelete_alarmwas withdrawn. Measured again on 2026-09-26, the same request deleted the alarm (a GET of it then answered 404, and the archived alarm count fell by one), anddelete_alarmis back as an opt-in write tool.
Troubleshooting
Quick Fixes
Server won't start:
# Clean and rebuild
npm run clean
npm run build
# If build fails, try:
npm install
npm run buildAuthentication errors:
Check your MSP token is valid
Verify Box ID format (long UUID)
Confirm MSP domain is correct
No data returned:
Try broader queries: "last week" vs "last hour"
Check if Firewalla is online
Test with: "show me basic statistics"
Slow responses:
Add limits: "top 10 devices"
Use shorter time ranges
Restart the server
Debug Mode
Enable detailed logging:
DEBUG=mcp:* npm run mcp:startFor more detailed troubleshooting, see TROUBLESHOOTING.md
Contributing
Fork the repository
Create a feature branch
Make your changes
Add tests for new functionality
Run the test suite
Submit a pull request
What's New
Release notes for every version, newest first, are in CHANGELOG.md. Changes that are merged but not yet released are listed there under Unreleased.
License
Support
For issues and questions:
Check the troubleshooting guide
Review the technical specifications
Open an issue on GitHub
GitHub Repository
Repository: https://github.com/amittell/firewalla-mcp-server
Quick Links
Repository Stats
Available Tools
28 toolscreate_target_listA
Create a new target list (POST /v2/target-lists); each call creates another list. owner global makes it shareable across all boxes, a box GID ties it to that box.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Target list name (required, max 24 chars) | |
| notes | No | Additional description (optional) | |
| owner | Yes | Owner: "global" or box GID (required) | |
| targets | Yes | Array of domains, IPs, or CIDR ranges (required) | |
| category | No | Content category (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnlyHint=false, idempotentHint=false), the description spells out 'each call creates another list,' making non-idempotency concrete for an agent. It also discloses the scope consequence of owner selection: 'global' shares the list across all boxes, while a box GID ties it to that box.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences deliver the core operation, endpoint, non-idempotency, and owner semantics without unnecessary information. The most decision-relevant details are front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with a fully documented input schema, the description sufficiently covers the operation, endpoint, non-idempotent behavior, and owner-scope options. Minor omissions such as return format are acceptable because no output schema is declared and the schema already describes parameter constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all five parameters, giving a baseline of 3. The description adds semantic value by explaining the owner field's effect ('shareable across all boxes' vs 'ties it to that box') rather than only restating accepted values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Create a new target list' and identifies the exact REST endpoint (POST /v2/target-lists), naming the action and resource unambiguously. The phrase 'each call creates another list' also distinguishes it from update/delete target-list siblings by emphasizing non-idempotent creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies this tool is for creating a fresh target list, but it does not explicitly state when not to use it or mention update_target_list/delete_target_list as alternatives for modifying or removing lists. Usage context is inferable rather than directly instructed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_target_listADestructiveIdempotent
Permanently delete a target list (DELETE /v2/target-lists/{id}); cannot be undone. The tool does not check whether a rule still targets the list.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Target list ID to delete (required) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation destructive and non-read-only; the description adds the permanence warning and the important caveat that no rule-reference check is performed. This is meaningful behavioral context beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the action and endpoint, then add the two essential caveats. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive delete with no output schema, the description covers the core action, permanence, and the rule-reference caveat. It does not describe response/error behavior, but that is a minor gap given the tool's simplicity and annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter id is fully described in the schema (100% coverage), so the description adds no additional parameter-level detail. Baseline 3 is appropriate because the schema carries the semantic weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Permanently delete'), the resource ('target list'), the HTTP endpoint, and a key consequence ('cannot be undone'). This clearly distinguishes it from sibling create/update/get target-list tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case: deleting a target list by ID, and warns that it is permanent and does not check for rule references. It does not explicitly name alternatives or conditions when to avoid use, so guidance is implied rather than fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_active_alarmsARead-only
Retrieve active security alarms from the Firewalla MSP API (GET /v2/alarms): status:1 is added unless the query names a status (status:2 for archived alarms). Without a ts: qualifier the API covers the last 30 days. Returns up to limit alarms and a cursor for the next page, or groups with groupBy. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Results per page (optional, default: 200, API maximum: 500) | |
| query | No | Search query for filtering alarms. Active alarms only (status:1) unless the query names a status. Use type:N where N is: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Examples: type:8 (video), type:10 (porn), region:US, device.ip:192.168.* | |
| cursor | No | Pagination cursor from previous response | |
| sortBy | No | Sort alarms (default: ts:desc) | |
| groupBy | No | Fields to group by, comma-separated, e.g. "type", "status", "device" or "type,box". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true and openWorldHint, but the description adds extensive behavioral detail: the default status:1 filter unless overridden, the default time range, pagination via cursor, grouping behavior and response format, and scope handling. This goes far beyond annotations and clearly discloses operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise yet information-dense. It front-loads the main purpose and then covers default status, time range, pagination, grouping, and scope in a few well-structured sentences. Every sentence contributes essential information with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a retrieval tool with 5 optional parameters and no output schema, the description covers all critical aspects: default filters, pagination mechanism, grouping output shape, and scope. It does not detail alarm object fields, but that is not necessary for invoking the tool correctly; the agent knows what to expect in the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. However, the description adds meaning by explaining the default status injection, the query syntax with examples (type codes, region, device.ip), and the grouping output structure. This adds value beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the verb 'Retrieve' and the resource 'active security alarms from the Firewalla MSP API' with the endpoint (GET /v2/alarms). It distinguishes itself by detailing default status, time window, pagination, grouping, and scope, making it clear what this tool does and how it differs from a generic search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use the tool: it covers active alarms by default, explains the default 30-day window, and notes scoping to a specific box ID. It does not explicitly name alternative tools or when not to use it, but the context is strong enough for an agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_alarm_trendsARead-only
Alarms generated per day, from GET /v2/trends/alarms: one point per day for the last 30 days, the last point being today so far. period (default 30d) returns the days that overlap it. The trends API takes no box, so it covers every box (or the group) even with FIREWALLA_BOX_ID set.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Get trends for a specific box group | |
| period | No | Return the days that overlap this period (default: 30d). The API has no finer resolution than a day, so 1h returns today so far and 24h returns yesterday and today | 30d |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond readOnlyHint/openWorldHint, it discloses daily granularity, that the last point is today so far, period overlap behavior, and that the API ignores box scoping. This is valuable behavioral context, though the exact response shape is left unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences front-load the core result and endpoint, then add scoping and period behavior without redundancy. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only two-parameter tool, the description covers endpoint, time window, granularity, period semantics, and box/group scoping. No output schema exists, but the daily-point framing gives enough shape to understand the return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters are fully described in the schema (100% coverage), including enum values and period behavior. The description adds the note that the trends API takes no box, but this is more environmental context than parameter-level semantics; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it returns daily alarm counts from GET /v2/trends/alarms, with an explicit time window and granularity. The phrase 'Alarms generated per day' identifies the metric and resource, differentiating it from alarm search/detail tools, though it does not name a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: this is a trends API that covers every box (or group) even when FIREWALLA_BOX_ID is set, which signals when to use it over box-scoped alarm tools. It does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bandwidth_usageARead-only
Top devices by upload plus download over the period, summed client-side from up to 10 times limit (1,000 at most) of the period's most recent flows (GET /v2/flows, 500 per request), so on a busy network the totals cover a sample. Scoped to box, else FIREWALLA_BOX_ID, else every box.
| Name | Required | Description | Default |
|---|---|---|---|
| box | No | Only flows of this box (box gid). Defaults to FIREWALLA_BOX_ID; without either, every box. | |
| limit | No | Number of top devices to return | |
| period | Yes | Time period for bandwidth calculation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark read-only and open-world, so the description carries the behavioral burden. It adds a substantial behavioral profile: client-side summation over at most 10x limit flows, cap of 1,000, sample-based incompleteness, and box scoping precedence. This is exactly the information an agent needs to interpret results appropriately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the outcome before the algorithm details. It is dense but every clause contributes; only the fairly long single sentence and the embedded API endpoint reference add slight complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only aggregate tool with no output schema, the description covers the key operational details: period, scope, sample limitation, and result ordering. It could still mention return shape or what happens with no flows, but the definition is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all three params, so baseline is 3. Description adds meaning by relating `limit` to the client-side sampling cap (up to 10x limit, max 1,000) and by giving the `box` resolution order (box → FIREWALLA_BOX_ID → every box), which the schema doesn't provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific result: top devices by combined upload+download over a period. It also differentiates itself from sibling tools by describing a client-side summation algorithm and scope resolution, though it never names a specific sibling tool or explicitly contrasts with alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: use this to get top bandwidth consumers over a period, scoped to a box or all boxes, and warns that on busy networks totals are sampled. However, it doesn't state when not to use it or explicitly mention alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_boxesARead-only
List the Firewalla boxes this MSP token can see (GET /v2/boxes, optionally one box group), with online status, model and version. Not limited by FIREWALLA_BOX_ID.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Get boxes within a specific group (requires group ID) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds useful behavioral context: the endpoint, the optional group filter, the returned fields, and the scope being all boxes visible to the MSP token rather than a single box.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, information-dense sentence front-loads the action and resource, then adds the endpoint, optional filter, output fields, and scope limitation. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with zero required parameters and one optional parameter, the description covers what the tool returns, its scope, and the relevant endpoint. No output schema exists, but the returned fields are stated directly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema already documents the optional group parameter and the requirement for a group ID. The description only restates that a group is optional, adding minimal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') with a clear resource ('Firewalla boxes this MSP token can see'), names the exact endpoint, and states what information is returned. It is distinct from siblings like get_device_status or search_devices.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the access scope explicit ('MSP token can see') and clarifies that it is not limited by FIREWALLA_BOX_ID, which helps agents decide when this is the right tool. It does not explicitly name alternatives or exclusions, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_device_statusARead-only
Check online/offline status of devices on the Firewalla network. Reads the device list from GET /v2/devices (box, else FIREWALLA_BOX_ID, else every box; group limits it to a box group) and returns up to limit devices, sorted by name.
| Name | Required | Description | Default |
|---|---|---|---|
| box | No | Get devices under a specific Firewalla box (requires box ID) | |
| group | No | Get devices under a specific box group (requires group ID) | |
| limit | Yes | Maximum number of devices to return (required) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description reveals useful behavior: it reads from GET /v2/devices, explains the box/group selection precedence, and states the result is limited and sorted by name. This gives the agent meaningful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary purpose, and packs the key behavioral details into a compact parenthetical. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool, the description is quite complete: it covers purpose, data source, scoping rules, limit behavior, and sort order. With no output schema, it could say slightly more about the exact shape of the returned device objects, but the core behavior is sufficiently clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning by explaining parameter precedence: box takes priority, then FIREWALLA_BOX_ID, then every box, while group further restricts. It also clarifies that limit caps the returned device count.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: checking online/offline status of Firewalla network devices. It does not explicitly differentiate from related siblings like get_offline_devices or search_devices, so it misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Check online/offline status of devices' implies a clear use case, and the selection logic for box/group provides scoping context. However, it does not state when to prefer this tool over siblings or mention any exclusions, so guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_flow_dataARead-only
Query network traffic flows from the Firewalla MSP API (GET /v2/flows). Without a ts: qualifier the API covers the last 24 hours. Returns up to limit flows and a cursor for the next page, or groups with groupBy. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (optional, default: 200, API maximum: 500) | |
| query | No | Search query for flows. Supports region:US for geographic filtering, protocol:tcp, status:blocked, domain:*, category:social, etc. | |
| cursor | No | Pagination cursor from previous response | |
| sortBy | No | Sort flows (default: "ts:desc") | |
| groupBy | No | Fields to group by, comma-separated, e.g. "category", "domain", "device", "box" or "device,category". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for "device,category") and the rest are the group's summed connection count and bytes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation read-only and open-world; the description adds materially: default 24h lookback unless a ts: qualifier is used, paginated cursor behavior, grouping semantics, and environment scoping. No contradictions with annotations. It doesn't cover error/rate-limit behavior but that is not required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences with no filler; the endpoint and core behavior are front-loaded, and each sentence adds a distinct operational fact. Could be slightly shorter, but all content earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 optional params, no output schema, and read-only annotations, the description covers the main invocation concerns: time window, pagination, grouping, scoping. It still omits an explicit differentiation from search_flows and does not describe the shape of a returned flow record for non-grouped results, which leaves minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% description coverage, so the baseline is 3. The description goes beyond the schema by explaining the ts: qualifier's effect on the time window and by detailing groupBy's shape (key, count, download, upload, total) and peculiar key behavior for device vs device,category. That compensation justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb ('Query'), resource ('network traffic flows'), and the underlying API endpoint (GET /v2/flows), making the operation unambiguous. It does not explicitly distinguish itself from sibling search_flows, which also appears to target flow data, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to prefer this tool over search_flows, get_recent_flow_activity, or get_flow_insights. It states behavioral conditions (24h window, FIREWALLA_BOX_ID scoping) but not selection criteria or exclusions. This is the main clarity gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_flow_insightsARead-only
Get category-based flow analysis for a period: top content categories and their domains, top devices by bandwidth, and optionally blocked traffic. Ideal for answering questions like "what porn sites were accessed" or "what social media was used". Computed client-side from the period's largest flows (GET /v2/flows by total bytes: up to 500 for categories, 200 for devices) and, with include_blocked, the 50 most frequent blocked flows, so on a busy network it covers the largest flows, not all of them. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | Time period for analysis (default: 24h) | 24h |
| categories | No | Filter to specific content categories (optional) | |
| include_blocked | No | Include blocked traffic analysis (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and openWorldHint, and the description adds substantial behavioral detail: it is computed client-side from the largest flows, with explicit caps (500 categories, 200 devices, 50 blocked flows), and it explicitly warns that on a busy network it covers the largest flows, not all of them. It also discloses FIREWALLA_BOX_ID scoping. This goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. The main purpose is front-loaded, followed by usage examples and then necessary computation/scoping caveats. Every sentence contributes to correct tool selection and invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining what the tool returns, and it does: categories, domains, devices, and optional blocked traffic. It also covers limits, scoping, and the optional parameter behavior, making it complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining that include_blocked yields the 50 most frequent blocked flows and that categories map to top content categories and domains. It does not add much about the period parameter, but the schema already covers it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Get category-based flow analysis for a period,' and enumerates concrete outputs (top content categories and their domains, top devices by bandwidth, optionally blocked traffic). It also gives example questions ('what porn sites were accessed') that make the tool's niche unmistakable among flow-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by saying it is 'Ideal for answering questions like...' and by explaining the client-side computation scope. It does not explicitly name alternative tools or state when not to use it, but the context is strong enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_network_rulesARead-only
Retrieve firewall rules and conditions (GET /v2/rules). The API returns every matching rule; the tool returns the first limit. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | Maximum number of rules to return (required) | |
| query | No | Search conditions for filtering rules |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description discloses two important behaviors: the tool returns only the first 'limit' results even though the API returns all matching rules, and it is scoped to FIREWALLA_BOX_ID when set. This adds meaningful context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three compact sentences with no wasted words. It front-loads the primary action, then explains the truncation behavior and scoping, making every sentence valuable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read-only tool with 100% schema coverage and clear annotations, the description is nearly complete. It covers the endpoint, result truncation, and scoping, though it could be stronger by distinguishing itself from search_rules.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline of 3 applies. The description clarifies the limit parameter's effect by noting the tool returns only the first 'limit' results, but it does not add substantive meaning beyond the schema's descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retrieve') and identifies the resource ('firewall rules and conditions') with an endpoint reference (GET /v2/rules). It is clear, but it does not explicitly differentiate from the sibling search_rules tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus search_rules or get_network_rules_summary. The description provides behavioral context such as limiting results and scoping by FIREWALLA_BOX_ID, but it does not state exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_network_rules_summaryARead-only
Get overview counts of network rules by action, direction, status and target type (convenience wrapper): reads the rules from GET /v2/rules and counts them locally. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| rule_type | No | Filter by rule type | |
| active_only | No | Only include active rules in summary (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, and the description is consistent with them. It adds valuable behavior beyond annotations: the tool reads from GET /v2/rules and counts locally, and it reveals global scope when FIREWALLA_BOX_ID is unset. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, each earning its place: the first communicates purpose and mechanism, the second covers scoping behavior. There is no filler, repetition of schema details, or irrelevant context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only summary tool with zero required parameters, the description adequately states what is counted, the data source, and the box-scoping behavior. Since there is no output schema, describing the exact return shape in more detail would make it fully complete, but 'counts by action, direction, status and target type' conveys the essential result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both rule_type and active_only are already documented in the input schema. The description does not add extra meaning about accepted values, filtering behavior, or edge cases, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a specific verb-resource combination: 'Get overview counts of network rules by action, direction, status and target type.' It further clarifies the aggregate nature by calling itself a 'convenience wrapper' that reads GET /v2/rules and counts locally, making it clearly distinguishable from sibling tools like get_network_rules or search_rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'convenience wrapper' label and 'overview counts' wording imply this is for summary-level insights rather than raw rule listings, but the description never explicitly names alternatives or states when not to use it. The only explicit guidance is scoping behavior (FIREWALLA_BOX_ID vs. every box), which is execution context rather than tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_offline_devicesARead-only
List offline devices from the full device list (GET /v2/devices), most recently seen first by default, up to limit; total_offline_devices counts all of them. Scoped to box, else FIREWALLA_BOX_ID, else every box.
| Name | Required | Description | Default |
|---|---|---|---|
| box | No | Filter devices under a specific Firewalla box | |
| limit | No | Maximum number of offline devices to return | |
| sort_by_last_seen | No | Sort devices by last seen time (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=true; the description adds scoping precedence, default sorting, and total_offline_devices count behavior. These details go beyond annotations and help an agent anticipate response semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with zero filler; the action is front-loaded and every clause conveys scoping, ordering, limit, and count behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with 3 optional params and no output schema, it covers return semantics (total count, sort, limit) and scoping. Does not define 'offline' or pagination, but those are minor given the annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds box scoping precedence ('Scoped to box, else FIREWALLA_BOX_ID, else every box') beyond the schema's 'filter' wording, and confirms limit/sort defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List offline devices from the full device list'), names the exact endpoint (GET /v2/devices), and clarifies the filtering semantics. The scoping rules distinguish it from generic device listing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage context is present through sort order, limit, and scoping fallback, but there is no explicit when-to-use guidance or comparison with the sibling search_devices. No exclusions or alternatives are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recent_flow_activityARead-only
Get a snapshot of the 50 most recent network flows (one GET /v2/flows request) with protocol, region and blocked/allowed counts; the minutes they span depend on how busy the network is. Use this for: "what's happening right now?", current security threats, immediate network issues. DO NOT use for: historical analysis, more than 50 flows, or daily/weekly patterns; use search_flows with time queries like "ts:>24h" for those. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, and the description adds meaningful behavioral context: exactly one API request, variable time span based on network busyness, and FIREWALLA_BOX_ID scoping with fallback to every box. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three well-organized sentences: definition, use/don't-use guidance, and scope. Every sentence earns its place, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description adequately covers return content, limits, request behavior, and scope. An agent has all the information needed to select and invoke this tool accurately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so there is no parameter burden on the description. It still adds relevant invocation context by explaining the implicit FIREWALLA_BOX_ID environment scoping, which is directly useful for calling the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get a snapshot of the 50 most recent network flows (one GET /v2/flows request)'. It names the output dimensions (protocol, region, blocked/allowed counts), and its contrast with search_flows distinguishes it from the closest sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool ('what's happening right now?', current security threats, immediate network issues) and when not to use it (historical analysis, >50 flows, daily/weekly patterns), pointing to search_flows with a concrete query example 'ts:>24h'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rule_trendsARead-only
Rules created per day for the last 30 days, from GET /v2/trends/rules; period and group work as in get_alarm_trends. When that endpoint answers 400 (it did when measured), each UTC day counts the rules in GET /v2/rules created on it, scoped to FIREWALLA_BOX_ID when set (rules deleted since are not counted), and the response says so.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Get trends for a specific box group | |
| period | No | Return the days that overlap this period (default: 30d). The API has no finer resolution than a day | 30d |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations (readOnlyHint and openWorldHint). It discloses the fallback mechanism when the primary endpoint returns 400, the scoping behavior with FIREWALLA_BOX_ID, the exclusion of deleted rules, and that the response explicitly indicates the fallback. This level of behavioral detail is exceptional and helps an agent anticipate edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, complex sentence that leads with the core functionality, then details the fallback and scoping. It is information-dense without redundancy, though its length and nested clauses slightly reduce skimmability. Overall, it is concise given the amount of context it conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional parameters and no output schema, the description covers all necessary runtime behavior: the data source, the aggregation window, the fallback logic, scoping rules, and handling of deleted rules. It also notes when the response will indicate a fallback. Nothing an agent needs to invoke it correctly or interpret its behavior is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides complete descriptions for both parameters (group and period), including defaults and semantics (e.g., 'The API has no finer resolution than a day'). The description adds no additional parameter-specific meaning beyond referencing get_alarm_trends for semantics, which is implicitly covered. With 100% schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it returns rules created per day over the last 30 days, sourced from a specific endpoint. It also differentiates from the sibling get_alarm_trends by explicitly referencing how period and group semantics align. The verb+resource+metric are all concrete, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description references get_alarm_trends for how period and group work, which gives users a model for usage. It also explains when a fallback occurs (HTTP 400) and how it scopes by FIREWALLA_BOX_ID. However, it does not explicitly state when to prefer this tool over get_alarm_trends or other siblings, nor outline any exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_simple_statisticsARead-only
Get account-wide counts from GET /v2/stats/simple: online boxes, offline boxes, alarms and rules, optionally for one box group. Not limited by FIREWALLA_BOX_ID.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Get statistics for specific box group |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=true. The description reinforces and extends these by explaining the account-wide scope, the optional group parameter, and the fact that it is not constrained by FIREWALLA_BOX_ID. It also implies the return shape by listing the counted entity types.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler. The endpoint and data scope are front-loaded, the optional parameter is mentioned, and the scope caveat is placed last for emphasis.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only statistics tool with one optional parameter and no output schema, the description is complete: it explains what is counted, where the data comes from, and how scope is determined. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single optional 'group' parameter at 100%, so the baseline is 3. The description adds only 'optionally for one box group', which slightly clarifies usage but does not significantly extend beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Get account-wide counts from GET /v2/stats/simple' and enumerates the exact data categories (online boxes, offline boxes, alarms and rules). It also distinguishes this from box-scoped siblings by adding 'Not limited by FIREWALLA_BOX_ID'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly communicates when the tool is appropriate: account-wide counts rather than per-box statistics, with an optional box-group filter. It does not explicitly name alternative tools like get_statistics_by_box or get_statistics_by_region, so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_specific_alarmARead-only
Get detailed information for one Firewalla alarm (GET /v2/alarms/{gid}/{aid}). Alarm IDs are per box: pass gid on a multi-box account, or each box is checked, one request per box, until one has the alarm.
| Name | Required | Description | Default |
|---|---|---|---|
| gid | No | Box the alarm belongs to (the gid field of get_active_alarms or search_alarms results). Defaults to FIREWALLA_BOX_ID; without either, each box on the account is checked. | |
| alarm_id | Yes | Alarm ID (required for API call): the aid from get_active_alarms or search_alarms, as a number or a string |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses the multi-box scan behavior: without gid, 'each box is checked, one request per box, until one has the alarm.' This is a meaningful behavioral trait involving multiple HTTP requests and stop-at-first-match logic that annotations do not convey. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and endpoint, followed by the essential per-box ID caveat. No filler or redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read-only getter with full schema coverage and readOnlyHint, the description covers the endpoint and the only non-obvious behavior (multi-box lookup). It does not describe the return shape or not-found behavior, but there is no output schema and the operation is simple enough that this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the per-box semantics of alarm IDs, but the input schema already documents the gid fallback and the aid source, so it adds little beyond what structured data already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get detailed information for one Firewalla alarm' plus the REST endpoint (GET /v2/alarms/{gid}/{aid}). This clearly differentiates it from sibling list tools like get_active_alarms and search_alarms by targeting a single alarm.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when the tool is relevant and explains the gid/fallback behavior ('pass gid on a multi-box account, or each box is checked...'). It does not explicitly name alternatives or exclusions, but the schema notes that alarm_id comes from get_active_alarms or search_alarms, which is enough to route an agent from a list result to this details call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_specific_target_listARead-only
Retrieve one target list by ID, including its targets (GET /v2/target-lists/{id}).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Target list ID (required) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds that the tool returns the list 'including its targets', which is useful but not a behavioral disclosure beyond what the annotations imply. It does not mention any side effects, error conditions, or pagination, but for a read-only retrieval this is acceptable. It adds some value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise and front-loaded with the verb and resource. It includes the API endpoint for reference, which is helpful without being verbose. Every word earns its place, and there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple retrieval tool with one parameter, no output schema, and read-only annotations, the description is complete. It states what it does, what it returns, and the parameter is fully documented. There are no missing pieces an agent would need to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents the only parameter (id) with 100% coverage, so the baseline is 3. The description does not add any extra semantics about the parameter format (e.g., UUID vs. string) or constraints beyond what the schema states. It simply reiterates the purpose. No additional value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Retrieve'), a precise resource ('one target list by ID'), and clarifies the returned content ('including its targets'). It clearly differentiates from sibling tools like get_target_lists (plural) and search_target_lists by implying single-ID retrieval. This is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it: when you have a specific target list ID and want its details. It does not explicitly mention alternatives or when not to use it, but the context is clear enough that an agent can infer it is for single-list retrieval rather than listing or searching. No explicit exclusions are given, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statistics_by_boxARead-only
Top boxes by blocked flows (the default) or by Security Activity alarms, from GET /v2/stats/{type}, with each box's details from GET /v2/boxes; each box's value is the statistic, over about the last 30 days when measured. Not limited by FIREWALLA_BOX_ID.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Statistics type to retrieve | topBoxesByBlockedFlows |
| group | No | Get statistics for specific box group | |
| limit | No | Maximum number of results (optional, default: 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, lowering the burden on the description. The description usefully adds the approximate 30-day measurement window, the fact that results are not limited by FIREWALLA_BOX_ID, and the semantics that each box's value is the statistic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence where every clause earns its place: default vs. alternative statistic, source endpoints, value semantics, time window, and scope limitation. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with three optional parameters and no output schema, the description gives enough to invoke it correctly and interpret the result concept. It could be slightly more explicit about the exact response shape, but the endpoint references and value semantics largely compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents type, group, and limit. The description reinforces the default type and the two allowed enum values, but does not add meaningful parameter-level details beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific purpose: top boxes ranked by blocked flows or Security Activity alarms, sourced from a specific endpoint and enriched with box details. It also distinguishes this from other statistics tools by emphasizing per-box scope and the 'Not limited by FIREWALLA_BOX_ID' qualifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clarifies the default statistic type and that the tool is not limited by FIREWALLA_BOX_ID, giving some usage context. However, it does not explicitly say when to prefer this tool over siblings like get_statistics_by_region or get_simple_statistics, nor does it state any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statistics_by_regionARead-only
Top regions by blocked flows, from GET /v2/stats/topRegionsByBlockedFlows, optionally for one box group; the API returned no more than 5 regions. Not limited by FIREWALLA_BOX_ID.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Get statistics for specific box group | |
| limit | No | Maximum number of regions (optional, default: 5; the API returned no more than 5 when a larger limit was tried) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already signaling a safe read operation, the description adds useful behavioral detail: the API caps results at 5 regions even when a larger limit is requested, and the query is not constrained by FIREWALLA_BOX_ID. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence front-loads the result semantics and packs in endpoint, optional filter, API cap, and scope limitation without wasted words. Every clause adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only statistics tool with two optional parameters and no output schema, the description plus schema fully covers what an agent needs: what is returned, the optional group filter, the limit behavior, and the scope. The readOnly and openWorld annotations complement the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and both parameters (group, limit) are already described with meaningful detail including default and minimum. The description's mentions of optional box-group filtering and the 5-region cap reinforce the schema but do not add new parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly identifies the result as top regions ranked by blocked flows, cites the exact endpoint, and disambiguates scope with optional box group and the absence of a box-ID restriction. The noun-phrase style is still unambiguous because the endpoint and resource are explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this tool for regional blocked-flow statistics, optionally filtered by one box group, and it explicitly notes the query is not limited by FIREWALLA_BOX_ID. It does not name sibling alternatives like get_statistics_by_box, but the scope guidance is strong enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_target_listsARead-only
Retrieve target lists (GET /v2/target-lists). Without owner the API returns the MSP's global lists and the Firewalla-managed lists; owner selects global lists, a box's lists, or several. entry_count is the number of entries in each list; the API does not return the entries of Firewalla-managed lists, so their targets is null. Returns up to limit lists.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | Yes | Maximum number of target lists to return (required) | |
| owner | No | Only lists with this owner: "global" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. "global,<box_gid>". Default: global and Firewalla-managed lists. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint and openWorldHint. The description adds valuable context about the response: entry_count is provided, and Firewalla-managed lists have null targets because entries are not returned. It also notes the limit behavior. This complements the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the endpoint and core behavior. It avoids unnecessary filler and includes important details about response fields and limits. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list-retrieval tool with no output schema, the description covers the key aspects: owner selection, default scope, entry_count, null targets for Firewalla-managed lists, and the limit parameter. Nothing essential is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are well-documented in the schema itself. The description adds minimal extra value: it repeats the default owner behavior and mentions the limit behavior. Since the schema already explains owner options and default, the description does not significantly enhance parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves target lists, specifies the endpoint, and explains the behavior with and without the owner parameter. It distinguishes the scope (global, box, or multiple owners) and the default behavior, making it unambiguous what this tool does relative to siblings like get_specific_target_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: to fetch target lists with optional owner filtering, and describes default behavior. It does not explicitly mention alternatives or when not to use it, but the context is clear enough for an agent to select it for list retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_ruleAIdempotent
Pause an active firewall rule on the box until resume_rule reactivates it (POST /v2/rules/{id}/pause, no body). The MSP API takes no duration, so the pause does not expire on its own. Checks the rule's status first and changes nothing if it is already paused.
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | Yes | Rule ID to pause |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already include idempotentHint=true and destructiveHint=false, but the description adds significant context: the API takes no duration and the pause does not expire on its own, and it checks the rule's status first, changing nothing if already paused. This explains the idempotent behavior in detail and provides endpoint information, exceeding what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all substantive: action, endpoint, no-duration behavior, and idempotency check. No filler or redundancy, and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with annotations covering safety and idempotency, the description includes the endpoint, no-duration constraint, and status check. It provides everything an agent needs to call it correctly, including when it is safe to invoke.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the description does not add any meaning to the rule_id parameter beyond the schema's 'Rule ID to pause'. No additional constraints, format hints, or examples are provided, so it relies on the schema as expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Pause') and resource ('active firewall rule'), and explicitly references 'resume_rule' as the counterpart, distinguishing it from sibling tools like search_rules. It is unambiguous about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions that the pause lasts until resume_rule reactivates it, which implies when to use the reverse tool, and notes that it checks status and does nothing if already paused, indicating safe calling conditions. However, it does not explicitly list alternatives or exclusions beyond resume_rule, so it's clear but not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_ruleAIdempotent
Resume a paused firewall rule on the box, restoring it to active (POST /v2/rules/{id}/resume, no body). Checks the rule's status first and changes nothing if it is already active.
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | Yes | Rule ID to resume |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses the HTTP method and body requirement, and explains the status-checking behavior that makes the operation effectively idempotent. This adds meaningful behavioral context that the annotations alone do not provide, such as 'changes nothing if it is already active.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences deliver the action, resource, endpoint, request body requirement, and idempotency behavior with no filler. The key purpose is front-loaded, and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description is complete: it states what happens, how it happens, and the edge case of an already-active rule. Annotations cover the mutation and idempotency profile, so nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter rule_id is already described as 'Rule ID to resume.' The description adds the path template showing where the ID goes, but does not add new semantic detail about the parameter's format, constraints, or behavior. Baseline 3 is appropriate given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Resume'), a specific resource ('paused firewall rule'), and the intended outcome ('restoring it to active'). It also names the exact endpoint (POST /v2/rules/{id}/resume), which unambiguously distinguishes it from sibling tools like pause_rule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when a paused firewall rule should be restored to active. It also provides useful guidance that the tool checks status first and is a no-op if already active, so callers need not pre-check state. It does not explicitly name alternatives, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_alarmsBRead-only
Search alarms using full-text or field filters. Alarm types: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Reads GET /v2/alarms, 500 per request, following the cursor up to limit. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (optional, default: 200, API maximum: 500) | |
| query | Yes | Search query using Firewalla syntax. Supported fields: type:1-16 (see alarm types above), status:1/2 (active/archived), device.ip:192.168.*, region:US (country code), box.id:box_gid, device.name:*. Examples: "type:8 AND region:US" (video from US), "type:10 AND status:1" (active porn alerts), "device.ip:192.168.* AND status:1" (active alarms from the LAN), "porn" (free text: a term without a qualifier searches alarm text) | |
| cursor | No | Pagination cursor from previous response | |
| sortBy | No | Sort alarms (default: ts:desc) | |
| groupBy | No | Fields to group by, comma-separated, e.g. "type", "status", "device" or "type,box". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint. The description adds valuable behavior: it mentions the HTTP endpoint (GET /v2/alarms), pagination limit of 500 per request, cursor following up to the limit, and scoping to a box ID. This goes beyond the annotation flags and helps the agent anticipate API behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph that lists all 16 alarm types, making it somewhat long but logically structured. It front-loads the core purpose, then provides the type mapping and behavioral details. While it could be more concise (e.g., moving the type list to a reference section), it is organized and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema, the description covers input semantics, pagination, and scoping but does not describe the response format (e.g., what fields are returned, how groups look). This is a gap because agents need to know how to parse results. Given the complexity of 5 parameters, this omission prevents full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all parameters have descriptions. The description adds the alarm type mapping (1-16) referenced by the query parameter, which is essential for constructing valid queries. It also explains pagination behavior with cursor and limit, adding meaning beyond the schema's generic descriptions. This is a meaningful supplement to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Search alarms using full-text or field filters.' It specifies the resource (alarms) and the action (search), and even lists alarm types, which helps understanding. However, it does not explicitly differentiate from sibling tools like get_active_alarms or get_specific_alarm, so it falls short of the 5-level criterion that requires distinguishing from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention conditions for choosing this over get_active_alarms or get_specific_alarm, nor does it state any exclusions or prerequisites. The only context given is scoping to FIREWALLA_BOX_ID, which is a parameter detail, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_devicesARead-only
Search devices by name, IP, MAC or status (convenience wrapper with client-side filtering): reads the device list from GET /v2/devices (box, else FIREWALLA_BOX_ID, else every box) and filters it locally.
| Name | Required | Description | Default |
|---|---|---|---|
| box | No | Filter devices under a specific Firewalla box | |
| limit | No | Maximum number of devices to return | |
| query | Yes | Search query using Firewalla syntax. Supported fields: mac:AA:BB:CC:DD:EE:FF, ip:192.168.1.*, name:*iPhone*, online:true/false, mac_vendor:Apple, gid:box_gid, network.name:*, group.name:*. Examples: "online:false AND mac_vendor:Apple", "ip:192.168.1.* AND name:*laptop*", "mac:AA:* OR name:*phone*" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only and open-world. The description adds valuable behavioral context beyond annotations: it fetches the complete device list and filters client-side, and describes the box/FIREWALLA_BOX_ID/every-box precedence. This helps the agent understand the tool's scope and performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the purpose and then packs in the essential implementation detail. Every phrase earns its place, and the key behavior (local filtering) is highlighted early.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with no output schema, the description covers the data source, filtering approach, and scope selection logic. It doesn't describe the return shape, but that is partially inferred from the query schema and the tool's search nature. The main minor gap is the lack of an explicit return-value description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining the box parameter's fallback behavior (box, else FIREWALLA_BOX_ID, else every box), which is not present in the schema. Limit and query are well-documented in the schema itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search devices by name, IP, MAC or status'. It clearly identifies the tool as a convenience wrapper and differentiates from sibling search tools by naming the device-list endpoint it wraps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: it reads the full device list from GET /v2/devices and filters locally, with explicit fallback logic for box selection. It doesn't explicitly name alternatives, but no sibling tool searches devices, so the usage context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_flowsARead-only
Search network flows with advanced query filters. Use this for: historical analysis, specific time ranges, complex filtering, or when you need more than 50 flows. Supports pagination, time-based queries (e.g., "ts:>1h" for the last hour, or Unix seconds such as "ts:1735689600-1735693200"), and all flow fields including geographic filtering. For quick "what's happening now" snapshots, use get_recent_flow_activity instead. Reads GET /v2/flows, 500 per request, following the cursor up to limit. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum results (optional, default: 200, API maximum: 500) | |
| query | Yes | Search query using Firewalla syntax. Supported fields: protocol:tcp/udp, direction:inbound/outbound/local, status:blocked/ok, total:>1MB (download + upload in B/KB/MB/GB/TB), download:>10MB, upload:>10MB, domain:*.example.com, region:US (country code), category:social/games/porn/etc, box.id:box_gid, device.ip:192.168.*, source.ip:*, destination.ip:*, ts:>1h. Examples: "region:US AND protocol:tcp", "status:blocked AND region:CN", "category:social OR category:games" | |
| cursor | No | Pagination cursor from previous response | |
| sortBy | No | Sort flows (default: "ts:desc") | |
| groupBy | No | Fields to group by, comma-separated, e.g. "category", "domain", "device", "box" or "device,category". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for "device,category") and the rest are the group's summed connection count and bytes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and openWorldHint, but the description adds substantive behavioral details: it reads GET /v2/flows, caps at 500 per request, follows the cursor up to limit, and scopes results to FIREWALLA_BOX_ID when set. These clarify pagination, rate limits, and scoping without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: purpose first, then usage conditions, then capability highlights, then a pointer to an alternative, then technical specifics. Every sentence serves a clear function with no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex search tool with no output schema, the description covers everything needed to invoke it correctly: purpose, when to use, pagination behavior, scoping, and even the endpoint. Grouping behavior is conveyed via the schema's groupBy description, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds marginal value by illustrating time-based query syntax ('ts:>1h', Unix seconds) and explaining how cursor and limit interact ('following the cursor up to limit'). These details are not fully captured in the schema's parameter descriptions, so a slight premium is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Search network flows with advanced query filters', giving a specific verb+resource. It then enumerates concrete uses (historical analysis, specific time ranges, complex filtering, >50 flows) that sharply differentiate it from the sibling get_recent_flow_activity, making the tool's role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('Use this for: historical analysis, specific time ranges, complex filtering, or when you need more than 50 flows') and when not to, naming the alternative ('For quick "what's happening now" snapshots, use get_recent_flow_activity instead'). This direct routing leaves no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_rulesARead-only
Search firewall rules by target, action or status; the MSP API applies the query (GET /v2/rules). Supports all rule fields. Scoped to FIREWALLA_BOX_ID when set, otherwise every box.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of rules to return | |
| query | Yes | Search query using Firewalla syntax. Supported fields: action:allow/block/timelimit, target.type:domain/ip/device, target.value:*.facebook.com, status:active/paused, direction:bidirection/inbound/outbound, protocol:tcp/udp, box.id:box_gid, scope.type:device/network, notes:"description text". Examples: "action:block AND target.value:*.social.com", "status:paused", "target.type:domain AND action:block" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is known. The description adds behavioral value beyond that: it identifies the API endpoint (GET /v2/rules), notes that all rule fields are supported, and explains the FIREWALLA_BOX_ID scoping behavior that affects which rules are returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the core purpose appears in the first clause, followed by only necessary scope and capability details. There is no redundant filler, and each sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with a rich schema, the description adequately covers the query scope and the box-scoping behavior. It does not describe the response payload, and there is no output schema to fill that gap, but the return of firewall rules is strongly implied; a brief note on result format would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% parameter coverage, including a detailed query syntax and examples for both query and limit. The description's reference to 'target, action or status' is only a light restatement of the schema, not additional semantic meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Search firewall rules by target, action or status', which names a specific verb, resource, and query dimensions. Adding 'Supports all rule fields' clarifies the scope and distinguishes it from simple list/get siblings like get_network_rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides useful context, especially 'Scoped to FIREWALLA_BOX_ID when set, otherwise every box', and implies query-based use. However, it does not explicitly state when to prefer search_rules over alternatives such as get_network_rules or search_flows, nor does it give when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_target_listsARead-only
Search target lists (convenience wrapper with client-side filtering): reads GET /v2/target-lists, sending owner if given (without it, the global and Firewalla-managed lists), and applies the query locally.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of target lists to return | |
| owner | No | Only lists with this owner, sent to the API: "global" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. "global,<box_gid>". Default: global and Firewalla-managed lists. | |
| query | Yes | Search query for target lists. Supported fields: name:*Social*, owner:global/box_gid, category:social/games/ad/porn/etc, targets:*.facebook.com, notes:"description text", target_count:>100 (entries: n, >n, >=n, <n, <=n or a range 10-50), last_updated:>2026-09-01 (a date or Unix seconds, with the same comparisons). Examples: "category:social", "owner:global AND name:*Block*", "targets:*.gaming.com", "target_count:>1000", "last_updated:<2026-01-01" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds meaningful behavior: the query is not sent to the server; filtering happens locally after reading the endpoint, and owner defaults to global and Firewalla-managed lists when omitted. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that is front-loaded with the tool's purpose and contains no filler. Every clause contributes behavioral or scoping information, making it easy to extract the key facts quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with full schema coverage, the description covers the endpoint, owner default, and local-filtering behavior. It does not describe the return shape, but no output schema exists and the resource name makes the return type predictable; the absence of explicit sibling routing is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds one useful semantic beyond the schema: the query is applied client-side while owner is passed through to the API. It does not add detail about limit or query syntax, but the schema already documents those thoroughly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly identifies the action ('Search') and resource ('target lists') and specifies that it is a convenience wrapper that reads GET /v2/target-lists and applies the query locally. This differentiates it from a direct list call or single-list retrieval without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is a wrapper over GET /v2/target-lists for client-side search/filtering, and the owner parameter controls scope. It does not explicitly name sibling alternatives such as get_target_lists or get_specific_target_list, nor state when not to use it, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_target_listADestructiveIdempotent
Update an existing target list (PATCH /v2/target-lists/{id}). Only the fields given are sent; targets, when given, is the complete new list and is not merged with the current targets.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Target list ID (required) | |
| name | No | Updated target list name (max 24 chars) | |
| notes | No | Updated description | |
| targets | No | Updated array of domains, IPs, or CIDR ranges | |
| category | No | Updated content category |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive and idempotent, and the description adds precise behavioral context: only supplied fields are sent, and targets is a complete replacement, not a merge. This meaningfully clarifies what gets overwritten and prevents an incorrect mental model of additive target updates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The endpoint and operation are front-loaded, and the second sentence delivers the most important behavioral nuance without unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with destructive potential and no output schema, the description covers the critical invocation semantics: PATCH behavior, partial field sending, and full target-list replacement. It does not mention the response body or failure cases, but these are not essential for correctly calling the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying the targets parameter's full-replacement semantics and the general partial-update behavior, which are not visible from parameter descriptions alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Update an existing target list (PATCH /v2/target-lists/{id})'. It conveys this is a modification operation on a single existing list, but it does not explicitly contrast it with create_target_list, delete_target_list, or get_specific_target_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong usage guidance by explaining PATCH semantics: 'Only the fields given are sent' and that targets, when provided, replaces the full list rather than merging. It does not explicitly state when to prefer this over create/delete/search alternatives, but it provides clear context for correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v1.5.0- Changed
get_active_alarms2 fields changed- changed
Input schema / properties / groupBy / descriptionPrevious value: -"Group alarms by field (e.g., type, box)"New value: +"Fields to group by, comma-separated, e.g. \"type\", \"status\", \"device\" or \"type,box\". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query for filtering alarms (default: status:1 for active). Use type:N where N is: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Examples: type:8 (video), type:10 (porn), region:US, source_ip:*"New value: +"Search query for filtering alarms. Active alarms only (status:1) unless the query names a status. Use type:N where N is: 1=Security Activity, 2=Abnormal Upload, 3=Large Bandwidth Usage, 4=Monthly Data Plan, 5=New Device, 6=Device Back Online, 7=Device Offline, 8=Video Activity, 9=Gaming Activity, 10=Porn Activity, 11=VPN Activity, 12=VPN Connection Restored, 13=VPN Connection Error, 14=Open Port, 15=Internet Connectivity Update, 16=Large Upload. Examples: type:8 (video), type:10 (porn), region:US, device.ip:192.168.*"
- Changed
get_alarm_trends1 field changed- added
Input schema / properties / periodAdded value: +{ + "default": "30d", + "description": "Return the days that overlap this period (default: 30d). The API has no finer resolution than a day, so 1h returns today so far and 24h returns yesterday and today", + "enum": [ + "1h", + "24h", + "7d", + "30d" + ], + "type": "string" +}
- Changed
get_bandwidth_usage1 field changed- changed
Input schema / properties / box / descriptionPrevious value: -"Filter devices under a specific Firewalla box"New value: +"Only flows of this box (box gid). Defaults to FIREWALLA_BOX_ID; without either, every box."
- Changed
get_flow_data2 fields changed- changed
Input schema / properties / groupBy / descriptionPrevious value: -"Group flows by specified values (e.g., \"domain,box\")"New value: +"Fields to group by, comma-separated, e.g. \"category\", \"domain\", \"device\", \"box\" or \"device,category\". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for \"device,category\") and the rest are the group's summed connection count and bytes." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query for flows. Supports region:US for geographic filtering, protocol:tcp, blocked:true, domain:*, category:social, etc."New value: +"Search query for flows. Supports region:US for geographic filtering, protocol:tcp, status:blocked, domain:*, category:social, etc."
- Changed
get_rule_trends1 field changed- added
Input schema / properties / periodAdded value: +{ + "default": "30d", + "description": "Return the days that overlap this period (default: 30d). The API has no finer resolution than a day", + "enum": [ + "1h", + "24h", + "7d", + "30d" + ], + "type": "string" +}
- Changed
get_specific_alarm3 fields changed- changed
Input schema / properties / alarm_id / descriptionPrevious value: -"Alarm ID (required for API call)"New value: +"Alarm ID (required for API call): the aid from get_active_alarms or search_alarms, as a number or a string" - changed
Input schema / properties / alarm_id / typePrevious value: -"string"New value: +[ + "string", + "number" +] - added
Input schema / properties / gidAdded value: +{ + "description": "Box the alarm belongs to (the gid field of get_active_alarms or search_alarms results). Defaults to FIREWALLA_BOX_ID; without either, each box on the account is checked.", + "type": "string" +}
- Changed
get_statistics_by_region1 field changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum number of results (optional, default: 5)"New value: +"Maximum number of regions (optional, default: 5; the API returned no more than 5 when a larger limit was tried)"
- Changed
get_target_lists1 field changed- added
Input schema / properties / ownerAdded value: +{ + "description": "Only lists with this owner: \"global\" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. \"global,<box_gid>\". Default: global and Firewalla-managed lists.", + "type": "string" +}
- Changed
pause_rule3 fields changed- removed
Input schema / properties / boxRemoved value: -{ - "description": "Box GID for context (required by API)", - "type": "string" -} - removed
Input schema / properties / durationRemoved value: -{ - "default": 60, - "description": "Duration in minutes to pause the rule (optional, default: 60, range: 1-1440)", - "maximum": 1440, - "minimum": 1, - "type": "number" -} - changed
Input schema / requiredPrevious value: -[ - "rule_id", - "box" -]New value: +[ + "rule_id" +]
- Changed
resume_rule2 fields changed- removed
Input schema / properties / boxRemoved value: -{ - "description": "Box GID for context (required by API)", - "type": "string" -} - changed
Input schema / requiredPrevious value: -[ - "rule_id", - "box" -]New value: +[ + "rule_id" +]
- Changed
search_alarms2 fields changed- changed
Input schema / properties / groupBy / descriptionPrevious value: -"Group alarms by specified fields (comma-separated)"New value: +"Fields to group by, comma-separated, e.g. \"type\", \"status\", \"device\" or \"type,box\". The API then returns groups instead of alarms: groups of { key, count }, where key holds the group fields (gid for box, device.id for device) and count the alarms in the group." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query using Firewalla syntax. Supported fields: type:1-16 (see alarm types above), resolved:true/false, status:1/2 (active/archived), source_ip:192.168.*, region:US (country code), gid:box_id, device.name:*, message:\"text search\". Examples: \"type:8 AND region:US\" (video from US), \"type:10 AND status:1\" (active porn alerts), \"source_ip:192.168.* AND NOT resolved:true\""New value: +"Search query using Firewalla syntax. Supported fields: type:1-16 (see alarm types above), status:1/2 (active/archived), device.ip:192.168.*, region:US (country code), box.id:box_gid, device.name:*. Examples: \"type:8 AND region:US\" (video from US), \"type:10 AND status:1\" (active porn alerts), \"device.ip:192.168.* AND status:1\" (active alarms from the LAN), \"porn\" (free text: a term without a qualifier searches alarm text)"
- Changed
search_devices2 fields changed- changed
Input schema / properties / query / descriptionPrevious value: -"Search query using Firewalla syntax. Supported fields: mac:AA:BB:CC:DD:EE:FF, ip:192.168.1.*, name:*iPhone*, online:true/false, vendor:Apple, gid:box_id, network.name:*, group.name:*. Examples: \"online:false AND vendor:Apple\", \"ip:192.168.1.* AND name:*laptop*\", \"mac:AA:* OR name:*phone*\""New value: +"Search query using Firewalla syntax. Supported fields: mac:AA:BB:CC:DD:EE:FF, ip:192.168.1.*, name:*iPhone*, online:true/false, mac_vendor:Apple, gid:box_gid, network.name:*, group.name:*. Examples: \"online:false AND mac_vendor:Apple\", \"ip:192.168.1.* AND name:*laptop*\", \"mac:AA:* OR name:*phone*\"" - removed
Input schema / properties / statusRemoved value: -{ - "default": "any", - "description": "Filter by online status", - "enum": [ - "online", - "offline", - "any" - ], - "type": "string" -}
- Changed
search_flows2 fields changed- changed
Input schema / properties / groupBy / descriptionPrevious value: -"Group flows by specified values (e.g., \"domain,box\")"New value: +"Fields to group by, comma-separated, e.g. \"category\", \"domain\", \"device\", \"box\" or \"device,category\". The API then returns groups instead of flows: groups of { key, count, download, upload, total }, where key holds the group fields (gid for box; for device the device with its name, but only its id for \"device,category\") and the rest are the group's summed connection count and bytes." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query using Firewalla syntax. Supported fields: protocol:tcp/udp, direction:inbound/outbound/local, blocked:true/false, bytes:>1MB, domain:*.example.com, region:US (country code), category:social/games/porn/etc, gid:box_id, device.ip:192.168.*, source_ip:*, destination_ip:*. Examples: \"region:US AND protocol:tcp\", \"blocked:true AND bytes:>1MB\", \"category:social OR category:games\""New value: +"Search query using Firewalla syntax. Supported fields: protocol:tcp/udp, direction:inbound/outbound/local, status:blocked/ok, total:>1MB (download + upload in B/KB/MB/GB/TB), download:>10MB, upload:>10MB, domain:*.example.com, region:US (country code), category:social/games/porn/etc, box.id:box_gid, device.ip:192.168.*, source.ip:*, destination.ip:*, ts:>1h. Examples: \"region:US AND protocol:tcp\", \"status:blocked AND region:CN\", \"category:social OR category:games\""
- Changed
search_rules1 field changed- changed
Input schema / properties / query / descriptionPrevious value: -"Search query using Firewalla syntax. Supported fields: action:allow/block/timelimit, target.type:domain/ip/device, target.value:*.facebook.com, status:active/paused, direction:bidirection/inbound/outbound, protocol:tcp/udp, gid:box_id, scope.type:device/network, notes:\"description text\". Examples: \"action:block AND target.value:*.social.com\", \"status:paused\", \"target.type:domain AND action:block\""New value: +"Search query using Firewalla syntax. Supported fields: action:allow/block/timelimit, target.type:domain/ip/device, target.value:*.facebook.com, status:active/paused, direction:bidirection/inbound/outbound, protocol:tcp/udp, box.id:box_gid, scope.type:device/network, notes:\"description text\". Examples: \"action:block AND target.value:*.social.com\", \"status:paused\", \"target.type:domain AND action:block\""
- Changed
search_target_lists3 fields changed- removed
Input schema / properties / categoryRemoved value: -{ - "description": "Filter by category", - "type": "string" -} - changed
Input schema / properties / owner / descriptionPrevious value: -"Filter by owner (global or box gid)"New value: +"Only lists with this owner, sent to the API: \"global\" (MSP lists), a box gid (that box's lists), or several comma-separated, e.g. \"global,<box_gid>\". Default: global and Firewalla-managed lists." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query for target lists. Supported fields: name:*Social*, owner:global/box_gid, category:social/games/ad/porn/etc, targets:*.facebook.com, notes:\"description text\". Examples: \"category:social\", \"owner:global AND name:*Block*\", \"targets:*.gaming.com\""New value: +"Search query for target lists. Supported fields: name:*Social*, owner:global/box_gid, category:social/games/ad/porn/etc, targets:*.facebook.com, notes:\"description text\", target_count:>100 (entries: n, >n, >=n, <n, <=n or a range 10-50), last_updated:>2026-09-01 (a date or Unix seconds, with the same comparisons). Examples: \"category:social\", \"owner:global AND name:*Block*\", \"targets:*.gaming.com\", \"target_count:>1000\", \"last_updated:<2026-01-01\""
5 tool updates
v1.3.0- Changed
search_alarms1 field changed- changed
Input schema / requiredPrevious value: -[]New value: +[ + "query" +]
- Changed
search_devices1 field changed- changed
Input schema / requiredPrevious value: -[]New value: +[ + "query" +]
- Changed
search_flows1 field changed- changed
Input schema / requiredPrevious value: -[]New value: +[ + "query" +]
- Changed
search_rules2 fields changed- added
Input schema / properties / limitAdded value: +{ + "description": "Maximum number of rules to return", + "type": "number" +} - changed
Input schema / requiredPrevious value: -[]New value: +[ + "query" +]
- Changed
search_target_lists1 field changed- changed
Input schema / requiredPrevious value: -[]New value: +[ + "query" +]
28 tool updates
- First observed
create_target_list - First observed
delete_target_list - First observed
get_active_alarms - First observed
get_alarm_trends - First observed
get_bandwidth_usage - First observed
get_boxes - First observed
get_device_status - First observed
get_flow_data - First observed
get_flow_insights - First observed
get_network_rules - First observed
get_network_rules_summary - First observed
get_offline_devices - First observed
get_recent_flow_activity - First observed
get_rule_trends - First observed
get_simple_statistics - First observed
get_specific_alarm - First observed
get_specific_target_list - First observed
get_statistics_by_box - First observed
get_statistics_by_region - First observed
get_target_lists - First observed
pause_rule - First observed
resume_rule - First observed
search_alarms - First observed
search_devices - First observed
search_flows - First observed
search_rules - First observed
search_target_lists - First observed
update_target_list
TDQS
Scored across 28 tools
Several tools overlap in purpose: get_offline_devices and get_device_status both read device lists and report online/offline status; get_flow_data, search_flows, and get_recent_flow_activity all query flows with different scopes; get_network_rules and search_rules both retrieve rules. Descriptions help distinguish them, but the boundaries are not always obvious.
Most tools follow a consistent get_/search_/create_/update_/delete_/pause_/resume_ verb pattern with clear noun objects. Minor deviations like get_network_rules_summary and get_specific_target_list vs get_specific_alarm are acceptable, but the mix of get_ and search_ for similar resources (flows, alarms, rules) creates slight inconsistency.
28 tools is on the heavy side for a single server, though the domain (Firewalla MSP API) is broad with devices, alarms, flows, rules, target lists, and statistics. The count is borderline: many tools are convenience wrappers or variations on the same resource, which could be consolidated.
The tool set covers the main Firewalla MSP resources well: devices, alarms, flows, rules, target lists, and statistics. Minor gaps exist (e.g., no create/update/delete for rules or alarms, no device management actions), but the core read and search operations are comprehensive.
Maintenance
Related MCP Connectors
MCP Server for agents to onboard, pay, and provision services autonomously with InFlow
Crypto transaction firewall and risk tools for MCP agents.
A paid remote MCP for ClawManager, built to return verdicts, receipts, usage logs, and audit-ready J
Unified MCP Server is a remote MCP connector for AI agents and vertical AI products that provides access to 22,000+ authorized SaaS tools across 400+ integrations and 24 categories directly inside LLMs (Claude, GPT, Gemini, Cohere). Tools operate only on explicitly authorized customer connections, enabling agents to safely read and write against live third-party systems.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceProvides real-time access to Firewalla firewall data through 28 specialized tools for network monitoring, security analysis, bandwidth tracking, and firewall rule management. Enables users to query security alerts, analyze network flows, monitor device status, and manage firewall configurations through natural language.516 npm1-
- AlicenseNot gradedqualityNot gradedmaintenanceEnables interaction with Firewalla network security devices for network monitoring, device management, traffic analysis, and security rule configuration through MCP tools.-
- AlicenseAqualityBmaintenanceA secure MCP server for managing OPNsense firewalls through AI assistants. Provides 81 tools across system, firewall, network, DNS, DHCP, VPN, HAProxy, services, diagnostics, and security domains.81232 PyPI21MIT
- AlicenseNot gradedqualityDmaintenanceProvides AI assistants and MCP clients with API access to manage Firewalla MSP resources including boxes, devices, alarms, rules, flows, and target lists.5 npmGPL 3.0