Skip to main content
Glama
antonio-mello-ai

io.github.antonio-mello-ai/mcp-pfsense

mcp-pfsense

PyPI Python License: MIT

MCP server for managing pfSense firewalls through AI assistants like Claude, ChatGPT, and Copilot.

Requires: pfrest package installed on your pfSense instance (provides the REST API).

Features

20 tools across 7 categories:

Category

Tools

Description

System

get_system_status, get_interfaces

Version, CPU, memory, uptime, temperature, network interfaces

Firewall

list_firewall_rules, add_firewall_rule, delete_firewall_rule, list_firewall_aliases

Rule management with interface filtering, alias listing

DHCP

list_dhcp_leases, list_dhcp_static_mappings, add_dhcp_static_mapping, delete_dhcp_static_mapping

Active leases, IP reservations

DNS

list_dns_host_overrides, add_dns_host_override, delete_dns_host_override

Unbound DNS Resolver host overrides

Pending changes

get_pending_changes, apply_changes

See what is staged per subsystem (firewall, dhcp, dns) and apply it

Monitoring

get_gateway_status, get_arp_table, list_services, get_firewall_logs

Gateway health, connected devices, service status, recent raw firewall log entries

Services

restart_service

Restart any pfSense service

Safety

  • Two-step confirmation for destructive operations (delete rules, delete mappings, restart services, apply changes): the tool returns a warning on first call and only executes when called again with confirm=true.

  • Writes are staged, not live. Like the pfSense WebGUI, add_* and delete_* store the change in the config but do not activate it. The tool response says so (applied: false, plus a pending note). Activate with apply_changes(subsystem, confirm=true) — which reloads that subsystem, including anything a human left staged in the WebGUI — or pass apply=true on the write itself when you explicitly want a one-shot change. Nothing the assistant does reaches the packet filter without one of those two explicit steps.

  • delete_dhcp_static_mapping takes the mapping's interface (its parent_id in list_dhcp_static_mappings) and mapping_id; a mapping is addressed by both.

Related MCP server: io.github.abl030/pfsense-mcp

Installation

# Using uvx (recommended)
uvx mcp-pfsense

# Using pip
pip install mcp-pfsense

Prerequisites

  1. pfSense with pfrest package installed

  2. A user account with API access (typically admin)

Configuration

Set environment variables:

Variable

Required

Default

Description

PFSENSE_HOST

Yes

—

pfSense hostname or IP

PFSENSE_PASSWORD

Yes

—

API user password

PFSENSE_USERNAME

No

admin

API username

PFSENSE_PORT

No

443

API port

PFSENSE_SCHEME

No

https

http or https

PFSENSE_VERIFY_SSL

No

false

Verify SSL certificate

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "pfsense": {
      "command": "uvx",
      "args": ["mcp-pfsense"],
      "env": {
        "PFSENSE_HOST": "10.10.10.1",
        "PFSENSE_PASSWORD": "your-password"
      }
    }
  }
}

Claude Code

claude mcp add pfsense -- uvx mcp-pfsense

Then set environment variables in your shell or .env file.

Usage Examples

Once connected, ask your AI assistant:

  • "What's the pfSense system status?"

  • "Show me all firewall rules on the LAN interface"

  • "List active DHCP leases"

  • "Add a DNS entry for nas.home.lan pointing to 10.10.10.50"

  • "What devices are connected to the network?" (ARP table)

  • "Show gateway health and latency"

  • "Show the 50 most recent firewall log entries" (read-only get_firewall_logs)

  • "Create a firewall rule to allow TCP port 8080 on LAN"

  • "Reserve IP 10.10.10.60 for MAC aa:bb:cc:dd:ee:20"

API Compatibility

  • pfSense: 2.7.x and 2.8.x

  • pfrest: REST API v2 — any v2.x release, except list_dhcp_static_mappings, which needs v2.7.0 or later (it uses the /services/dhcp_server/static_mappings collection endpoint added in that release).

  • Python: 3.11+

The endpoint, parameters and encoding each tool uses are pinned by tests/test_client_endpoints.py and tests/test_wire_format.py, derived from the pfrest v2 endpoint definitions. Versions before 0.2.0 called several endpoints that do not exist in pfrest v2 (see Troubleshooting).

Note: pfrest runs on nginx (port 80 by default), separate from the pfSense WebGUI (lighttpd on port 443). If your pfrest is configured on a non-standard port, set PFSENSE_PORT and PFSENSE_SCHEME accordingly.

Troubleshooting

Only get_system_status and get_arp_table work; everything else returns 400/404

mcp-pfsense 0.1.1 and earlier called singular endpoints for listing (/interface, /firewall/rule, /firewall/alias) and legacy paths that pfrest v2 does not serve (/status/dhcp_leases, /services/dhcpd/static_mapping, /services/unbound/host_override, /status/gateway, /status/service for GET). Upgrade to 0.2.0 or later.

403 on list_services or other reads

pfrest checks the privileges of the API user per endpoint. Grant the user the api-v2-* privileges for the endpoints you need (or page-all for full access) under System → User Manager.

ModuleNotFoundError: No module named 'mcp.server.fastmcp'

The MCP Python SDK 2.0 removed the module that mcp-pfsense 0.1.1 and earlier import, so fresh installs (uvx mcp-pfsense, pip install) failed on startup. Upgrade to 0.2.0 or later, which pins mcp<2. If you must stay on an older mcp-pfsense: uvx --with "mcp<2" mcp-pfsense.

A rule / mapping / override was created but is not in effect

That is the default: writes are staged (see Safety). Check with get_pending_changes(subsystem) and activate with apply_changes(subsystem, confirm=true), or in the WebGUI. If a write returns 200 but nothing is stored at all, the pfrest read_only setting is on (System → REST API → Settings).

Development

git clone https://github.com/antonio-mello-ai/mcp-pfsense.git
cd mcp-pfsense
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

# Run tests
pytest

# Lint and type check
ruff check .
mypy src/

License

MIT

Available Tools

20 tools
add_dhcp_static_mappingA

Create a DHCP static mapping (IP reservation) for a MAC address.

Staged until apply_changes('dhcp') is called or apply=true is passed.

ParametersJSON Schema
NameRequiredDescriptionDefault
macYes
applyNo
descrNo
ipaddrYes
hostnameNo
interfaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and it uses it well: it discloses the non-obvious staged-apply behavior and the effect of the apply parameter. It could add conflict/idempotency behavior, but the core side-effect model is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences: the first states the purpose, the second states the critical staging behavior. No filler or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the output schema likely covers return values, but with six parameters and no annotations the description leaves parameter semantics mostly to inference. The staging behavior is the essential context and is present, but an agent still lacks guidance on required interface/IP values or duplicate handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains MAC address context and the apply boolean. The required interface and ipaddr parameters plus hostname and descr receive no semantic explanation beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair: "Create a DHCP static mapping (IP reservation) for a MAC address." It is clearly a creation operation and distinct from siblings like delete_dhcp_static_mapping and list_dhcp_static_mappings, though it does not explicitly call out the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The staging note is practical usage guidance: the mapping is not persisted until apply_changes('dhcp') is called or apply=true is passed. This tells the agent when follow-up is required, but it does not explicitly contrast with sibling tools such as delete_dhcp_static_mapping or list_dhcp_static_mappings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_dns_host_overrideA

Create a DNS host override entry in Unbound DNS Resolver.

Staged until apply_changes('dns') is called or apply=true is passed.

ParametersJSON Schema
NameRequiredDescriptionDefault
ipYes
hostYes
applyNo
descrNo
domainYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It explicitly states that the operation is staged and requires a separate apply step, which is a non-obvious and important behavioral detail. It doesn't mention error handling or idempotency, but for a create operation, the staging mechanism is the most critical behavior and is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise: two sentences with no wasted words. The core purpose is stated first, and the staging behavior is added as a clear second sentence. This is an efficient use of space and easy for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The staging behavior is covered, and an output schema exists (though not shown here), which may describe return values. However, the description does not explain the parameters at all, which is a significant gap given the schema provides no descriptions. For a tool with five parameters, three required, the agent needs more guidance to use it correctly. The description is adequate for purpose but incomplete for parameter usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It does not. The agent is left to infer what 'host', 'domain', and 'ip' mean, and there is no guidance on format, examples, or how 'apply' and 'descr' are used. The names are somewhat self-explanatory, but without details, the agent may pass incorrect values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create') and resource ('DNS host override entry') with an explicit target ('Unbound DNS Resolver'). This is immediately distinguishable from sibling tools like list_dns_host_overrides and delete_dns_host_override, which operate on the same resource but with different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a crucial usage context: the operation is staged until apply_changes('dns') is called or apply=true is passed. This tells the agent when the change takes effect and how to trigger it. It doesn't explicitly mention alternative tools, but the staging guidance is a strong usage signal that helps the agent decide how to invoke this tool correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_firewall_ruleA

Add a firewall rule. Type is 'pass', 'block', or 'reject'.

The rule is staged (not active) until apply_changes('firewall') is called or apply=true is passed here.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYes
applyNo
descrNo
sourceNoany
dstportNo
protocolNo
interfaceYes
ipprotocolNoinet
destinationNoany

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It reveals a critical non-obvious behavior: the rule is staged, not active, until committed via apply_changes or apply=true. It does not disclose other effects like rule ordering or validation, but the most important mutation behavior is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, front-loaded, and every sentence earns its place. It states the essential purpose, the allowed values, and the crucial staging behavior without repetition or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core call semantics and staging behavior, and the presence of an output schema reduces the need to document return values. However, most parameters remain semantically unexplained, so an agent may struggle to construct correct non-trivial rules involving protocols, ports, or interfaces.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains the type values and apply behavior. The meanings and formats of source, destination, protocol, dstport, interface, ipprotocol, and descr are left undocumented. This is insufficient for a 9-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action: 'Add a firewall rule' with a concrete resource and enumerates the valid type values. This is distinct from sibling tools like delete_firewall_rule and list_firewall_rules, and the staging note separates it from apply_changes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when the rule takes effect and explicitly references apply_changes('firewall') and the apply=true alternative. It does not spell out when to avoid this tool or choose a sibling, but the staging vs. immediate-apply distinction is practical guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

apply_changesA

Apply ALL staged changes of a subsystem ('firewall', 'dhcp' or 'dns'). Requires confirm=true.

This reloads the subsystem, activating every pending change — including any a human staged in the pfSense WebGUI and has not reviewed yet.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo
subsystemYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It discloses that the tool reloads the subsystem and activates every pending change, including those staged by humans, which is a significant side effect. It also mentions the confirm requirement. However, it does not mention whether the operation is reversible, what happens on failure, or if it could cause downtime. For a mutation tool, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence front-loads the action and the confirm requirement. The second sentence adds important context about the scope (all pending changes including human-staged). Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, the description covers the key aspects: what it does, what it requires, and a warning about the scope. It does not explain the output format, but since the tool has an output schema, that is not the description's job. It could mention irreversibility or potential service disruption, but the warning about applying ALL changes is a strong implicit caution. Overall, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It adds meaning for both parameters: it specifies the valid values for 'subsystem' and states that 'confirm' must be true, which is not in the schema (schema only shows type and default). This goes beyond the schema and helps an agent use the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the verb 'apply' and the resource 'ALL staged changes of a subsystem', and explicitly enumerates the valid subsystems ('firewall', 'dhcp' or 'dns'). This distinguishes it from the sibling get_pending_changes (which retrieves changes) and restart_service (which restarts a service).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states a key precondition ('Requires confirm=true') and warns that it applies all pending changes, including human-staged ones. However, it does not explicitly say when to use this tool vs alternatives (e.g., after reviewing pending changes, or when you only want to apply your own changes). It gives context but not explicit routing or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_dhcp_static_mappingA

Delete a DHCP static mapping. Requires confirm=true.

interface is the mapping's DHCP server — the parent_id value returned by list_dhcp_static_mappings; mapping_id is its id there. Staged until apply_changes('dhcp') is called or apply=true is passed.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNo
confirmNo
interfaceYes
mapping_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full transparency burden. It clearly discloses that the delete requires confirmation, is staged rather than immediately destructive, and can be made permanent via apply_changes or apply=true. This is strong behavioral disclosure for a destructive mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: action, required flag, parameter semantics, then staging behavior. Each sentence earns its place and the backtick formatting makes the key identifiers and values easy to scan. There is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter mutation tool with no annotations and an output schema present, the description is complete: it states what is deleted, the required confirmation, how to identify the correct interface and mapping, and how to commit or stage the change. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does: interface is explained as the mapping's DHCP server and the parent_id from list_dhcp_static_mappings, while mapping_id is its id in that listing. It also explains the roles of confirm and apply. This is far more than the schema alone provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and resource: 'Delete a DHCP static mapping.' This directly distinguishes it from sibling delete tools like delete_firewall_rule and delete_dns_host_override. It also adds essential behavior context (confirm and staging) rather than merely restating the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational guidance: confirm=true is required, and the delete is staged until apply_changes('dhcp') or apply=true. It also points to list_dhcp_static_mappings for finding identifiers. It does not explicitly state when not to use this tool, but the prerequisites and workflow are clear enough for an agent to select and invoke it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_dns_host_overrideA

Delete a DNS host override by ID. Requires confirm=true.

Staged until apply_changes('dns') is called or apply=true is passed.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNo
confirmNo
override_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior, and it does: it reveals that the deletion is staged until explicitly applied, and that confirmation is mandatory. This adds significant insight beyond the schema, though it does not mention error handling or whether the deletion is reversible. The staging mechanism is a critical behavioral trait that the agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The main action is front-loaded, followed by the mandatory confirmation and the staging behavior. Every word adds value, and the structure is clear and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deletion tool with moderate complexity, the description covers the essential operational details: what it does, the confirmation requirement, and the staging mechanism. The output schema exists, so return format is handled. It does not mention reversibility or downstream effects, but these are not critical for invoking the tool correctly. It is sufficient for an agent to use it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clarifies the confirm parameter (required) and apply parameter (staging vs immediate), and the override_id is implicitly the target by 'by ID'. However, it does not explain that confirm is a safety flag or the default behavior of apply=false, leaving some room for interpretation. It covers the key parameters but not exhaustively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (delete), a specific resource (DNS host override), and an identifier (by ID). It distinguishes from sibling tools like add_dns_host_override and list_dns_host_overrides by the action and the mention of ID-based deletion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states the prerequisite 'Requires confirm=true' and explains the staging behavior with apply_changes('dns') or apply=true. This gives clear context on when the action takes effect, though it does not explicitly list alternatives or when not to use it. The staging behavior is a strong usage differentiator from direct-apply tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_firewall_ruleA

Delete a firewall rule by its ID (the id from list_firewall_rules). Requires confirm=true.

Staged until apply_changes('firewall') is called or apply=true is passed.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNo
confirmNo
rule_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses two important behavioral traits: the confirmation requirement and the staging/apply mechanism. It also implies the destructive nature of the operation. It does not describe side effects like cascading dependencies, but the disclosure is adequate for the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, with the primary action and confirmation requirement stated first, followed by the staging behavior. It is concise with no filler words—every phrase adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for a delete operation with a staging mechanism. It explains the two ways to apply changes and the required confirm flag. It does not cover error handling or preconditions (e.g., rule must exist), but the output schema likely covers return values. Given the complexity and the presence of an output schema, this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It clarifies rule_id as the ID from list_firewall_rules, confirm as required (true), and apply as the trigger for immediate application versus staging. This provides semantic meaning for all three parameters, though it does not explicitly state defaults or types (which are in the schema).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (delete a firewall rule), identifies the resource by ID, and references the source for obtaining the ID (list_firewall_rules). It is specific and unambiguous, distinguishing it from other delete tools that target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit conditions for use: it requires confirm=true and explains that the deletion is staged until apply_changes or apply=true. This tells the agent when the action will take effect, but it does not explicitly contrast with alternative delete tools (e.g., delete_dhcp_static_mapping) or state when not to use it. The usage guidance is solid for the tool's own mechanics but lacks explicit sibling differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_arp_tableA

Get ARP table showing connected devices (IP, MAC, interface).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It correctly identifies the operation as retrieving data (non-destructive), but omits details like authentication requirements, data freshness, or behavior when no devices are connected. Adequate but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the tool's purpose and output.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters and an output schema (present), the description sufficiently explains what the tool does and what it returns. It covers the essential information but does not address edge cases or provide extra context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so baseline is 4. The description does not need to add parameter information, and it correctly omits it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves the ARP table and specifies the data shown (IP, MAC, interface). It is a specific verb+resource combination that distinguishes it from sibling tools like get_interfaces or list_dhcp_leases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide any when-to-use or when-not-to-use guidance, nor does it mention alternatives among the sibling tools. Usage is implied but not explicitly addressed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_firewall_logsA

Read-only view of recent firewall log entries.

Each entry contains its ID and raw log text from pfrest. limit caps the number of entries (default 50). This tool only reads logs; it never writes to the firewall.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It explicitly states read-only behavior and never writes, which is key. It also describes entry contents (ID and raw log text) and the limit's effect. However, it doesn't disclose ordering, time window of 'recent', or any potential side effects beyond not writing. Given the simplicity, the coverage is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no wasted words. The purpose is front-loaded in the first sentence, and the second adds entry details and parameter explanation. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and an output schema, the description covers the parameter and basic return info. The output schema likely details the full response. The only gap is the vagueness of 'recent' (no time window or ordering), but that's not critical for basic usage. The tool is complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage, so the description must explain the limit parameter. It does so clearly: '`limit` caps the number of entries (default 50).' This adds meaning about the effect and default, compensating for the schema's silence. It doesn't give range or examples, but the purpose is clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides a 'read-only view of recent firewall log entries', specifying the resource and action. It is distinct from siblings like get_system_status or get_interfaces, though it doesn't explicitly name alternatives. The verb 'view' and resource 'firewall log entries' make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes the tool 'only reads logs' and 'never writes', which is a behavioral exclusion but not a usage guideline. It doesn't say when to use this tool versus other read tools (e.g., get_gateway_status, list_firewall_rules) or mention any prerequisites or context where this is preferred. There is no guidance on when not to use it or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_gateway_statusA

Get gateway status including latency, packet loss, and online state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses the returned data (latency, packet loss, online state) and implies read-only behavior. Does not mention errors or side effects, but adequate for a simple read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that is front-loaded with the verb and resource, containing only essential information. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool is simple with no parameters and an output schema exists. Description adds the specific fields returned, which is complete context for an agent to understand the tool's functionality.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist; schema coverage is 100% (vacuous). Baseline 4 for zero parameters. No additional parameter information needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the verb 'Get' and the resource 'gateway status', listing specific attributes (latency, packet loss, online state). It distinguishes from sibling tools like 'get_system_status' which likely returns different data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use vs alternatives such as 'get_system_status'. However, the tool has no parameters and the name 'get_gateway_status' implies it is for gateway-specific status, so usage is implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_interfacesA

List all network interfaces with status and configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description alone must convey behavioral traits. It states 'List' implying a read-only operation, but lacks details on side effects, authentication needs, or output specifics beyond what the output schema provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence of 6 words, front-loaded with verb and resource, no waste. Appropriate for a simple list operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters and an output schema present, the description is nearly complete. It states what it lists and includes status/configuration. Could mention that it returns all interfaces, but output schema likely covers return structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so schema coverage is 100%. The description correctly implies no parameters are needed, which is sufficient. Baseline 4 for zero-parameter tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists all network interfaces with status and configuration, distinguishing it from siblings like get_arp_table (ARP table) and get_gateway_status (gateway status).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not specify any prerequisites, exclusions, or comparisons with sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pending_changesA

Check whether a subsystem ('firewall', 'dhcp' or 'dns') has staged, unapplied changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
subsystemYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It says 'check' which implies a non-mutating operation, but it does not explicitly state read-only behavior, potential side effects, or error handling. It also does not describe what happens if the subsystem is invalid or if there are no pending changes, leaving the agent to infer from the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence that front-loads the action and clearly defines the parameter scope. No redundant words or filler. The structure is optimal for a simple check tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and an output schema, the description covers the essential purpose and parameter values. It does not explain the return format, but the output schema likely handles that. It also omits usage context (e.g., calling before apply_changes), but that is covered under usage guidelines. Given the simplicity, it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only a string parameter with no description (0% coverage). The description compensates by enumerating the valid values ('firewall', 'dhcp' or 'dns'), giving the agent precise knowledge of what the parameter expects. This adds significant meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check whether') and a clear resource ('a subsystem has staged, unapplied changes'). It explicitly lists the valid subsystem values, making the purpose unambiguous and distinct from sibling mutation tools like add/delete/apply.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention that this should be called before apply_changes, nor does it state any exclusions or conditions that would lead an agent to choose a different tool. The description only states what it does, not when.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_system_statusA

Get pfSense system status including version, CPU, memory, uptime, and temperature.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must cover behavior. It does not mention if the operation is read-only, required permissions, or side effects. For a status tool, read-only implied but not explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with verb, no unnecessary words. Efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and existence of output schema, description adequately lists included status fields. Could mention real-time vs cached, but not a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, so baseline is 4 per rules. Description adds value by listing output content, compensating for lack of param info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states verb 'Get' and resource 'pfSense system status' with specific items (version, CPU, memory, uptime, temperature). It is distinct from sibling tools which focus on DHCP, DNS, firewall, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives (e.g., get_gateway_status for gateway info). Sibling list exists but no exclusions or context provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dhcp_leasesA

List active DHCP leases showing IP, MAC, hostname, and lease times.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It indicates a read-only, non-destructive operation. Does not mention auth or rate limits, but for a listing tool this is adequate. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with verb and resource, no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and existence of output schema, description is complete. It explains what the tool returns and suffices for a simple listing operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, and schema coverage is 100%. Baseline is 3 as description adds no parameter-specific info, which is fine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists active DHCP leases and specifies the fields (IP, MAC, hostname, lease times). It distinguishes from sibling tools like list_dhcp_static_mappings which deal with static mappings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for viewing active leases. While no explicit when-not or alternatives are given, the context of sibling tools makes it clear. For a simple no-parameter tool, this is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dhcp_static_mappingsA

List DHCP static mappings (IP reservations), optionally filtered by interface.

Each mapping carries parent_id (its DHCP server / interface) — pass that as interface to delete_dhcp_static_mapping.

ParametersJSON Schema
NameRequiredDescriptionDefault
interfaceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description makes the read-only operation clear via 'List' and adds a behavioral/structural detail: each mapping carries parent_id representing its DHCP server/interface, and that value is reusable for deletion. No annotations are provided, but this disclosure is adequate for a simple list operation with an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, with the core purpose front-loaded and the cross-tool usage tip in the second sentence. No filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-optional-parameter list tool with an output schema, the description covers the operation and the only param's semantics. It could add an explicit 'returns all if interface is omitted' clarification, but the schema default already conveys that, so it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the name/type of interface with 0% description coverage, while the description gives its real meaning: a DHCP server/interface identifier taken from parent_id, used to filter the listing. This fully compensates for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List DHCP static mappings (IP reservations)', with an optional interface filter. The parenthetical differentiates static reservations from sibling list_dhcp_leases, so the agent can select the right list tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly explains how to use the optional interface filter and connects it to the delete sibling by instructing the agent to pass each mapping's parent_id as the interface value. It does not explicitly enumerate when-not-to-use or alternatives, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dns_host_overridesA

List DNS Resolver host overrides (local DNS entries).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only states the basic operation and does not mention required permissions, scope of results (all entries vs. filtered), pagination, or the return format. The output schema exists but is not referenced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It conveys the purpose efficiently and is appropriately sized for a simple list operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (zero parameters, clear sibling differentiation), the description is adequate for basic understanding but lacks usage guidance and behavioral details. It is missing information to fully inform an agent for correct invocation in varied contexts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters and 100% schema description coverage. According to the rules, a baseline of 4 applies for zero parameters, and the description adds no unnecessary information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List') and the resource ('DNS Resolver host overrides'), with a parenthetical explanation for clarity. It is distinct from sibling tools like add_dns_host_override and delete_dns_host_override.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives (e.g., when to list vs. add/delete). The description implies the use case by naming the action, but there is no explicit context or exclusionary language.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_firewall_aliasesA

List firewall aliases (IP groups, port groups, URL lists).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral disclosure. It only states the action without mentioning read-only nature, authentication needs, pagination, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence that is front-loaded with the verb and resource. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema (not shown) and no parameters, the description is largely complete for a simple list tool. It could mention that it returns all aliases or any prerequisites, but it's sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so schema coverage is 100%. The description adds semantic value by enumerating alias types (IP groups, port groups, URL lists), which goes beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists firewall aliases and specifies the types (IP groups, port groups, URL lists). This is specific and distinct from sibling tools like list_firewall_rules or list_dhcp_leases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like list_firewall_rules or list_services. The description does not mention context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_firewall_rulesC

List firewall rules, optionally filtered by interface.

ParametersJSON Schema
NameRequiredDescriptionDefault
interfaceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose any behavioral traits such as read-only nature, permissions, or pagination, which would be expected for a list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that is front-loaded with the main action. It contains no waste, though it could be slightly more informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and an output schema, the description is minimally complete. However, it lacks context about output structure or limitations, which the output schema does not fully replace.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions optional filtering by interface but does not explain the format or constraints of the parameter, adding little value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (list) and resource (firewall rules), and mentions optional filtering. It is specific but does not explicitly differentiate from sibling tools like list_firewall_aliases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description only states what it does without providing context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_servicesA

List all services and their running status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, but description clearly states read-only behavior (list and status). No side effects or permissions mentioned, which is acceptable for a simple list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with essential information, no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and an output schema, the description is complete. It adequately covers the tool's purpose and behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, so the description correctly provides no parameter info. Baseline 4 for zero-parameter tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and resource 'services' with additional detail 'running status', distinguishing it from sibling tools like restart_service.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or alternatives guidance, but the context of sibling tools implies this is for viewing services status, not modifying them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restart_serviceB

Restart a service by name. Requires confirm=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
confirmNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavior. It only states the action and a constraint (confirm=true), but does not disclose that restart is disruptive, may require permissions, or what happens on failure/success.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no wasted words, front-loaded with the primary action and a critical requirement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a disruptive action with no annotations and no output schema details provided, the description is too minimal. It lacks prerequisites, side effects, and post-condition info beyond the bare minimum.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must clarify parameters. It adds meaning for 'confirm' (must be true) but only implicitly for 'name' (by name). No details on valid service names or how to obtain them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Restart') and resource ('service'), and specifies the parameter ('by name'), distinguishing it from sibling tools like list_services which merely list services.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions the requirement for confirm=true, which provides guidance on invocation, but does not discuss when to use this tool versus alternatives, nor any prerequisites or conditions for restart.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.3.0
    • Changedadd_dhcp_static_mapping1 field changed
      • addedInput schema / properties / apply
        Added value: +{
        +  "default": false,
        +  "title": "Apply",
        +  "type": "boolean"
        +}
    • Changedadd_dns_host_override1 field changed
      • addedInput schema / properties / apply
        Added value: +{
        +  "default": false,
        +  "title": "Apply",
        +  "type": "boolean"
        +}
    • Changedadd_firewall_rule1 field changed
      • addedInput schema / properties / apply
        Added value: +{
        +  "default": false,
        +  "title": "Apply",
        +  "type": "boolean"
        +}
    • Addedapply_changes
    • Changeddelete_dhcp_static_mapping3 fields changed
      • addedInput schema / properties / apply
        Added value: +{
        +  "default": false,
        +  "title": "Apply",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / interface
        Added value: +{
        +  "title": "Interface",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "mapping_id"
        -]New value: +[
        +  "interface",
        +  "mapping_id"
        +]
    • Changeddelete_dns_host_override1 field changed
      • addedInput schema / properties / apply
        Added value: +{
        +  "default": false,
        +  "title": "Apply",
        +  "type": "boolean"
        +}
    • Changeddelete_firewall_rule1 field changed
      • addedInput schema / properties / apply
        Added value: +{
        +  "default": false,
        +  "title": "Apply",
        +  "type": "boolean"
        +}
    • Addedget_firewall_logs
    • Addedget_pending_changes
  2. 17 tool updatesv0.1.1
    • First observedadd_dhcp_static_mapping
    • First observedadd_dns_host_override
    • First observedadd_firewall_rule
    • First observeddelete_dhcp_static_mapping
    • First observeddelete_dns_host_override
    • First observeddelete_firewall_rule
    • First observedget_arp_table
    • First observedget_gateway_status
    • First observedget_interfaces
    • First observedget_system_status
    • First observedlist_dhcp_leases
    • First observedlist_dhcp_static_mappings
    • First observedlist_dns_host_overrides
    • First observedlist_firewall_aliases
    • First observedlist_firewall_rules
    • First observedlist_services
    • First observedrestart_service

TDQS

A3.5/5.0

Scored across 20 tools

Disambiguation4/5

Most tools target a distinct resource and action, such as firewall rules, DHCP mappings, DNS overrides, logs, and ARP. A few read-only tools like get_arp_table and list_dhcp_leases could be confused when looking for connected devices, but their descriptions clarify the difference.

Naming Consistency3/5

The set largely uses snake_case verb_noun naming, but read operations are split inconsistently between get_ and list_ (e.g., get_interfaces vs list_firewall_rules). Add/delete pairs are consistent, while apply_changes and get_pending_changes follow a different style.

Tool Count4/5

20 tools is on the heavier side, but it is justified by pfSense's broad scope covering firewall, DHCP, DNS, system status, services, and networking. The count is slightly over the ideal range but not bloated for the domain.

Completeness3/5

Firewall rules, DHCP static mappings, and DNS overrides each support add/list/delete, and the staged-apply workflow is covered. However, there are no update/edit tools for any managed resource, which is a notable gap for a management server.

Maintenance

ActivityMaintained
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables natural language interaction and management of pfSense firewalls through Claude and other GenAI applications using the Model Context Protocol. It provides advanced tools for firewall rule configuration, interface management, and intelligent log analysis via a REST API integration.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that gives AI agents full control over pfSense firewalls via the REST API v2, with 677 tools covering firewall rules, NAT, VPN, services, routing, certificates, users, diagnostics, and more.
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    A secure MCP server for managing OPNsense firewalls through AI assistants. Provides 81 tools across system, firewall, network, DNS, DHCP, VPN, HAProxy, services, diagnostics, and security domains.
    81
    232 PyPI
    21
    MIT