Skip to main content
Glama
README.md
# netauto

[![tests](https://github.com/ters-golemi/netauto/actions/workflows/tests.yml/badge.svg?branch=main)](https://github.com/ters-golemi/netauto/actions/workflows/tests.yml)
![read-only](https://img.shields.io/badge/devices-read--only%20by%20default-brightgreen)
![MCP](https://img.shields.io/badge/interface-MCP%20%2B%20web%20GUI-blue)
![python](https://img.shields.io/badge/python-3.11%2B-blue)
![license](https://img.shields.io/badge/license-Apache--2.0-blue)

Multi-vendor network automation, exposed to agents over MCP. **Read-only by
default**: it reads device state, audits configuration, and diffs proposed
changes, with no path to commit configuration. Its one device write is firmware
upgrade — a single, deliberately gated operation, off unless `allow_writes` is
enabled.

Covers Cisco (IOS, IOS-XE, NX-OS, IOS-XR), Juniper Junos, Arista EOS,
HPE Aruba (AOS-CX, AOS-Switch, Central), Cisco Meraki, Fortinet FortiOS and
Palo Alto Networks PAN-OS.

## Installing

For a server install on Ubuntu — service user, systemd, TLS, firewall — follow
[INSTALL.md](INSTALL.md). The quick start below is for a workstation.

### Sizing the VM

**Ubuntu 24.04 LTS.** Python **3.11 is a hard floor**: the Meraki SDK requires
it, so `pip install` fails outright on anything older rather than degrading.
That rules out 22.04 with its stock Python 3.10 unless you install a newer
interpreter alongside it.

| | Minimum | Room to grow | What actually uses it |
|---|---|---|---|
| vCPU | 1 | 2 | Work is I/O-bound — threads sit waiting on SSH, not computing. A second core keeps the UI responsive while a topology collection or a Word export is running. |
| RAM | 2 GB | 4 GB | The service stays under 100 MB; the rest is headroom for concurrent sessions and in-memory run history. |
| Disk | 5 GB | 10 GB | The venv is ~190 MB and the checkout ~11 MB. Everything else is Ubuntu itself and logs. |

**The service is smaller than the VM.** With every driver imported and every
page served it measures about 70 MB resident. Nothing is preallocated and
there is no database — netauto's state is an inventory file, a users file and
an append-only log.

**What the headroom is actually for.** Two things scale with your fleet rather
than with traffic:

- **Concurrent device sessions.** Topology collection and the metrics collector
  fan out to `max_concurrency` (default 8) sessions at once. Audits do not —
  the Audit page walks its targets one at a time, so a large group takes longer
  rather than more memory.
- **Run history, held in memory on purpose.** Up to 20 workflow runs and 20 lab
  jobs are kept, each carrying the full running configuration of every device it
  touched. Those contain password hashes, community strings and key material,
  and netauto keeps them off disk deliberately — the cost is RAM, and a restart
  discards them.

**Disk grows in one place.** `activity.log` is append-only JSON lines and
nothing rotates it. On a busy server give it room, or add a logrotate rule.

**Where the VM sits matters more than how big it is.** It needs reachability to
the devices you intend to manage — ideally from a management VLAN rather than a
general user network — and it holds device credentials in its environment.
Anyone who can reach it with an account can read your network's configuration.
Treat it as infrastructure, not as a convenience box.

**Two things this VM is not sized for**, both deliberately:

- **Bringing labs up.** The [Lab page](#labs-netlab) needs a hypervisor or
  container runtime, device images, and privileges the hardened service user
  does not have — that belongs on a separate lab host or a CI runner, and
  [INSTALL.md](INSTALL.md) says why. *Auditing* a lab that is already up costs
  this VM nothing extra: it reads the snapshot and connects read-only, exactly
  as it does to managed gear.
- **Grafana and Prometheus.** [Optional and separate](#metrics-and-grafana),
  brought up with `docker compose`. Putting them on this VM means Docker on it
  too, plus whatever retention you give the time series.

## Quick start

```bash
cd ~/Work/netauto
cp config.example.yaml config.yaml
cp inventory/devices.example.yaml inventory/devices.yaml   # then edit
export CORE_SW_USERNAME=admin CORE_SW_PASSWORD='...'       # per device prefix
.venv/bin/python -m pytest tests/ -q
```

The MCP server is registered in `~/Work/.mcp.json` and starts automatically in
Claude Code sessions rooted at `~/Work`. To run it by hand:

```bash
.venv/bin/python -m netauto.mcp_server
```

## The safety model

Four independent layers, because one is not enough:

1. **One write, gated three ways; no config commit path.** netauto reads,
   except for a single deliberate write — firmware upgrade. `Driver.apply_config`
   still raises `WriteDisabled` and no driver overrides it, so there is no path
   to commit arbitrary configuration. The firmware install is separate: it never
   goes through the command guard, and fires only past three independent gates —
   `allow_writes` enabled in `config.yaml` (off by default), the driver
   declaring the `upgrade` capability, and the caller confirming with the device
   name. Off by default, netauto is read-only; the Software Upgrade workflow then
   produces a runbook and touches nothing.
2. **Command allowlist.** `net_run_show` accepts only commands matching a
   per-vendor read allowlist, rejects a deny-list of state-changing verbs, and
   refuses command chaining (`;`, `&&`, `|`, newlines, `$(...)`, backticks).
   Unrecognised commands are refused rather than forwarded — default deny.
   Validation happens *before* connecting, so a bad command costs no session.
3. **Credentials never touch disk.** The inventory names an environment prefix;
   secrets resolve from the process environment at connect time. A device with
   prefix `CORE_SW` needs `CORE_SW_USERNAME` and `CORE_SW_PASSWORD`. No tool
   returns a credential.
4. **Ad-hoc targets are gated.** Connecting to a host that is not in the
   inventory makes the server offer its stored credentials to whatever answers,
   so the address must be one this process found in a sweep within the last
   hour, and must not be routable on the internet. Anything else is an
   inventory entry, which takes access to the server's filesystem.

## MCP tools

| Tool | Does |
|---|---|
| `net_list_devices` | Inventory, filterable by tag or platform |
| `net_supported_platforms` | The 12 platform strings with drivers |
| `net_device_facts` | Vendor, model, OS version, serial, hostname |
| `net_get_config` | Running/startup/candidate config as text |
| `net_run_show` | One allowlisted read-only command |
| `net_audit` | Hardening ruleset against a device or tag group |
| `net_config_diff` | Candidate vs running — review only, no commit |
| `net_discover_local` | ARP sweep of a directly attached segment, optionally probing management ports |
| `net_scan_ports` | Which of SSH, telnet and friends given hosts answer on |
| `net_connect_adhoc` | Read from a discovered host, without adding it to the inventory |
| `net_topology` | LLDP/CDP graph, as JSON or an editable draw.io file |

## Agents

The repo holds the canonical definitions in `agents/`. Claude Code loads them
from the project's `.claude/agents/` directory, so install them with:

```bash
mkdir -p ~/Work/.claude/agents
cp agents/*.md ~/Work/.claude/agents/
cp mcp.json.example ~/Work/.mcp.json   # then edit the paths if not under ~/Work
```

Three subagents:

- **network-auditor** — compliance across the estate; ranks by exploitability,
  reports unreachable devices separately from passing ones.
- **network-troubleshooter** — fault diagnosis; proposes fixes as diffs and
  stops there.
- **network-documenter** — inventories and baselines keyed to hardware rather
  than to leased addresses.

## Web GUI

A self-hosted browser interface so a team can use the toolkit without the CLI.
It runs *inside* the network it manages -- it needs reachability to the gear and
holds the device credentials in its own environment.

```bash
.venv/bin/python -m netauto.web.manage add adis --admin    # first account
export NETAUTO_SECRET_KEY="$(python3 -c 'import secrets;print(secrets.token_urlsafe(48))')"
export NETAUTO_WEB_HOST=0.0.0.0        # omit for localhost only
./run-web.sh
```

![The Overview page: four counts, the inventory broken down by vendor family
and by tag, and the thirteen platform strings that have drivers](docs/dashboard.png)

*Synthetic inventory. The counts are of what is in the inventory file; the
thirteen platforms are what netauto has drivers for, whether or not you run any.*

![The Devices page: the inventory with each device's platform, guard family,
address and tags, and an Inspect button per row](docs/devices.png)

*Synthetic inventory. The `arista_eos` rows show family `cisco` because the
command guard judges Arista by Cisco syntax — the family is the allowlist that
applies, not the vendor.*

Pages: overview, device list, per-device facts and running config, a read-only
command box, compliance audit by group, ARP discovery with an optional
SSH/telnet port check, ad-hoc sessions against discovered hosts, LLDP
topology with a
draw.io export, Workflows, a Lab tab for building and auditing virtual networks
with netlab, an activity log, and a Metrics tab embedding the
Grafana dashboard when one is configured.

### Accounts

Per-user, so every action is attributable. Accounts live in `users.yaml` as
bcrypt hashes at `0600`; the file is re-read on each lookup, so adding or
removing a user takes effect without a restart.

```bash
python -m netauto.web.manage add <name> [--admin]
python -m netauto.web.manage list
python -m netauto.web.manage passwd <name>
python -m netauto.web.manage remove <name>
python -m netauto.web.manage admin <name> [--revoke]
```

Passwords are prompted for, never passed as arguments, so they stay out of
shell history and the process table. Minimum length is 10 characters, and the
last admin cannot be removed. Admins can read the activity log; standard
accounts cannot.

### Activity log

Every device-touching action is appended to `activity.log` as JSON lines with
the account that made it -- logins, failed logins, device inspections, commands
run, commands refused, audits and discovery sweeps. Greppable directly, or
viewable at `/activity` by an admin.

![The Activity page: each action with the account that made it, the target and
the detail — a refused command showing why it was refused](docs/activity.png)

*Synthetic history. This is the only screenshot taken as an administrator,
which is why an Activity tab appears in the navigation — a standard account
does not get one, or the page behind it.*

### Access control

Bcrypt verification with a dummy comparison for unknown users so response time
does not reveal which accounts exist. Signed `HttpOnly` `SameSite=Strict`
session cookies, a CSRF token on the login form, and a five-attempt lockout
keyed to *(client address, username)* -- so locking out one account cannot lock
out the team.

**It serves plain HTTP.** On a shared network, passwords and retrieved configs
cross the wire in clear text. For anything beyond a trusted management VLAN,
put it behind a reverse proxy with TLS and set `https_only=True` on the session
middleware in `netauto/web/app.py`.

**The write is gated and enumerated.** A test pins every POST route the app
exposes to a justified allowlist in `tests/conftest.py`. The only one that can
*write* to a managed device is the workflow-start route, and only the Software
Upgrade workflow writes through it — past `allow_writes`, the upgrade capability
and a device-name confirmation. The Lab routes POST too, but they orchestrate
ephemeral netlab infrastructure (`/lab/up`, `/lab/down`) or read lab devices
read-only (`/lab/audit`) — never a managed-device write. Every managed-device
route that reads is a GET.

## Workflows

Three multi-step pipelines per supported platform -- 36 in all, under
the **Workflows** tab. They are per-platform because the useful part is the
show-command set, and that does not generalise.

![The Workflows tab: each platform lists all three pipelines with every step
named, how many inventory devices it would run against, and an optional tag
filter](docs/workflows.png)

*Synthetic inventory. Each platform's card states its device count, and a
platform with none says so instead of offering a run.*

**Device Configuration Check** connects, pulls the running configuration and
the platform's read-only command set, compares both against the vendor-guide
ruleset, and reports what to improve -- each finding with a severity, the
evidence from the config, a recommendation, and the guide it comes from.

**Network Documentation Maker** does all of that, then maps LLDP/CDP
neighbours and writes an editable Word document: inventory, findings, and a
topology diagram embedded as a picture.

**Software Upgrade** is the one workflow that can write. It snapshots the device
(running version and a configuration backup), checks the version against a
recommended target and the upgrade path, and produces a copy-pasteable runbook.
The firmware install happens only when armed — `allow_writes` on, an image and
server supplied — and otherwise the runbook is the whole output. This is the
single exception to read-only, gated as the safety model describes.

```
Configuration check:   connect -> config + show output -> compare -> report
Documentation maker:   ... the above ... -> neighbours -> assemble -> .docx
Software upgrade:      connect -> snapshot -> verify path -> [install] -> runbook
```

Runs happen in the background: start one, watch the steps, come back to it.
A run over an estate takes minutes because each device is a real session, so
holding an HTTP request open for it was never going to work.

Three things worth knowing before relying on it:

**The command sets are data, and the guard still decides.** Every command a
workflow can send lives in `netauto/workflows/spec.py` as a plain tuple, and
the test suite runs every one of them through `assert_read_only` for its
platform. A workflow cannot widen what netauto may send to a device; only an
edit to `READ_ALLOW` can, and that is a deliberate change to the guard itself.
The guard rejects `|` as chaining, which is why nothing here pipes -- no
`| display set`, no `| section`.

**Three platforms have no CLI.** Meraki, Aruba Central and AOS-CX expose no
command interface through their drivers, so those workflows pull state through
the API and mark the show-command step "skipped" with the reason. An empty
command set for them is a statement about the platform, not an omission.

**Runs are held in memory and nowhere else.** They contain full running
configurations -- password hashes, community strings, key material -- and
netauto keeps that off disk. Documents are built on demand and streamed. The
cost is real: restarting the service discards run history, and you re-run the
workflow.

Workflows evaluate a larger ruleset than the Audit page: `checks/vendor.py`,
which is the built-in rules annotated with their sources plus the guidance that
only makes sense once you have the show output too. The Audit page and the
Prometheus metrics keep running `BUILTIN` unchanged, so this cannot move a
dashboard or fire an alert.

## Labs (netlab)

Build a virtual network with [netlab](https://netlab.tools), then read and audit
it with everything above. netauto inspects devices but cannot stand a network
up to inspect; netlab builds and configures virtual topologies but does not
audit them afterward. The **Lab** tab joins the two, so a design can be checked
on a throwaway copy before it reaches real hardware.

The loop is **build → read → audit**, and it reuses what is already here: a lab
that is up has its nodes mapped to netauto platforms and appears as an ordinary
inventory, so Audit, Topology and the command box work against it unchanged —
no new drivers.

- **Bring up / tear down** a topology from the page; netlab up/down run as
  background jobs, one lab per topology at a time.
- **Audit this lab** runs the same compliance ruleset the Audit page runs, and
  shows the findings per device. An unreachable node is an error, not a silent
  pass — and lab findings are kept out of the compliance metrics, because a lab
  is not the fleet.

![The Lab page with the two sample topologies: mixed-vendor down with a Bring up
button, and spine-leaf up with its four nodes mapped to arista_eos and their
management IPs, a Tear down and an Audit this lab button](docs/lab.png)

*The repo ships two topologies — `spine-leaf` (all containers, run it today) and
`mixed-vendor` (four vendors, needs VM images). The node addresses above are
synthetic; the platform mapping and the page are real.*

![The Lab audit results for the mixed-vendor lab: four tiles, then each node
with its rules, severities and verdicts, failures highlighted and the config
line that decided each — juniper_junos and arista_eos shown, cisco_nxos and
cisco_ios below them on the page](docs/lab-audit.png)

*Auditing the `mixed-vendor` lab: the same rules and finding layout the Audit
page produces for managed devices, here across four vendors at once. Synthetic
findings; the rules, severities and page are real.*

The same loop headless, as a CI design-regression gate:

```bash
python -m netauto.lab.ci labs/spine-leaf/topology.yml --fail-on high
# exit 0 clean · 1 findings (design regressed) · 2 could not build/audit
```

netlab is **optional and never a dependency**: it needs libvirt or containerlab
plus device images, all host-level, so netauto never imports it — the wrapper
runs the installed binary and, where it is absent, the page says so and the
up/down actions refuse rather than pretend. The read-only stance is untouched:
netlab writes only to the throwaway VMs and containers it creates and destroys,
and netauto's connection *to* lab devices stays read-only — it audits them, it
does not configure them.

Only the platforms netauto already drives are mapped (Cisco IOS/NX-OS/IOS-XR,
Arista EOS and Juniper Junos); other netlab kinds — including container-native
ones like FRR, VyOS and SR Linux — are listed but skipped, with the reason. Lab devices share one credential prefix
(`LAB` by default), so export `LAB_USERNAME` and `LAB_PASSWORD` before auditing.
See [docs/netlab-integration.md](docs/netlab-integration.md) for the design, and
[docs/lab-audit.ci.yml](docs/lab-audit.ci.yml) for a copyable CI workflow.

## Discovery, port scanning and ad-hoc sessions

Two stages, because they answer different questions. `arp-scan` finds live
hosts on a directly attached segment -- authoritative there, since hosts answer
ARP even when they drop ICMP, and useless off it. Tick *check ports* and each
host that answered is then dialled on the listed ports, SSH and telnet by
default:

![The Discover page: five hosts on a swept segment, SSH and telnet checked on
each, with banners, and a link to the two already in the inventory](docs/discover.png)

*Synthetic hosts — the addresses, MACs and banners above are made up.*

```
# in the GUI: Discover -> 192.168.1.0/24, check ports, 22,23
# as an agent tool:
net_discover_local(cidr="192.168.1.0/24", probe_ports="22,23")
net_scan_ports(hosts="192.168.1.1, 192.168.1.9")   # no ARP, so any routed address
```

The probe is an ordinary TCP `connect()` and, at most, a read of whatever the
service volunteers first. It sends nothing, needs no privileges, and learns
nothing a client dialling the port would not. SSH names its software before
the client speaks, so an open 22 usually comes back with the far end's banner
-- often enough to tell a Cisco from an OpenSSH host. Telnet opens with binary
option negotiation, which is reported as no banner rather than as mojibake.

An open telnet port is flagged separately from an open SSH port: it carries
credentials in clear text, so it is a finding rather than an inventory fact.
Addresses already in the inventory are named in a column of their own, which
makes the interesting rows the ones that are *not* -- hosts on the wire that
nobody documented.

Bounded on purpose: at most a /22 per sweep, 16 ports, and 4096 probes in
total. An unbounded range times an unbounded port list is how a scan turns
into an hour-long page load.

### Connecting to what the sweep found

A discovered host that answers on 22 gets a *connect* link, which opens the
ordinary device page against it -- facts, running configuration, and the same
guarded command box a managed device gets. The Device is built for that one
request and stored nowhere; the platform is pre-selected from the SSH banner
and MAC vendor, and the credentials come from an environment prefix the server
already carries, chosen by name. No secret is typed into the browser.

![The connect page for a discovered host: what the sweep saw, a platform
pre-selected from the MAC vendor, and a credentials prefix chosen by
name](docs/connect.png)

The prefixes offered are read out of the environment, so the list is exactly
what this server can authenticate with -- export `LAB_USERNAME` and
`LAB_PASSWORD` before starting it and `LAB` appears. Only names are read.

What this deliberately is not: a shell. The guard polices one command at a
time, and an interactive session is a stream, so there is no terminal here and
the footer's promise holds on this page too. It is also not a way to manage a
device -- audits, metrics and workflows all read the inventory, so a host worth
keeping belongs in `inventory/devices.yaml`.

Agents get the same thing through `net_connect_adhoc`, gated identically and
per process -- an agent that has not swept has nothing it may connect to:

```python
net_discover_local(cidr="192.168.1.0/24", probe_ports="22")
net_connect_adhoc(ip="192.168.1.9", platform="aruba_aoscx", credentials="LAB")
net_connect_adhoc(ip="192.168.1.9", platform="aruba_aoscx", credentials="LAB",
                  command="show vlan", include_config=True)
```

Facts always come back, because they are how you confirm you reached what you
thought you did. A refused command or an unsupported config is reported beside
them rather than sinking the call.

The target gate is the part to understand before exposing this. See the safety
model above: a swept, unroutable address, or nothing.

## Topology diagrams

Builds a network diagram from what the devices themselves report over LLDP/CDP,
and exports it as a **draw.io** file — network stencils, orthogonal connectors
and port labels already placed. draw.io reads and writes `.vsdx`, so that file
is also the route to something editable in Visio.

![The Topology page after a collection: every link with both port names and
whether one end or both reported it, the devices that could not be asked, and
the neighbour with no inventory entry](docs/topology.png)

*Synthetic graph. Note the last row — a neighbour LLDP reported that has no
inventory entry of its own, which is usually the reason to run this.*

```bash
# in the GUI: Topology -> Discover links -> Download .drawio
.venv/bin/python -c "
from netauto import drawio, topology
from netauto.session import load_context
s, inv = load_context()
print(drawio.render(topology.build(inv, s)))" > topology.drawio
```

Every link is reported twice, once from each end. Links both ends agree on are
drawn solid; a link only one device reported is drawn dashed, because a
one-sided report usually means the far end has LLDP off rather than that the
cable is imaginary.

Neighbours with no inventory entry are drawn dashed and labelled *discovered* —
typically APs, phones and servers, occasionally a switch nobody wrote down.
Devices that could not report neighbours at all are listed separately with the
reason, so a thin diagram can be told from a small network.

Tiers come from inventory tags (`core`, `edge`, `distribution`, `access` and
their synonyms) and fall back to link count when a device carries none. LLDP
must be enabled on the devices; nothing appears for a link neither end
advertises.

Every member of a port-channel is kept as its own cable rather than merged
into one, since redundancy between core switches is usually the point of
looking. The exported file carries a caption with the collection time — an
undated network diagram is worse than none a year later.

## Metrics and Grafana

A Prometheus endpoint at `/metrics`, with a provisioned Grafana dashboard for
compliance drift and device reachability over time. See
[deploy/grafana/README.md](deploy/grafana/README.md).

```bash
cd deploy/grafana && docker compose up -d     # Grafana on 127.0.0.1:3000
```

Metrics are opt-in: with `NETAUTO_METRICS_TOKEN` unset the endpoint returns 404
and no collector runs. Prometheus cannot hold a session cookie, so that token
is the whole access control on the endpoint.

**Nothing polls your devices.** Audits are manual, and the audits you run
from the Audit page are what feed the metrics -- merged per device, so
auditing one switch does not blank the rest. `/metrics` serves that cache, so
scraping it costs nothing on the network. Set `NETAUTO_METRICS_INTERVAL` to a
number of seconds only if you do want a background sweep as well.

## Platforms and transports

| Platform string | Transport | Library |
|---|---|---|
| `cisco_ios`, `cisco_xe` | SSH CLI | napalm |
| `cisco_nxos`, `cisco_xr` | SSH CLI | napalm |
| `arista_eos` | eAPI | napalm |
| `juniper_junos` | NETCONF | napalm / PyEZ |
| `aruba_aoscx` | REST | pyaoscx |
| `aruba_osswitch` | SSH CLI | netmiko |
| `aruba_central` | Cloud REST | pycentral |
| `meraki` | Cloud REST | meraki |
| `fortinet_fortios` | REST | fortiosapi |
| `fortinet_cli` | SSH CLI | netmiko |
| `paloalto_panos` | SSH CLI | netmiko |

Meraki and Aruba Central are cloud tenants, not boxes: an inventory entry is an
organization or tenant, and neither has a CLI. `net_run_show` refuses on both
with an explanation rather than a generic failure.

## Compliance rules

15 rules in `netauto/checks/builtin.py`, scoped by platform family so a Junos
config is never judged by IOS syntax: telnet exposure, SSH version, cleartext
management, password encryption, enable secrets, session timeouts, root-login
policy, default SNMP communities, time sources, remote logging. This is what
the Audit page and the Prometheus metrics evaluate.

`netauto/checks/vendor.py` adds 15 more for the workflows -- AAA, legacy
services, management ACLs, BPDU guard, log timestamps, banners, Junos web
management and idle timeouts, Aruba loop protection, FortiOS trusted hosts and
remote logging, and writable SNMP -- and gives all 30 a `reference`, so a
recommendation can be traced to the guide it came from rather than read as an
opinion. `BUILTIN` itself carries none; the annotation is applied where the
workflows read the rules, which is why the Audit page and the metrics keep
evaluating exactly what they did before.

![The Audit page for one switch: four tiles, then every rule with its severity,
verdict and the configuration line that decided it — failures sorted to the top,
worst first](docs/audit.png)

*Synthetic switch, real rules: the config is invented, but every rule ID,
severity, detail and piece of evidence above is what netauto actually produced
from it.*

Rules take config text plus facts and return pass/fail with evidence. They
connect to nothing, so they unit-test against captured configs. Each vendor
rule is fired in both directions in `tests/test_vendor_rules.py`: a rule that
cannot fail and a rule that cannot pass both look healthy from the outside.

## Testing

The suite needs no hardware and no credentials. The command guard has the
heaviest coverage since it is the safety boundary — including chaining-escape
attempts and default-deny behaviour.

```bash
.venv/bin/python -m pytest tests/ -q
```

## What is not verified

The vendor drivers are written against each SDK's documented API and are
exercised by import and signature checks, and the **`fortinet_cli` path is
the first to have been run against real equipment** -- connect, authenticate,
`get system status`, full-configuration retrieval and the read-only guard,
against a FortiSwitch 108F on FortiSwitchOS 7.2.7. That run found the facts
parser returning `get system status` as one raw blob, since it was written
for FortiGate; it now parses the shared label set (model, os_version, serial,
hostname), tested against the captured 108F output. Every other driver path is
still exercised only by tests, and **local ARP discovery** was the only
real-equipment coverage before this. The port probe is tested against
real sockets on the loopback -- a listener that answers, one that stays silent,
a closed port -- but has not been pointed at production gear. An ad-hoc session
has never opened against a real device either: the target gate and the
transient Device are covered, and everything past them is the same driver code
as an inventory device, which is the code that has not met real gear. That includes the LLDP topology
path: the graph assembly and draw.io output are covered by tests against
captured neighbour tables, but no driver's `neighbors()` has met real gear. Cisco, Juniper, Aruba, Meraki,
Fortinet and PAN-OS paths need a first run against actual gear or a lab; expect to adjust
response parsing, particularly `aoscx_driver.get_config` and the Central
endpoint paths, which vary by firmware and region.

The workflows inherit all of that, and add their own. Every command set is
checked against the read-only guard and the pipelines are tested end to end
against fake drivers, but **no workflow has run against real equipment**. The
show commands are written from vendor documentation, so expect some to be
refused by a given model or software train -- the runner treats that as a
per-command gap rather than a failed device, which is exactly the case that
needs a real run to shake out. The rules themselves are tested against
representative config snippets, not captured production configs, so their
false-positive rate is unmeasured.

The **Lab (netlab) snapshot seam has been run against a real netlab install**
(netlab 26.08): `netlab create` produces the transformed topology, and netauto's
parser reads it and maps the nodes to platforms. The first real run corrected the
pinned command — the output form is `netlab create -o yaml=<file>` (an `=`, not
the `:` first assumed) — and confirmed netauto reads the management address from
`mgmt.ipv4` where real netlab records it. What is still unverified is a full
**`netlab up`** with a provider and device images, and auditing the resulting
live lab; once a lab is up, auditing it is the same driver code as an inventory
device, which has its own real-gear caveats above.

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md). The short version: the toolkit reads
devices and never changes them, four boundaries in the code exist to keep that
true, and the suite must pass without touching a network device.

## License

Apache License 2.0 — see [LICENSE](LICENSE) and [NOTICE](NOTICE).
Copyright 2026 Adis Cato.