tcp-tuner
by PiSpike
README.md
# TCP Congestion Tuner
An agentic TCP congestion-control tuning system where **Bob (IBM watsonx Code Assistant)** autonomously benchmarks, diagnoses, tunes Linux kernel TCP parameters, injects network faults, queries live Prometheus metrics, and commits runbooks to git — converging toward a user-defined SLO with zero human involvement per iteration.
## Architecture
```
┌─────────────────────────────────────────────────────────────┐
│ Bob (TCP Tuner Mode) │
│ "cwnd collapsing — switch to BBR, increase rmem_max" │
└──────────────────────────┬──────────────────────────────────┘
│ MCP tools (stdio)
┌──────────────────────────▼──────────────────────────────────┐
│ MCP Server (Node.js) │
│ run_benchmark │ get_sysctl_params │ apply_sysctl │
│ │ get_benchmark_history │
└──────────────────────────┬──────────────────────────────────┘
│ kubectl
┌──────────────────────────▼──────────────────────────────────┐
│ kind Kubernetes Cluster (3 nodes) │
│ │
│ ┌─────────────────┐ ┌──────────────────────────────┐ │
│ │ iperf3-server │◄───│ iperf3-client (Job) │ │
│ │ (Deployment) │ │ measures throughput/RTT │ │
│ └─────────────────┘ └──────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ sysctl-tuner DaemonSet (privileged, hostNetwork) │ │
│ │ worker-node-1 pod │ worker-node-2 pod │ │
│ │ reads/writes kernel TCP sysctl params │ │
│ └──────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
```
## Prerequisites
| Tool | Version | Install |
|------|---------|---------|
| Podman Desktop | 5+ | [podman-desktop.io](https://podman-desktop.io) |
| kubectl | v1.28+ | `winget install Kubernetes.kubectl` |
| kind | v0.20+ | `curl -Lo kind.exe https://kind.sigs.k8s.io/dl/v0.29.0/kind-windows-amd64` |
| Node.js | v22 LTS | `winget install OpenJS.NodeJS.LTS` |
| Python | 3.12+ | [python.org](https://python.org) |
## Quick Start
### 1. Spin up the cluster
```powershell
cd cluster
.\start.ps1
```
This creates a 3-node kind cluster (`1 control-plane + 2 workers`), deploys the iperf3 server, and starts the sysctl-tuner DaemonSet on both worker nodes.
### 2. Install and build the MCP server
```powershell
cd mcp-server
npm install
npm run build
```
### 3. Register the MCP server with Bob
Add to your Bob MCP config (`~/.bob/mcp-settings.json` or via Bob UI → Settings → MCP):
```json
{
"mcpServers": {
"tcp-tuner": {
"command": "node",
"args": ["<absolute-path-to-repo>/tcp-congestion-tuner/mcp-server/src/index.js"],
"cwd": "<absolute-path-to-repo>/tcp-congestion-tuner"
}
}
}
```
### 4. Load the custom Bob mode
Copy `.bob/custom_modes.yaml` to your Bob workspace config, or merge it into your existing custom modes file.
### 5. Start tuning
Open Bob, switch to **TCP Tuner** mode, and say:
> "Run a baseline benchmark, inspect the current sysctl configuration, and autonomously tune TCP parameters to maximise throughput."
Bob will run the full detect → diagnose → tune → validate loop.
---
## Manual CLI Usage
```powershell
# Run a benchmark
python scripts/benchmark.py run
# Show last 10 benchmark runs
python scripts/benchmark.py history
# Compare the last two runs
python scripts/benchmark.py compare
# Read current sysctl params
python scripts/benchmark.py sysctl get
# Apply a sysctl change manually
python scripts/benchmark.py sysctl set net.ipv4.tcp_congestion_control bbr
```
## MCP Tools Reference
| # | Tool | Description |
|---|------|-------------|
| 1 | `run_benchmark` | Launches iperf3 Job, returns throughput/RTT/retransmits + sysctl snapshot |
| 2 | `get_sysctl_params` | Reads all TCP sysctl values from a worker node |
| 3 | `apply_sysctl` | Writes a sysctl value on one or all worker nodes via privileged DaemonSet pod |
| 4 | `get_benchmark_history` | Returns past N runs for trend comparison |
| 5 | `inject_fault` | Injects packet loss + delay via `tc netem` on worker node interfaces |
| 6 | `clear_fault` | Removes netem qdiscs, restores clean network |
| 7 | `save_runbook` | Writes Markdown runbook to `runbooks/`, `git commit`, `git push`, returns SHA |
| 8 | `get_metrics` | Runs PromQL query against in-cluster Prometheus, returns live node metrics |
| 9 | `check_slo` | Evaluates last benchmark against SLO targets; returns pass/fail + next-action recommendation |
| 10 | `export_report` | Generates a self-contained HTML report with SVG trend charts from benchmark history |
| 11 | `benchmark_regression` | Compares latest run vs a golden baseline; returns pass/fail for use in CI |
## Observability
Deploy Prometheus + Grafana into the cluster:
```powershell
kubectl apply -f manifests/monitoring.yaml
```
Open the live dashboard (auto port-forwards and opens browser):
```powershell
.\cluster\port-forward.ps1
# Grafana: http://localhost:3000/d/tcp-tuner (admin / tcptuner)
# Prometheus: http://localhost:9090
```
The **TCP Congestion Tuner** dashboard shows:
- Network transmit/receive Mbps (all nodes)
- TCP retransmits/sec (kernel counter via node-exporter)
- CPU usage % (iperf3 load visibility)
## Known Limitations (kind cluster)
These are expected artifacts of running Kubernetes inside Podman containers — not bugs:
| Limitation | Explanation |
|---|---|
| Baseline retransmits ~1,000–2,000 on clean network | veth/bridge interfaces inside kind containers exhibit higher retransmit rates than bare-metal under high-throughput iperf3. The *reduction* from tuning is real; the absolute baseline is inflated. |
| `net.core.rmem_max` / `wmem_max` unavailable | These params are not namespaced and cannot be written from inside a container network namespace. `net.ipv4.tcp_rmem/wmem` are used instead. |
| `tc netem` fault affects control-plane pod too | kind's control-plane runs as a container on the same host network. Injecting fault on control-plane's eth0 can cause kubectl timeouts. Use `node_selector=worker` in production demos. |
| Throughput limited by loopback BDP | 50–75 Gbps is a loopback ceiling, not a real network limit. Buffer sizing changes show their full impact on real network paths with RTT > 1ms. |
## Key Kubernetes Concepts Demonstrated
- **DaemonSet** — sysctl-tuner runs on every worker node automatically
- **Privileged pods with hostNetwork** — required to read/write host kernel parameters
- **Job** — iperf3 client runs once and terminates cleanly
- **Headless Service** — iperf3 server addressable by DNS name within the cluster
- **Node-level sysctl tuning** — per-node kernel parameter management in K8s
## Tuning Playbook (What Bob Does)
| Change | Reason | Expected Impact |
|--------|--------|----------------|
| `tcp_congestion_control=bbr` | BBR tracks bottleneck bandwidth directly; better in high-BDP paths | +10–40% throughput, fewer retransmits |
| `rmem_max=134217728` | Larger receive buffers allow higher in-flight data | Higher throughput on high-latency links |
| `tcp_slow_start_after_idle=0` | Prevents cwnd reset after idle bursts | Important for bursty market-data feeds |
| `tcp_notsent_lowat=16384` | Reduces bufferbloat in the socket send queue | Lower RTT under load |
## Teardown
```powershell
cd cluster
.\teardown.ps1
```
## Bob Integration
This project is designed to be driven entirely by Bob in **TCP Tuner** mode. Three key prompts
cover the full lifecycle:
### Baseline + autonomous tuning
```
Run a baseline benchmark, read the current sysctl configuration, then autonomously tune TCP
parameters to maximise throughput. Show a before/after comparison table after each change.
Save a runbook to the runbooks/ folder when done.
```
### SLO-driven convergence loop
```
Tune until retransmits < 500 and throughput > 50000 Mbps. Run autonomously — apply one
sysctl change per iteration, benchmark, check SLO, repeat until SLO passes or playbook
exhausted. Save a runbook to git when done.
```
### Fault-injection resilience
```
Inject 3% packet loss. Run the fault-resilience tuning loop: benchmark under fault, tune
to compensate, clear the fault, confirm recovery, save runbook.
```
Bob uses the `tcp-tuner` skill (`.bob/skills/tcp-tuner/SKILL.md`) for structured agentic
best practices: 4 phases, 6 invariants, a full tuning playbook, and runbook templates.
---
## Project Structure
```
tcp-congestion-tuner/
├── cluster/
│ ├── kind-config.yaml # 1 control-plane + 2 worker nodes
│ ├── start.ps1 # cluster bootstrap script
│ ├── port-forward.ps1 # Grafana + Prometheus port-forward for demos
│ └── teardown.ps1 # cluster cleanup
├── manifests/
│ ├── iperf-server.yaml # iperf3 server Deployment + headless Service
│ ├── iperf-client-job.yaml # iperf3 benchmark Job
│ ├── tuner-daemonset.yaml # privileged sysctl-tuner DaemonSet (alpine + iproute2)
│ ├── monitoring.yaml # Prometheus + node-exporter DaemonSet + Grafana
│ └── alerts.yaml # TcpHighRetransmits + TcpThroughputDrop alert rules
├── mcp-server/
│ ├── package.json
│ ├── tsconfig.json
│ └── src/
│ └── index.ts # 9 MCP tools (TypeScript source)
├── scripts/
│ └── benchmark.py # CLI benchmark runner, comparator, rollback
├── tests/
│ └── test_benchmark.py # 119 pytest tests, 98% coverage
├── runbooks/ # Bob-generated runbooks (auto-committed by save_runbook)
├── .bob/
│ ├── custom_modes.yaml # TCP Tuner Bob mode with SLO convergence loop
│ └── skills/tcp-tuner/
│ └── SKILL.md # Reusable agentic best-practices skill
└── README.md
```
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues