Skip to main content
Glama
bhavyam2

Sentinel

by bhavyam2
README.md
# Sentinel

**A spend-policy and observability layer for [Agentcard](https://agentcard.sh)** — declarative
YAML policies, an enforcing MCP proxy, and an attribution/anomaly dashboard for AI agents
that hold real (or sandbox) money.

## The problem

Agentcard issues single-use virtual Visa cards to AI agents, and its own writing describes
exactly how that should be governed:

> "One card per task. Each task is a separate financial segment with its own card, its own
> budget, and its own blast radius." … "A task that should cost $12 gets a $15 card. Not a
> $100 card 'to be safe.'"
> — *[Financial Zero Trust for AI Agents](https://agentcard.sh/blog/financial-zero-trust-ai-agents)*

> "Agent expense systems need to attribute spend to tasks, workflows, and agent identities."
> … "Set velocity thresholds based on your historical patterns and alert when they are exceeded."
> — *[AI Agent Expense Management](https://agentcard.sh/blog/ai-agent-expense-management)*

The product today ships numeric limits: a budget, a per-transaction cap, and a human-approval
threshold. There is no merchant allowlist/denylist, no velocity limit, no time-window rule, no
per-agent policy scoping, and no attribution or anomaly layer — teams are told to build those
themselves. **Sentinel is that missing layer**, built the way the zero-trust post says it
should work:

- **Policy engine** — per-agent YAML rules (merchant allow/deny with wildcards, card velocity,
  allowed hours with timezones, per-card caps, daily budgets, approval thresholds) evaluated
  *before* any card exists, returning `ALLOW` / `DENY` / `NEEDS_APPROVAL` with a
  machine-readable reason chain of every rule checked, in order.
- **MCP proxy** — a stdio MCP server that wraps Agentcard's MCP server. Every tool passes
  through unchanged except `create_card` and `buy_checkout`, which are policy-checked first.
  Every decision — *including the denials Agentcard never sees* — lands in a SQLite ledger
  with agent id, task id, and the full rule trace.
- **Observability** — a webhook receiver joins Agentcard transaction events to the ledger, so
  every settled charge maps to *agent → task → policy decision*. One dashboard page: spend by
  agent and task, budget burn-down vs policy, the denied-spend log, and anomaly flags
  (daily-spend z-score, first-seen merchant, off-hours attempts).

## 60-second quickstart

```bash
git clone <this repo> sentinel && cd sentinel
./demo.sh        # creates a venv, seeds two agents, replays a scripted day,
                 # and serves the dashboard at http://127.0.0.1:8787/
```

No Agentcard account or credentials needed — the demo runs against a built-in mock that
implements Agentcard's documented API shapes (see `docs/NOTES.md`). Then try the killer
feature, a dry run that touches nothing:

```bash
.venv/bin/sentinel simulate --policy policies/example.yaml \
    --agent research-bot --amount 1800 --merchant OPENAI \
    --at "2026-07-15T23:30 America/Los_Angeles"
```

```
  ✓ merchant-deny        merchant 'OPENAI' matches no deny pattern
  ✓ merchant-allow       merchant 'OPENAI' matches allow pattern 'OPENAI'
  ✗ hours                local time 23:30 PDT is outside allowed window 06:00-22:00
  ...
Verdict: DENY
```

Exit codes: `0` ALLOW · `1` NEEDS_APPROVAL · `2` DENY. Add `--json` for machines,
`--state '{"spent_today_cents": 9000}'`-style files for what-ifs, or `--db sentinel.db`
to evaluate against real ledger history.

## Architecture

```mermaid
flowchart LR
    subgraph agent side
        A[AI agent / MCP client]
    end
    subgraph Sentinel
        P[MCP proxy<br/>sentinel proxy]
        E[Policy engine<br/>policies/*.yaml]
        L[(SQLite ledger<br/>decisions + transactions)]
        O[Observe server<br/>webhooks + dashboard]
    end
    U[Agentcard MCP server<br/>mcp.agentcard.sh<br/>or built-in mock]
    W[Agentcard webhooks<br/>transaction.*]

    A -- "create_card / buy_checkout<br/>(+ agent_id, task_id)" --> P
    A -- "all other tools" --> P
    P -- "evaluate" --> E
    P -- "ALLOW → forward<br/>(sentinel args stripped)" --> U
    P -- "every decision,<br/>incl. denials" --> L
    W -- "signed events" --> O
    O -- "join to decisions" --> L
    O -- "dashboard :8787" --> A
```

## Running it for real

1. Point the proxy at Agentcard (use a **sandbox** `sk_test_`-provisioned OAuth token; never
   start with live money):

   ```bash
   cp .env.example .env    # fill in AGENTCARD_MCP_URL + AGENTCARD_TOKEN
   sentinel proxy --policy policies/example.yaml --db sentinel.db
   ```

   Register that command as the MCP server in your agent's config *instead of* Agentcard's —
   e.g. in `claude_desktop_config.json` / `.mcp.json`:

   ```json
   {
     "mcpServers": {
       "agentcard": {
         "command": "/path/to/.venv/bin/sentinel",
         "args": ["proxy", "--policy", "policies/example.yaml", "--db", "sentinel.db"],
         "env": { "AGENTCARD_MCP_URL": "https://mcp.agentcard.sh/mcp",
                  "AGENTCARD_TOKEN": "…", "SENTINEL_AGENT_ID": "research-bot" }
       }
     }
   }
   ```

   Leave `AGENTCARD_MCP_URL`/`AGENTCARD_TOKEN` unset and the proxy uses the clearly-labeled
   mock — useful for CI and local development.

2. Run the observe server and register its URL as a webhook endpoint in the Agentcard
   dashboard (subscribe to `transaction.*`):

   ```bash
   AGENTCARD_WEBHOOK_SECRET=whsec_… sentinel observe --policy policies/example.yaml --db sentinel.db
   ```

The proxy adds three optional arguments to the intercepted tools: `agent_id` and `task_id`
(attribution, threaded into the ledger and joined to settled transactions) and
`merchant_hint` on `create_card` — Agentcard's `create_card` has **no merchant parameter**
(cards are open-loop until first charge), so the hint is the only pre-spend merchant control;
it is checked against policy and stripped before forwarding. A `NEEDS_APPROVAL` verdict
returns a hold; a human retries the call with `sentinel_approved: true` to release it.
`sentinel_approved` can never override a `DENY`.

## Policy reference

```yaml
defaults:                      # applied to agents not listed (omit = unknown agents denied)
  max_per_card_cents: 1000

agents:
  research-bot:
    max_per_card_cents: 2500           # hard cap per card / checkout
    daily_budget_cents: 10000          # forwarded spend per local calendar day
    velocity: { max_cards: 5, per: 1h }  # trailing-window card-creation limit (s/m/h/d)
    merchants:
      allow: ["OPENAI", "ANTHROPIC", "AWS"]   # if present, everything else is denied
      deny: ["*GAMBLING*", "*CRYPTO*"]        # deny always wins over allow
    hours: { allow: "06:00-22:00", tz: "America/Los_Angeles" }  # overnight windows OK
    require_approval_over_cents: 1500  # NEEDS_APPROVAL above this
```

Rule precedence (first deny wins; the full chain is always reported):
`merchant-deny → merchant-allow → hours → velocity → max-per-card → daily-budget →
approval-threshold`. Merchant patterns are case-insensitive; `*`/`?` are shell-style
wildcards against the whole descriptor, and plain patterns match the descriptor prefix or
any token prefix (`OPENAI` matches `OPENAI *CHATGPT SUBSCR` and `PAYPAL *OPENAI`).

## Tests

```bash
.venv/bin/pytest --cov=sentinel.policy   # engine is 96%+ covered
```

143 tests cover rule precedence, wildcard matching, timezone/DST/overnight-window edges,
velocity window boundaries, ledger joins and dedupe, webhook signatures against known HMAC
vectors, and an end-to-end run of the proxy over a real in-memory MCP session.

## What this doesn't do (honest edition)

- **It is not a network-level control.** Sentinel gates what flows through *its* proxy. An
  agent holding raw Agentcard credentials — or a card number already minted — can spend
  without Sentinel ever knowing. Pair it with Agentcard's own budget/limit settings as the
  backstop; Sentinel is the fine-grained layer, not the outer wall.
- **Merchant rules on `create_card` are advisory-by-construction.** Agentcard's API takes no
  merchant at card-creation time, so `merchant_hint` is only as truthful as the agent
  supplying it. The hint is still valuable (it's checked, logged, and auditable against the
  settled descriptor later), and `buy_checkout` merchant checks are real. Post-hoc, webhook
  descriptors expose any mismatch on the dashboard.
- **Merchant matching is string matching.** `AWS` will not match a descriptor that reads
  `AMAZON WEB SERVICES`; card-network descriptors are messy. Write patterns against
  descriptors you've actually seen (the dashboard shows raw ones).
- **Approvals are honor-system at the MCP boundary.** `sentinel_approved: true` is meant to
  be attached by a human-in-the-loop client, not the agent itself. If your agent framework
  can't guarantee that, treat NEEDS_APPROVAL as DENY.
- **Anomaly detection is deliberately simple** (z-score with ≥3 days history, first-seen
  merchant, off-hours). A constant-spend history (σ=0) yields no z-flag; day one yields no
  flags at all. It's a tripwire, not a fraud model.
- **The live-API path is untested against production.** Card/transaction operations are
  MCP-only (no public REST spec), and exercising the real server requires interactive OAuth
  onboarding — so CI runs against `MockAgentcardClient`, which mirrors the documented shapes
  in `docs/NOTES.md`. The policy engine, ledger, CLI, webhook receiver, and dashboard are
  fully real regardless of upstream.
- **Single-process ledger.** SQLite (WAL) is plenty for one proxy + one observe server; it is
  not a multi-region audit store.

## License

MIT — see [LICENSE](LICENSE).

TDQS

B3.3/5.0

Scored across 10 tools

Disambiguation5/5

Each tool targets a distinct action: card management (create, list, close, get details), transaction queries (single card vs all), balance check, and purchasing flow (add to cart, view cart, checkout). The 'buy_' prefix clearly separates merchant actions from card operations, eliminating ambiguity.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern using snake_case (e.g., list_cards, close_card, buy_add_to_cart). The only minor variance is 'check_balance' instead of 'get_balance', but it remains a clear verb_noun structure, maintaining high consistency.

Tool Count5/5

With 10 tools, the server covers the essential operations for virtual card management and merchant purchasing without being overwhelming. Each tool serves a necessary function in the workflow, making the count well-scoped for the domain.

Completeness3/5

The tool set covers core card lifecycle and basic transactions, but lacks merchant management (e.g., linking merchants), wallet funding, and cart modification (remove/update items). These gaps may cause agents to get stuck when the workflow requires actions outside the provided surface.

Maintenance

ActivityMaintained
ResponsivenessSyncing