Skip to main content
Glama
README.md
# Nemotron Scout

**An autonomous opportunity-discovery agent, exposed to Alexa+ as a self-hosted MCP server.**

Ask it out loud *"what can I earn from this week?"* and it reads you back the
live, ranked shortlist — with the effort, the realistic time to first money, and a
ready-to-send message for the one you pick.

Built for the **Amazon Developer Hackathon (Build, Ship, Shape)**, Alexa+ track.
Runs on any OpenAI-compatible model provider; the shipped default uses
**NVIDIA Nemotron 3**.

## Demo video

[▶ Watch on YouTube](https://www.youtube.com/watch?v=T9_shkAgB3s) — 2 min 20 s,
in English, unedited real run. Screens in `demo/`, rebuild with
`python demo/make_video.py`.

---

## The problem

Money-relevant opportunities are scattered across places nobody monitors: a
bounty posted in a GitHub issue, a hackathon that opens quietly, a paid issue on
a job board. Reading them all is unpaid work, so people read none of them.

## What Scout does

A four-stage agent pipeline turns raw public signals into a decision:

```
  collect          extract           critique          plan
┌───────────┐   ┌──────────────┐   ┌─────────────┐   ┌──────────────┐
│ Devpost   │   │ Nemotron     │   │ Nemotron    │   │ Nemotron     │
│ HackerN.  │──▶│ 3 Nano/Omni  │──▶│ 3 Super     │──▶│ 3 Super     │
│ RemoteOK  │   │ typed JSON   │   │ kills hype, │   │ steps +      │
│ GitHub    │   │ extraction   │   │ adjusts     │   │ outreach +   │
└───────────┘   └──────────────┘   │ score       │   │ kill criteria│
                                  └─────────────┘   └──────────────┘
                                        │
                                        ▼
                            deterministic 0-100 score
                            speed · effort · competition
                            evidence · confidence
```

The split is deliberate: **the model judges, the arithmetic decides.** Every
point of a score is explainable, and a rerun on the same inputs produces the same
ranking.

The critic is adversarial on purpose. It sees only extracted fields, never the
extractor's reasoning, and it is the stage that throws out enthusiasm dressed up
as evidence.

## Why it is an MCP server

Alexa+ integrations are built on MCP. This repo ships a real MCP server over
**Streamable HTTP** (spec `2025-11-25`), not a wrapper around one:

| Method | Behaviour |
| --- | --- |
| `initialize` | negotiates `protocolVersion`, issues `Mcp-Session-Id` |
| `tools/list` | 4 tools with full JSON Schema |
| `tools/call` | runs a tool, returns text + `structuredContent` |
| `resources/read` | latest report as a markdown resource |
| `ping` | keepalive |
| notifications | answered `202 Accepted`, empty body |

### Tools

| Tool | Purpose |
| --- | --- |
| `scout_opportunities` | Ranked paid opportunities, filterable by kind and score |
| `scout_action_plan` | Steps, first-hour checklist, ready-to-send message, kill criteria |
| `scout_voice_answer` | Spoken-friendly answer: no markdown, no URLs, safe to read aloud |
| `scout_status` | Providers, models, collector health, last run |

## Quickstart

Zero heavyweight dependencies — Python 3.10+ and `requests`.

```bash
pip install -r requirements.txt

export OPENROUTER_API_KEY=sk-or-v1-...
# or any OpenAI-compatible provider:
#   export SCOUT_PROVIDER=bedrock  AWS_REGION=us-east-1  AWS_BEARER_TOKEN_BEDROCK=...
#   export SCOUT_PROVIDER=nebius   NEBIUS_API_KEY=...

python -m scout.cli collect --limit 12   # raw signals, no model calls
python -m scout.cli run --limit 24 --top 3
python -m unittest discover -s tests     # 8 tests, no network
```

### Run the MCP server

```bash
python -m scout.mcp_server --port 8765
#   MCP endpoint   http://127.0.0.1:8765/mcp
#   Health         http://127.0.0.1:8765/healthz
#   Demo UI        http://127.0.0.1:8765/
```

The bundled page at `/` is the simulated Alexa+ experience: it speaks to the
server over the same MCP protocol a real Alexa+ client would use, so the demo
shows protocol traffic rather than a mock.

### Exposing it publicly

A tool call spends real LLM credits, so bind a token whenever the server is not
on loopback:

```bash
SCOUT_MCP_TOKEN=$(python -c 'import secrets;print(secrets.token_urlsafe(24))') \
  python -m scout.mcp_server --port 8765            # token required on /mcp

cloudflared tunnel --url http://127.0.0.1:8765       # quick tunnel, prints a trycloudflare.com URL
```

`/healthz` and `/` stay open so a probe and the demo page keep working; only
`/mcp` needs the header:

```bash
curl -sN -X POST https://<tunnel>/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'MCP-Protocol-Version: 2025-11-25' \
  -H "Authorization: Bearer $SCOUT_MCP_TOKEN" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize",
       "params":{"protocolVersion":"2025-11-25","clientInfo":{"name":"probe"}}}'
```

The server warns on startup if you bind to a public interface without a token.

#### Letting a judge in without handing over the key

A token nobody has is the same as no demo. `--public-readonly` splits the two
concerns: a caller with no token can connect, list tools and read the **cached**
run, but only a token holder can start a live search that spends credits.

```bash
python -m scout.mcp_server --port 8765 --token "$SCOUT_MCP_TOKEN" --public-readonly
```

```
$ curl -s -X POST https://<tunnel>/mcp -d '{"jsonrpc":"2.0","id":1,
    "method":"tools/call","params":{"name":"scout_opportunities","arguments":{}}}'
Read-only public access: this is the last real run, not a live search.
6 ranked opportunities.
...
```

With no cache on disk the call is refused with `-32001` and a message saying a
token is required, rather than quietly spending a run. `scout_status`,
`scout_action_plan` and `scout_voice_answer` are free of charge, so they are
served either way.

The cached run is loaded at startup, so a restart does not present a judge with
an empty server.

### Free-tier limits are real

OpenRouter's free tier is capped at **50 requests per day**, and one pipeline
run costs 7 or more. When the cap is hit the server does not simply fail: it
serves the last good run from `reports/last_run.json` and says so in the tool
text, so a demo never dead-ends. Point `SCOUT_CACHE_FILE` elsewhere to move that
file.

### Probe it with curl

```bash
curl -sN -X POST http://127.0.0.1:8765/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize",
       "params":{"protocolVersion":"2025-11-25","capabilities":{},
                 "clientInfo":{"name":"probe","version":"1"}}}'
# -> data: {...}  and header  Mcp-Session-Id: scout-...
```

## Configuration

Everything is an environment variable; see `.env.example`.

| Variable | Default | Meaning |
| --- | --- | --- |
| `SCOUT_PROVIDER` | `openrouter` | `openrouter` / `bedrock` / `nebius` |
| `SCOUT_API_KEY` | — | Overrides the provider's own key variable |
| `SCOUT_MODEL_EXTRACT` | Nemotron 3 Nano/Omni (free) | Extraction model |
| `SCOUT_MODEL_REASON` | Nemotron 3 Super 120B (free) | Critic and planner model |
| `SCOUT_COLLECT_LIMIT` | `24` | Signals gathered per run |
| `SCOUT_SHORTLIST` | `8` | Items surviving extraction |
| `SCOUT_PLAN_TOP` | `3` | Items that get a full execution plan |
| `SCOUT_BASE_URL` | per provider | Any OpenAI-compatible host, comma-separated |
| `SCOUT_MAX_TOKENS` | `1200` | Output budget per call; raise it for reasoning models |
| `SCOUT_MAX_CONCURRENCY` | `4` | Parallel model calls; lower it for tight per-minute limits |
| `SCOUT_MCP_TOKEN` | — | Bearer token required on `/mcp` |
| `SCOUT_CACHE_FILE` | `reports/last_run.json` | Where the last good run is cached |

Model calls fall back through a chain of alternatives, so a rate-limited or
unavailable model degrades one stage instead of failing the run. The chain is
OpenRouter-specific on purpose: those slugs mean nothing on another host, where
they would only 404 and bury the real error.

Two details that cost real debugging time and are now handled for you:

- **Reasoning models need a real token budget.** `gpt-oss-120b` spends part of
  its allowance on hidden thinking; with a low provider default the visible
  content comes back empty and the parse fails, which reads like a schema bug
  rather than a truncation. `SCOUT_MAX_TOKENS` is always sent explicitly.
- **Rate limits are honoured.** A `429` is retried using the provider's
  `Retry-After`, with a longer default pause, because a per-minute budget does
  not recover in two seconds.

## Design notes

- **Free-tier by default.** The shipped OpenRouter defaults are `:free` NVIDIA
  models, so a fresh account with `$0` credits can run the full pipeline. The
  trade-off is the 50 requests/day cap; see the caching note above.
- **One run, many callers.** Tool calls are single-flighted, so a burst of
  parallel requests triggers one pipeline run, not one per request. A public
  endpoint cannot multiply your bill.
- **Partial model output is tolerated.** The extractor may return a bare JSON
  array, omit `kind`, or send `"12h"` where a number was expected; the schema
  layer normalises all of it instead of raising.
- **Degrades to the last good run.** If the backend is unavailable, the server
  replays the cached run and labels it as cached rather than inventing data.
- **Collectors never take the run down.** A failing source is reported in
  `collectors` and contributes nothing.
- **No LLM at all in offline mode.** `--offline` collects and reports, which is
  what CI and the unit tests use.
- **The score is not the model.** `scout/scoring.py` is pure arithmetic with
  documented priors, unit-tested separately from any network call.

## Project layout

```
scout/
  config.py     provider presets, env resolution
  llm.py        OpenAI-compatible client, JSON repair, model fallbacks
  sources.py    Devpost / Hacker News / RemoteOK / GitHub collectors
  agents.py     the three prompts and their strict JSON contracts
  scoring.py    deterministic 0-100 ranking
  pipeline.py   orchestration, markdown + JSON reporting
  mcp_server.py MCP over Streamable HTTP
  server.py     plain JSON/HTML server
  cli.py        collect | models | run
web/index.html  simulated Alexa+ demo
tests/          offline unit tests
```

## License

MIT — see `LICENSE`.