Skip to main content
Glama

Nemotron Scout

An autonomous opportunity-discovery agent, exposed to Alexa+ as a self-hosted MCP server.

Ask it out loud "what can I earn from this week?" and it reads you back the live, ranked shortlist — with the effort, the realistic time to first money, and a ready-to-send message for the one you pick.

Built for the Amazon Developer Hackathon (Build, Ship, Shape), Alexa+ track. Runs on any OpenAI-compatible model provider; the shipped default uses NVIDIA Nemotron 3.

Demo video

▶ Watch on YouTube — 2 min 20 s, in English, unedited real run. Screens in demo/, rebuild with python demo/make_video.py.


Related MCP server: bounty-radar

The problem

Money-relevant opportunities are scattered across places nobody monitors: a bounty posted in a GitHub issue, a hackathon that opens quietly, a paid issue on a job board. Reading them all is unpaid work, so people read none of them.

What Scout does

A four-stage agent pipeline turns raw public signals into a decision:

  collect          extract           critique          plan
┌───────────┐   ┌──────────────┐   ┌─────────────┐   ┌──────────────┐
│ Devpost   │   │ Nemotron     │   │ Nemotron    │   │ Nemotron     │
│ HackerN.  │──▶│ 3 Nano/Omni  │──▶│ 3 Super     │──▶│ 3 Super     │
│ RemoteOK  │   │ typed JSON   │   │ kills hype, │   │ steps +      │
│ GitHub    │   │ extraction   │   │ adjusts     │   │ outreach +   │
└───────────┘   └──────────────┘   │ score       │   │ kill criteria│
                                  └─────────────┘   └──────────────┘
                                        │
                                        ▼
                            deterministic 0-100 score
                            speed · effort · competition
                            evidence · confidence

The split is deliberate: the model judges, the arithmetic decides. Every point of a score is explainable, and a rerun on the same inputs produces the same ranking.

The critic is adversarial on purpose. It sees only extracted fields, never the extractor's reasoning, and it is the stage that throws out enthusiasm dressed up as evidence.

Why it is an MCP server

Alexa+ integrations are built on MCP. This repo ships a real MCP server over Streamable HTTP (spec 2025-11-25), not a wrapper around one:

Method

Behaviour

initialize

negotiates protocolVersion, issues Mcp-Session-Id

tools/list

4 tools with full JSON Schema

tools/call

runs a tool, returns text + structuredContent

resources/read

latest report as a markdown resource

ping

keepalive

notifications

answered 202 Accepted, empty body

Tools

Tool

Purpose

scout_opportunities

Ranked paid opportunities, filterable by kind and score

scout_action_plan

Steps, first-hour checklist, ready-to-send message, kill criteria

scout_voice_answer

Spoken-friendly answer: no markdown, no URLs, safe to read aloud

scout_status

Providers, models, collector health, last run

Quickstart

Zero heavyweight dependencies — Python 3.10+ and requests.

pip install -r requirements.txt

export OPENROUTER_API_KEY=sk-or-v1-...
# or any OpenAI-compatible provider:
#   export SCOUT_PROVIDER=bedrock  AWS_REGION=us-east-1  AWS_BEARER_TOKEN_BEDROCK=...
#   export SCOUT_PROVIDER=nebius   NEBIUS_API_KEY=...

python -m scout.cli collect --limit 12   # raw signals, no model calls
python -m scout.cli run --limit 24 --top 3
python -m unittest discover -s tests     # 8 tests, no network

Run the MCP server

python -m scout.mcp_server --port 8765
#   MCP endpoint   http://127.0.0.1:8765/mcp
#   Health         http://127.0.0.1:8765/healthz
#   Demo UI        http://127.0.0.1:8765/

The bundled page at / is the simulated Alexa+ experience: it speaks to the server over the same MCP protocol a real Alexa+ client would use, so the demo shows protocol traffic rather than a mock.

Exposing it publicly

A tool call spends real LLM credits, so bind a token whenever the server is not on loopback:

SCOUT_MCP_TOKEN=$(python -c 'import secrets;print(secrets.token_urlsafe(24))') \
  python -m scout.mcp_server --port 8765            # token required on /mcp

cloudflared tunnel --url http://127.0.0.1:8765       # quick tunnel, prints a trycloudflare.com URL

/healthz and / stay open so a probe and the demo page keep working; only /mcp needs the header:

curl -sN -X POST https://<tunnel>/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'MCP-Protocol-Version: 2025-11-25' \
  -H "Authorization: Bearer $SCOUT_MCP_TOKEN" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize",
       "params":{"protocolVersion":"2025-11-25","clientInfo":{"name":"probe"}}}'

The server warns on startup if you bind to a public interface without a token.

Letting a judge in without handing over the key

A token nobody has is the same as no demo. --public-readonly splits the two concerns: a caller with no token can connect, list tools and read the cached run, but only a token holder can start a live search that spends credits.

python -m scout.mcp_server --port 8765 --token "$SCOUT_MCP_TOKEN" --public-readonly
$ curl -s -X POST https://<tunnel>/mcp -d '{"jsonrpc":"2.0","id":1,
    "method":"tools/call","params":{"name":"scout_opportunities","arguments":{}}}'
Read-only public access: this is the last real run, not a live search.
6 ranked opportunities.
...

With no cache on disk the call is refused with -32001 and a message saying a token is required, rather than quietly spending a run. scout_status, scout_action_plan and scout_voice_answer are free of charge, so they are served either way.

The cached run is loaded at startup, so a restart does not present a judge with an empty server.

Free-tier limits are real

OpenRouter's free tier is capped at 50 requests per day, and one pipeline run costs 7 or more. When the cap is hit the server does not simply fail: it serves the last good run from reports/last_run.json and says so in the tool text, so a demo never dead-ends. Point SCOUT_CACHE_FILE elsewhere to move that file.

Probe it with curl

curl -sN -X POST http://127.0.0.1:8765/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize",
       "params":{"protocolVersion":"2025-11-25","capabilities":{},
                 "clientInfo":{"name":"probe","version":"1"}}}'
# -> data: {...}  and header  Mcp-Session-Id: scout-...

Configuration

Everything is an environment variable; see .env.example.

Variable

Default

Meaning

SCOUT_PROVIDER

openrouter

openrouter / bedrock / nebius

SCOUT_API_KEY

—

Overrides the provider's own key variable

SCOUT_MODEL_EXTRACT

Nemotron 3 Nano/Omni (free)

Extraction model

SCOUT_MODEL_REASON

Nemotron 3 Super 120B (free)

Critic and planner model

SCOUT_COLLECT_LIMIT

24

Signals gathered per run

SCOUT_SHORTLIST

8

Items surviving extraction

SCOUT_PLAN_TOP

3

Items that get a full execution plan

SCOUT_BASE_URL

per provider

Any OpenAI-compatible host, comma-separated

SCOUT_MAX_TOKENS

1200

Output budget per call; raise it for reasoning models

SCOUT_MAX_CONCURRENCY

4

Parallel model calls; lower it for tight per-minute limits

SCOUT_MCP_TOKEN

—

Bearer token required on /mcp

SCOUT_CACHE_FILE

reports/last_run.json

Where the last good run is cached

Model calls fall back through a chain of alternatives, so a rate-limited or unavailable model degrades one stage instead of failing the run. The chain is OpenRouter-specific on purpose: those slugs mean nothing on another host, where they would only 404 and bury the real error.

Two details that cost real debugging time and are now handled for you:

  • Reasoning models need a real token budget. gpt-oss-120b spends part of its allowance on hidden thinking; with a low provider default the visible content comes back empty and the parse fails, which reads like a schema bug rather than a truncation. SCOUT_MAX_TOKENS is always sent explicitly.

  • Rate limits are honoured. A 429 is retried using the provider's Retry-After, with a longer default pause, because a per-minute budget does not recover in two seconds.

Design notes

  • Free-tier by default. The shipped OpenRouter defaults are :free NVIDIA models, so a fresh account with $0 credits can run the full pipeline. The trade-off is the 50 requests/day cap; see the caching note above.

  • One run, many callers. Tool calls are single-flighted, so a burst of parallel requests triggers one pipeline run, not one per request. A public endpoint cannot multiply your bill.

  • Partial model output is tolerated. The extractor may return a bare JSON array, omit kind, or send "12h" where a number was expected; the schema layer normalises all of it instead of raising.

  • Degrades to the last good run. If the backend is unavailable, the server replays the cached run and labels it as cached rather than inventing data.

  • Collectors never take the run down. A failing source is reported in collectors and contributes nothing.

  • No LLM at all in offline mode. --offline collects and reports, which is what CI and the unit tests use.

  • The score is not the model. scout/scoring.py is pure arithmetic with documented priors, unit-tested separately from any network call.

Project layout

scout/
  config.py     provider presets, env resolution
  llm.py        OpenAI-compatible client, JSON repair, model fallbacks
  sources.py    Devpost / Hacker News / RemoteOK / GitHub collectors
  agents.py     the three prompts and their strict JSON contracts
  scoring.py    deterministic 0-100 ranking
  pipeline.py   orchestration, markdown + JSON reporting
  mcp_server.py MCP over Streamable HTTP
  server.py     plain JSON/HTML server
  cli.py        collect | models | run
web/index.html  simulated Alexa+ demo
tests/          offline unit tests

License

MIT — see LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to find and query real-time GitHub coding bounties with built-in scam filtering, supporting listing, matching, and detailed bounty retrieval.
    4
    42 npm
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides MCP tools to search and retrieve open-source bounties, jobs, and challenges from a live feed.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables developers and AI agents to discover, rank, watch, and draft submissions for Gibwork coding bounties through MCP tools, with optional wallet-authenticated SDK support.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables any MCP-capable AI app to act as a personal job-search assistant by storing resumes and job postings, matching skills, tracking applications, and analyzing market skill demand.
    MIT