Skip to main content
Glama
martinvibes

skin-in-the-game

by martinvibes

Skin

skin in the game

An AI analyst that has to bet its own money on every call it makes.

Its track record is not a claim it writes about itself. It is settled positions on Binance prediction markets, scored with a rule it cannot game.

Live dashboard → · 2-minute demo →

Binance Agent OS Agentic Wallet MCP Server ci tests TypeScript license


The problem

Every AI agent will tell you what it thinks the market will do. None of them pay for being wrong.

An agent can say "BTC looks bullish, roughly 80% confidence" a hundred times. It is never scored, so the confidence number means nothing, and it costs nothing to inflate. This is the single largest unsolved problem with AI financial agents: there is no price on being wrong, so there is no information in anything they say.

Related MCP server: pmxt-mcp

The fix

Binance Agent OS shipped prediction markets to the Agentic Wallet. A prediction market quotes belief directly: an outcome token at $0.62 is a 62% probability. That makes it the only venue where an agent's forecast and its money are the same object.

So this agent is not allowed to have an opinion for free. It prices a market from realized volatility (a closed-form model, no LLM guess), and if it disagrees by more than 4 points it sizes a stake by the Kelly criterion and buys with real USDT. The conviction goes into an append-only journal before the outcome is known. When the market resolves, the call is scored by Brier score, a strictly proper rule, so the agent's best possible strategy is to state what it actually believes. It sweeps its own winnings too, because prediction markets do not pay out automatically and an agent that never claims goes broke while being right.

Every number on the dashboard traces back to a settled on-chain position. The agent does not get a vote in its own performance review.


What it looks like

The console

The dashboard is not a report you scroll. It is the agent's loop with a button on it. Run a scan prices five markets one at a time and says out loud why it refuses most of them; Stake makes you type the word stake before any money moves, exactly as the CLI does; Claim sweeps the winnings that prediction markets do not pay out on their own. Every number in it comes from the scan block of record.json, which the CLI produced by calling the same formOpinion and sizeStake a live run calls, so the demo cannot quietly disagree with the tool it is demonstrating.

The record

  SKIN  skin in the game   ● DEMO  synthetic fixtures · no wallet, no money moved

THE RECORD
computed from settled positions · the agent does not get a vote
──────────────────────────────────────────────────────────────────────────
  realized PnL                +$0.07  on $13.20 staked
  return on stake               0.5%

  Brier score                  0.261  roughly as useful as always saying 50%
                                      0 perfect · 0.25 = always saying "50%" · lower is better

  settled calls                   12
  won / lost                   7 / 5
  hit rate                     58.3%
──────────────────────────────────────────────────────────────────────────

That is the bundled demo record, and it is deliberately unflattering: up seven cents on $13.20, a Brier score barely better than a coin flip, and the biggest single loss is the call it was 88% sure about. A demo that showed a winning agent would be showing you the marketing, not the mechanism.

The point of the project is that this number is not editable. So it is not edited.

A betting slip

Every stake prints a receipt before any money moves: the full derivation, and which of the four ceilings actually bound the size.

┌┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┐
┆ Bitcoin Up or Down on September 8?                            PENDING  ┆
┆ side NO                                                                ┆
┆┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┆
┆ agent says            83.4%                                            ┆
┆ market says           74.0%                                            ┆
┆ edge                   9.4%                                            ┆
┆ payout odds           0.35:1                                           ┆
┆ full Kelly            36.2%                                            ┆
┆ applied (¼ Kelly)      9.0%                                            ┆
┆ bound by              per-call-cap                                     ┆
┆┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┆
┆ STAKE                 $1.00                                            ┆
┆┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┆
┆ BTC spot $78,426, realized vol 27% (200 × 1m candles). Zero-drift      ┆
┆ lognormal over 4.0 h puts P(BTC above $78,861) at 16.6%; the book      ┆
┆ quotes the NO token at 75.0% for this size.                            ┆
└┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┘

That is a real receipt. It filled as order 26090800001869670157 for 1.32 shares at $0.7545.

A scan

SCAN
12 market(s) · bankroll $14.99 · run cap $1.00 · ¼ Kelly
──────────────────────────────────────────────────────────────────────────
  side it would buy @ price it would pay · fair = what the model thinks it is worth
  pass      BNB Up or Down · 7:45AM-8AM ET            NO @ 22% · fair 37%    below-minimum
  pass      Ethereum Up or Down · 7AM ET              NO @ 99% · fair 100%   conviction-bounds
  pass      Bitcoin Up or Down · 7:45AM ET            NO @ 51% · fair 52%    no-edge
  pass      Bitcoin Up or Down · 7AM ET               NO @ 98% · fair 100%   conviction-bounds
  BET       Bitcoin Up or Down on September 8?        NO @ 74% · fair 83%    $1.00
  pass      Ethereum Up or Down on September 8?       NO @ 47% · fair 48%    no-edge
  pass      BNB Up or Down 1d                         YES @ 85% · fair 99%   conviction-bounds
──────────────────────────────────────────────────────────────────────────

Eleven of twelve live markets refused, each with a named reason. That is the system working, not the system failing to find anything.

The dashboard

tryskin.vercel.app

A React page rendering the same record.json the CLI exports. No backend and no key: it reads one file, so what it shows is exactly what the CLI produced, and it could not place an order if it wanted to. The mode badge is permanent, so a demo number can never be mistaken for a live one. Both charts are hand-drawn SVG; the whole page gzips to 70 kB.

The equity curve

Calibration

npm --prefix web install && npm run web:dev

The dashboard lives in web/, and this repo has two package.json files: the CLI's at the root and the web app's in web/. Set the Vercel project's Root Directory to web, or the build runs the root package.json's build script (which compiles the CLI) and deploys nothing servable. The symptom is a successful build followed by a 404 on every route.


Quick start

No wallet, no money, no signup. The demo rail is fully offline and is what a reviewer should run first:

git clone https://github.com/martinvibes/skin-in-the-game
cd skin-in-the-game
npm install

npm run skin -- record --demo     # the track record
npm run skin -- scan   --demo     # opinions and refusals, no money
npm run skin -- stake  --demo     # receipts + typed confirmation
npm run skin -- claim  --demo     # sweep unclaimed winnings
npm test                          # 68 tests
npm --prefix web install && npm run web:dev    # the dashboard

Going live

Live needs a Binance Agentic Wallet (created in the Binance mobile app; an agent cannot create one for you) and a little USDT on BSC. Stakes are about $1, so $10 to $15 runs this for real.

npx skills add binance/binance-skills-hub/skills/binance-web3/binance-agentic-wallet
npm install -g @binance/agentic-wallet
baw auth signin                   # approve on your phone

Once baw is on your PATH, live is the default and --demo is the opt-out. --live forces it and fails loudly if the CLI is missing, so live is never entered by accident either.

npm run skin -- doctor            # wallet, session, quota, balance
npm run skin -- scan              # the real book, re-priced against real quotes
npm run skin -- stake --budget 1 --per-call 1     # the only command that spends
npm run skin -- record            # settled record, calibration, open positions
npm run skin -- claim             # redeem settled wins to the wallet

Start with doctor. It probes every wallet call the staking path depends on and names the one that would have failed, which is a cheaper way to find a signed-out session than discovering it halfway through an order.

Prediction trading has its own daily quota in the Binance app, separate from the wallet's general limit. doctor reads it, and sizing treats it as a hard ceiling.


Binance Agent OS surfaces used

Surface

How this project uses it

Agentic Wallet (baw)

The entire trading path. prediction market list / detail / last-trade-price to read markets, prediction trade place-order to stake, prediction trade redeem to claim, prediction position list --tab ONGOING|PENDING_CLAIM and settled-history to build the record, wallet balance for the bankroll. See src/adapters/baw.ts.

MCP, as a server

Skin is itself an MCP server. skin mcp exposes eight read-only tools so another agent can query this one's record, calibration and current book, and cannot spend its money. See src/mcp/server.ts and the section below.

Binance MCP Server

Market data for the pricing model. An MCP-authenticated agent fetches klines and hands them over with --mcp-data, so volatility is computed under the user's own Agent OS session. See src/adapters/marketdata.ts and the walk-through below.

Skills Hub format

skill/SKILL.md is a publishable skill in the Hub's frontmatter format, with a references/ split. It is what turns the CLI into an agent workflow: a policy the agent has to follow, not just a binary it can call.

MPC Wallet limits

The wallet's own daily spend limit is read and treated as a hard ceiling on sizing, so the agent is structurally unable to exceed a boundary the human set in the Binance app.

The wallet's daily limit as a sizing input is the part worth pausing on. Most agent projects treat wallet limits as an error to handle. Here it is one of the four terms in the position-sizing minimum, on equal footing with Kelly.

Using the MCP Server

claude mcp add binance-mcp-server --transport http https://agent.binance.com/mcp/agentic

The agent fetches candles under its own authenticated session, writes them to a file as { "BTCUSDT": { "interval": "1m", "closes": [...] } }, and hands the path over with skin scan --mcp-data ./klines.json. A committed fixture exercises the path with no session and no network:

npm run skin -- scan --demo --mcp-data fixtures/klines.example.json

Two of those fixture markets have a strike sitting exactly on spot, and the model prices both at exactly 50%: the cheapest available check that the lognormal is wired up correctly. At the money, over any horizon, a zero-drift walk is a coin flip.

Without --mcp-data the CLI falls back to Binance's public market-data REST endpoint, which needs no credentials. Same arithmetic either way. It tries four hosts in order, including data-api.binance.vision, because api.binance.com does not resolve from every country this was developed in.


Skin as an MCP server

The agent keeping this record should not be the only one able to read it. skin mcp speaks MCP over stdio, so any host (Claude Code, Claude Desktop, Cursor, another desk's agent) can interrogate this one directly instead of trusting a screenshot of it.

Eight tools. None of them spends.

Running it

Zero setup, in this repo. A .mcp.json is committed, so a clone opened in Claude Code offers the server on its own. Approve it at the prompt and the tools are there.

One line, from the repo root:

npm install
claude mcp add skin -- npx tsx src/cli/index.ts mcp --demo

Any other MCP host (Claude Desktop, Cursor, an agent you wrote yourself) launches the same process. Give it an absolute path, because the host's working directory is not yours:

{
  "mcpServers": {
    "skin": {
      "command": "npx",
      "args": [
        "tsx",
        "/absolute/path/to/skin-in-the-game/src/cli/index.ts",
        "mcp",
        "--demo"
      ]
    }
  }
}

Run npm install in the clone once and that is the whole install. No API key, no wallet, no build step: --demo serves fixtures that ship in the source, so the tools answer with the network off.

One wrinkle worth knowing before you blame the config: launched from outside the repo, npx cannot see the local tsx and fetches its own copy, which takes a few seconds and can trip a host's first connect timeout. It is cached after that, so reconnecting works. Hosts that support a cwd field avoid it entirely by pointing at the clone.

Pointing it at the real wallet is one flag less. Drop --demo and the tools read live markets, the live journal and settled positions through baw, which requires baw auth signin to have happened. It stays read-only either way:

"args": ["tsx", "/absolute/path/to/skin-in-the-game/src/cli/index.ts", "mcp"]

Sizing defaults for skin_propose are passed exactly as the CLI takes them, after the mcp argument:

npx tsx src/cli/index.ts mcp --limit 12 --budget 6 --per-call 1.5 --kelly 0.25

Knowing it came up. Every byte of stdout belongs to the JSON-RPC stream, so the server greets you on stderr, where the host will show it and the protocol will never read it:

  SKIN MCP  demo mode · listening on stdio

    skin_record       settled calls, PnL, Brier score
    skin_calibration  where the numbers break down
    skin_slips        convictions written before the outcome
    skin_positions    money committed, not yet resolved
    skin_unclaimed    settled wins still on-chain
    skin_scan         every open market, with the working
    skin_opinion      one market: vol, horizon, probability
    skin_propose      a sized stake, and the command a human must run

  Eight tools, none of which can spend. There is deliberately no
  skin_stake: an agent can reason all the way to a position and still
  cannot open one.

  stdout is the protocol, so nothing more will be printed here.
  Waiting for a client. Ctrl-C to stop.

Silence after that line is what success looks like.

What the calling agent gets

Tool

Answers

skin_record

Settled calls, hit rate, realized PnL, ROI, Brier score

skin_calibration

Reliability curve: where the agent's numbers break down

skin_slips

Every conviction written down before the outcome was known

skin_positions

Money committed, not yet resolved, with the conviction behind it

skin_unclaimed

Settled wins still sitting on-chain

skin_scan

A view on every open market, with the full working

skin_opinion

One market: spot, volatility, horizon, probability, edge

skin_propose

A sized stake, and the command a human must run to place it

Each tool returns a one-line headline a model can act on, then the JSON behind it, so a host that only shows the first line still shows something true. In demo mode every headline is tagged [demo mode: synthetic fixtures, no wallet, no money moved], because the one thing worse than a fake number is a fake number a second agent repeats as real.

The server also ships MCP instructions, which most hosts load into the calling model's context on connect. They tell it how to judge this agent rather than how to relay it: start at skin_record, treat a Brier score above 0.25 as worse than a coin-flipper who always says 50%, and never tell a user a bet was placed.

Ask it these three things, in this order

  1. "How good is Skin's record? Is it worth listening to?" It answers off skin_record and skin_calibration, which are pure functions of settled positions. The demo record is mediocre and it will say so.

  2. "What does it like right now, and why?" skin_scan for the book, skin_opinion for the model's working on a single market: spot, realized volatility, horizon, probability, edge.

  3. "Go ahead and place it." It cannot, and it will tell you why.

That third turn is the demo.

There is no skin_stake tool, and that is the design

An MCP server is a surface an arbitrary model can drive, usually several turns removed from the person who owns the wallet. On a venue where anyone can create a market, market titles are attacker-controlled text arriving inside the same context window as the tool descriptions. Prompt injection there is not hypothetical.

So the spend boundary sits at the process edge rather than inside a tool description. skin_propose returns a fully sized stake (Kelly fraction, binding constraint, thesis, token id) and then this:

"execution": {
  "placed": false,
  "executableHere": false,
  "command": "npm run skin -- stake --live --budget 6 --per-call 1.5 --kelly 0.25",
  "confirmation": "That command prints the receipt, then waits for the operator
                   to type the word `stake` before any order is submitted."
}

A model can reason the whole way to a position. Only a human at a terminal can open it. test/mcp.test.ts asserts that invariant twice: once against the tool names, and once against the server's source, so an edit that reaches for client.placeOrder fails the suite rather than the review.


Why it refuses

Money moves only if a call survives all seven checks. Each one has a single job, and each produces a named verdict rather than a silent skip.

#

Check

Rejects

1

Horizon

A market resolving sooner than --min-horizon (default 2 minutes), which cannot be priced, quoted and confirmed before it settles → closing-soon

2

Model

A market whose title cannot be parsed, or with too little price history to measure volatility → no-price. Barrier questions ("will BTC hit $110k") are refused outright: they resolve on touching the level, and this model prices terminal probability only.

3

Edge

Disagreement with the market smaller than 4 points → no-edge

4

Bounds

Conviction outside [0.02, 0.98], where the model's own error exceeds its claimed edge → conviction-bounds

5

Sizing

Quarter-Kelly stake below the venue minimum → below-minimum. Never rounded up.

6

Budget

Run cap or the wallet's daily limit already spent → budget-exhausted

7

Re-quote

The edge was measured against the last traded price. The quote returns averagePrice, the price this order actually fills at, and the two are not close: a market listed at 0.81 quoted at 0.90 seconds later. Every call is re-tested against the price we are really paying, and abandoned if the edge that justified it has gone.

Six of the seven run before a quote is requested. The seventh runs with the quote in hand, immediately before the order goes out, because it is the only one that can see the real price.

A scan where nothing clears is a successful scan. The demo fixtures are tuned so all the verdicts are reachable, because a gate you cannot see fire is a gate you cannot trust.


How the model works

No language model touches this path. Everything is computed, so it can be rechecked.

Pricing. Threshold markets are barrier questions with a standard answer. Assume zero-drift geometric Brownian motion:

P(S_T > K) = Φ(d₂)        d₂ = [ln(S/K) − σ²T/2] / (σ√T)

Zero drift is the honest assumption. Over a 5-minute horizon any drift you could estimate is swamped by noise, and assuming one is exactly how forecasters talk themselves into positions. μ = 0 makes the price a martingale and leaves volatility as the only real input.

Volatility. Realized, from log returns of 1-minute klines, annualised. Under 10 returns it returns null and the market is declined, because a model that always produces an answer is a model that sometimes lies.

Sizing. Kelly, for a binary token bought at cost c:

b  = (1 − c) / c
f* = p − (1 − p)·c/(1 − c)

f* is zero exactly at p = c, so agreeing with the market produces a zero bet with no special-case rule. Quarter-Kelly by default: full Kelly is only optimal if p is exactly right, and p is a model output. Kelly is brutal under overestimated edge, and the edge estimate is the least reliable term in the equation.

Scoring. Brier score, (1/N)·Σ(conviction − actual)². It is strictly proper: minimised only by reporting true beliefs. The agent cannot improve it by sounding confident and cannot improve it by hedging to 50%. It returns null, never 0, when there is nothing to score, since zero would read as perfect.

Full derivations in skill/references/model.md.


Architecture

                      ┌──────────────────────────────┐
   Binance MCP  ─────▶│  marketdata.ts               │
   (klines)           │  realized volatility         │
                      └──────────────┬───────────────┘
                                     ▼
                      ┌──────────────────────────────┐
                      │  analyst.ts    P(S>K) = Φ(d₂)│──▶ conviction
                      └──────────────┬───────────────┘
                                     ▼
   baw prediction ───▶┌──────────────────────────────┐
   market list        │  sizing.ts   Kelly + 4 caps  │──▶ stake or decline
                      └──────────────┬───────────────┘
                                     ▼
                      ┌──────────────────────────────┐
                      │  journal.jsonl (append-only) │  ◀── written BEFORE
                      └──────────────┬───────────────┘       the outcome
                                     ▼
   baw trade      ◀───┌──────────────────────────────┐
   place-order        │  cli/index.ts                │
   redeem             └──────────────┬───────────────┘
                                     ▼
   baw position   ───▶┌──────────────────────────────┐
   settled-history    │  record.ts  Brier + calib.   │──▶ record.json ──▶ web
                      └──────────────────────────────┘

Three layers, and the boundary is enforced rather than aspirational: adapters/ is everything that touches the outside world, and all failure lives there; engine/ is pure and total, with no I/O, clock, randomness or network, which is why its unit tests are worth something; cli/ formats and confirms, and never computes.

PredictionClient is the seam. LiveClient shells out to baw, DemoClient returns fixtures, and nothing above the seam knows which it holds. That is what makes the demo rail a real exercise of the code rather than a mock of it.


Repo layout

src/
  domain/types.ts        the vocabulary: Call → StakeVerdict → Position → Record
  adapters/
    exec.ts              execFile process runner (never a shell string)
    baw.ts               Agentic Wallet client + defensive JSON pickers
    marketdata.ts        REST / MCP / static kline sources, realized vol
    demo.ts              12 settled fixtures + 5 markets, every verdict reachable
  engine/
    analyst.ts           erf, Φ, P(S>K), market-title parsing
    sizing.ts            Kelly, edge, the four ceilings
    record.ts            Brier, calibration bins, equity curve
    journal.ts           append-only JSONL, first-write-wins
    scan.ts              one scan as data, shared by CLI, dashboard and MCP
  cli/                   commands, receipts, doctor, ANSI-aware rendering
  mcp/server.ts          eight read-only tools · no tool can spend

skill/                   the Agent OS skill: the policy the agent follows
web/                     Vite + React dashboard ("Ledger Noir")
fixtures/                a klines payload for exercising the --mcp-data path
test/                    68 tests: engine, adapter, MCP surface, live gates

~3,300 lines of TypeScript, strict with noUncheckedIndexedAccess.


Commands

Command

Money

What it does

skin record

—

Realized PnL, Brier, hit rate, ledger, equity, calibration

skin scan

—

Form opinions on live markets, commit nothing

skin stake

spends

Scan, print receipts, place orders after typed confirmation

skin claim

moves

Redeem settled winnings after typed confirmation

skin positions

—

What is still open, with time to resolution

skin export

—

Dump the record as JSON for the dashboard

skin doctor

—

Probe every live wallet call before the first real stake

skin mcp

—

Serve the record to other agents over MCP, read-only

Flags: --demo --live --yes --json --budget --per-call --kelly --min-order --min-horizon --limit --mcp-data --deep --out. Full reference: skill/references/commands.md.


Safety

This spends real money from a real wallet, so:

  • Typed confirmation. stake and claim require the whole word typed. y is rejected. On a non-TTY they refuse outright unless --yes was passed, and --yes announces itself in the output.

  • No shell strings. exec.ts uses execFile with an argv array. A market title cannot become a command.

  • Market titles are never sent to a model. They are attacker-controlled text from a public venue. They are parsed by regex and rendered as data. This removes the prompt-injection surface rather than trying to filter it.

  • The wallet's daily limit is a sizing input, not an error to handle. The agent cannot raise it; that lives in the Binance app.

  • Errors are reported verbatim. BawError carries the CLI's own words. A wallet error is never paraphrased into a guess.

  • Nulls are never defaulted. A missing price is null and produces a decline. It never silently becomes 0, which would read as a 100% edge.

  • The journal is append-only, first-write-wins. Editing it is the one way to make the Brier score meaningless, which is why re-recording a token is refused rather than overwritten.


Testing

npm test          # 68 tests
npm run typecheck # tsc --noEmit, strict

math.test.ts covers the mathematics rather than the plumbing. mcp.test.ts pins the spend boundary twice: no tool name may contain a state-changing verb, and the server's own source may not reference client.placeOrder or client.redeem. adapter.test.ts runs the client against a fake baw binary, and every case in it is a shape this adapter got wrong against the live wallet: a CONNECTED wallet read as signed out, the general daily limit read in place of the prediction quota, USDT on the wrong chain counted as bankroll. All three failed silently. scan.test.ts covers the two gates only a live venue could have taught us, including the market that nearly cost a real dollar: a BNB coin flip listed at 3%, quoting at 53%.

CI runs all of it on every push, plus a smoke test that executes every command in demo mode on a clean machine with no wallet, no network and no journal in $HOME. It re-exports web/public/record.json and fails if it differs from the committed copy, so the dashboard payload cannot quietly drift from the fixtures that produce it.

One test caught a real bug: parseClaim('BTC Up or Down · 5 min') returned 'below', because the below pattern was tested first and matched the word "Down". That silently inverted every two-sided short-duration market, which is most of the tradable universe.


For judges and reviewers

Five minutes, in order:

  1. npm install && npm run skin -- record --demo: the whole thesis in one screen. The demo agent is barely profitable and its Brier score is mediocre. That is on purpose.

  2. npm run skin -- scan --demo: markets refused, each with a named reason.

  3. npm run skin -- stake --demo: the receipt and the typed confirmation. Try typing y; it will not accept it.

  4. claude mcp add skin -- npx tsx src/cli/index.ts mcp --demo: ask your own agent "how good is skin's record?" and watch it answer from settled positions. Then ask it to place a bet, and watch it discover it cannot. Setup for any other host is in Skin as an MCP server.

  5. src/engine/sizing.ts: ~120 lines, the heart of it. Four ceilings, the tightest binds, and the binding one is named in the output.

If you would rather watch than run: 2-minute demo.

What is different here, stated plainly: the agent's track record is adversarially verifiable. You do not have to trust the dashboard, the README or the agent. The positions are on-chain and the scoring rule is strictly proper. If the agent inflated its confidence, the Brier score would get worse, and the Brier score is the number it is judged on.


Honest limitations

Written before anyone asks.

  • The model is simple by choice. Zero-drift lognormal on realized vol. It has no view on order flow, funding, news, or microstructure. It will be systematically wrong around scheduled events. A better model would improve the PnL; it would not change what the project demonstrates.

  • Track record length. Brier over a dozen calls is indicative, not conclusive. The CLI says so itself: under five scored calls it prints "too few to judge" instead of a verdict.

  • Fees and slippage are inside the realized PnL because it is computed from settled positions, but they are not modelled ahead of a trade. On $1 stakes the effect is small; at size it would need to enter the edge threshold.

  • Orders are submitted, not filled. place-order returning success means accepted. The CLI says this and points at baw prediction order history --status FILLED.

  • Journal integrity is local. Append-only and first-write-wins, but it is a file on the operator's machine. Anchoring a hash on-chain would make it externally verifiable; the settled positions already are.


License

MIT. See LICENSE.

Built for the Binance Agent OS Mini Hackathon, Track A.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    An MCP server that streams the latest Binance announcements to your AI agents in real time, enabling instant analysis and action on market-moving updates.
    1
    1
    MIT
  • A
    license
    C
    quality
    C
    maintenance
    MCP server that provides a unified prediction market API for multiple venues like Polymarket and Kalshi, allowing AI agents to discover markets, fetch order books, and execute trades through a single interface.
    32
    262 npm
    8
    MIT