skin-in-the-game
by martinvibes
README.md
<div align="center">
<img src="docs/logo.svg" width="72" height="72" alt="">
# Skin
*skin in the game*
**An AI analyst that has to bet its own money on every call it makes.**
Its track record is not a claim it writes about itself.
It is settled positions on Binance prediction markets, scored with a rule it cannot game.
**[Live dashboard →](https://tryskin.vercel.app/)** · **[2-minute demo →](https://youtu.be/kcDXys0YT3k)**
[](https://agent.binance.com)
[](https://github.com/binance/binance-skills-hub)
[](https://agent.binance.com/mcp/agentic)
[](https://github.com/martinvibes/skin-in-the-game/actions)
[](test/)
[](tsconfig.json)
[](LICENSE)
</div>
---
## The problem
Every AI agent will tell you what it thinks the market will do. None of them
pay for being wrong.
An agent can say *"BTC looks bullish, roughly 80% confidence"* a hundred times.
It is never scored, so the confidence number means nothing, and it costs
nothing to inflate. This is the single largest unsolved problem with AI
financial agents: **there is no price on being wrong**, so there is no
information in anything they say.
## The fix
Binance Agent OS shipped prediction markets to the Agentic Wallet. A prediction
market quotes belief directly: an outcome token at $0.62 *is* a 62% probability.
That makes it the only venue where an agent's forecast and its money are the
same object.
So this agent is not allowed to have an opinion for free. It prices a market
from realized volatility (a closed-form model, no LLM guess), and if it
disagrees by more than 4 points it sizes a stake by the **Kelly criterion** and
buys with real USDT. The conviction goes into an append-only journal *before*
the outcome is known. When the market resolves, the call is scored by **Brier
score**, a strictly proper rule, so the agent's best possible strategy is to
state what it actually believes. It sweeps its own winnings too, because
prediction markets do not pay out automatically and an agent that never claims
goes broke while being right.
Every number on the dashboard traces back to a settled on-chain position.
**The agent does not get a vote in its own performance review.**
---
## What it looks like
[](https://tryskin.vercel.app)
The dashboard is not a report you scroll. It is the agent's loop with a button
on it. **Run a scan** prices five markets one at a time and says out loud why it
refuses most of them; **Stake** makes you type the word `stake` before any money
moves, exactly as the CLI does; **Claim** sweeps the winnings that prediction
markets do not pay out on their own. Every number in it comes from the `scan`
block of `record.json`, which the CLI produced by calling the same `formOpinion`
and `sizeStake` a live run calls, so the demo cannot quietly disagree with the
tool it is demonstrating.
### The record
```
SKIN skin in the game ● DEMO synthetic fixtures · no wallet, no money moved
THE RECORD
computed from settled positions · the agent does not get a vote
──────────────────────────────────────────────────────────────────────────
realized PnL +$0.07 on $13.20 staked
return on stake 0.5%
Brier score 0.261 roughly as useful as always saying 50%
0 perfect · 0.25 = always saying "50%" · lower is better
settled calls 12
won / lost 7 / 5
hit rate 58.3%
──────────────────────────────────────────────────────────────────────────
```
That is the bundled demo record, and it is deliberately **unflattering**: up
seven cents on $13.20, a Brier score barely better than a coin flip, and the
biggest single loss is the call it was 88% sure about. A demo that showed a
winning agent would be showing you the marketing, not the mechanism.
The point of the project is that this number is not editable. So it is not
edited.
### A betting slip
Every stake prints a receipt before any money moves: the full derivation, and
which of the four ceilings actually bound the size.
```
┌┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┐
┆ Bitcoin Up or Down on September 8? PENDING ┆
┆ side NO ┆
┆┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┆
┆ agent says 83.4% ┆
┆ market says 74.0% ┆
┆ edge 9.4% ┆
┆ payout odds 0.35:1 ┆
┆ full Kelly 36.2% ┆
┆ applied (¼ Kelly) 9.0% ┆
┆ bound by per-call-cap ┆
┆┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┆
┆ STAKE $1.00 ┆
┆┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┆
┆ BTC spot $78,426, realized vol 27% (200 × 1m candles). Zero-drift ┆
┆ lognormal over 4.0 h puts P(BTC above $78,861) at 16.6%; the book ┆
┆ quotes the NO token at 75.0% for this size. ┆
└┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┄┘
```
That is a real receipt. It filled as order `26090800001869670157` for 1.32
shares at $0.7545.
### A scan
```
SCAN
12 market(s) · bankroll $14.99 · run cap $1.00 · ¼ Kelly
──────────────────────────────────────────────────────────────────────────
side it would buy @ price it would pay · fair = what the model thinks it is worth
pass BNB Up or Down · 7:45AM-8AM ET NO @ 22% · fair 37% below-minimum
pass Ethereum Up or Down · 7AM ET NO @ 99% · fair 100% conviction-bounds
pass Bitcoin Up or Down · 7:45AM ET NO @ 51% · fair 52% no-edge
pass Bitcoin Up or Down · 7AM ET NO @ 98% · fair 100% conviction-bounds
BET Bitcoin Up or Down on September 8? NO @ 74% · fair 83% $1.00
pass Ethereum Up or Down on September 8? NO @ 47% · fair 48% no-edge
pass BNB Up or Down 1d YES @ 85% · fair 99% conviction-bounds
──────────────────────────────────────────────────────────────────────────
```
Eleven of twelve live markets refused, each with a named reason. That is the
system working, not the system failing to find anything.
### The dashboard
**[tryskin.vercel.app](https://tryskin.vercel.app)**
A React page rendering the same `record.json` the CLI exports. No backend and
no key: it reads one file, so what it shows is exactly what the CLI produced,
and it could not place an order if it wanted to. The mode badge is permanent,
so a demo number can never be mistaken for a live one. Both charts are
hand-drawn SVG; the whole page gzips to 70 kB.


```bash
npm --prefix web install && npm run web:dev
```
<details>
<summary>Deploying it yourself</summary>
The dashboard lives in `web/`, and this repo has **two** `package.json` files:
the CLI's at the root and the web app's in `web/`. Set the Vercel project's
**Root Directory to `web`**, or the build runs the root `package.json`'s
`build` script (which compiles the CLI) and deploys nothing servable. The
symptom is a successful build followed by a 404 on every route.
</details>
---
## Quick start
**No wallet, no money, no signup.** The demo rail is fully offline and is what
a reviewer should run first:
```bash
git clone https://github.com/martinvibes/skin-in-the-game
cd skin-in-the-game
npm install
npm run skin -- record --demo # the track record
npm run skin -- scan --demo # opinions and refusals, no money
npm run skin -- stake --demo # receipts + typed confirmation
npm run skin -- claim --demo # sweep unclaimed winnings
npm test # 68 tests
npm --prefix web install && npm run web:dev # the dashboard
```
### Going live
Live needs a Binance **Agentic Wallet** (created in the Binance mobile app; an
agent cannot create one for you) and a little USDT on BSC. Stakes are about $1,
so $10 to $15 runs this for real.
```bash
npx skills add binance/binance-skills-hub/skills/binance-web3/binance-agentic-wallet
npm install -g @binance/agentic-wallet
baw auth signin # approve on your phone
```
Once `baw` is on your PATH, **live is the default** and `--demo` is the opt-out.
`--live` forces it and fails loudly if the CLI is missing, so live is never
entered by accident either.
```bash
npm run skin -- doctor # wallet, session, quota, balance
npm run skin -- scan # the real book, re-priced against real quotes
npm run skin -- stake --budget 1 --per-call 1 # the only command that spends
npm run skin -- record # settled record, calibration, open positions
npm run skin -- claim # redeem settled wins to the wallet
```
Start with `doctor`. It probes every wallet call the staking path depends on
and names the one that would have failed, which is a cheaper way to find a
signed-out session than discovering it halfway through an order.
Prediction trading has its own daily quota in the Binance app, separate from
the wallet's general limit. `doctor` reads it, and sizing treats it as a hard
ceiling.
---
## Binance Agent OS surfaces used
| Surface | How this project uses it |
|---|---|
| **Agentic Wallet (`baw`)** | The entire trading path. `prediction market list / detail / last-trade-price` to read markets, `prediction trade place-order` to stake, `prediction trade redeem` to claim, `prediction position list --tab ONGOING\|PENDING_CLAIM` and `settled-history` to build the record, `wallet balance` for the bankroll. See [`src/adapters/baw.ts`](src/adapters/baw.ts). |
| **MCP, as a server** | Skin is itself an MCP server. `skin mcp` exposes eight read-only tools so another agent can query this one's record, calibration and current book, and cannot spend its money. See [`src/mcp/server.ts`](src/mcp/server.ts) and [the section below](#skin-as-an-mcp-server). |
| **Binance MCP Server** | Market data for the pricing model. An MCP-authenticated agent fetches klines and hands them over with `--mcp-data`, so volatility is computed under the user's own Agent OS session. See [`src/adapters/marketdata.ts`](src/adapters/marketdata.ts) and [the walk-through below](#using-the-mcp-server). |
| **Skills Hub format** | [`skill/SKILL.md`](skill/SKILL.md) is a publishable skill in the Hub's frontmatter format, with a `references/` split. It is what turns the CLI into an agent workflow: a policy the agent has to follow, not just a binary it can call. |
| **MPC Wallet limits** | The wallet's own daily spend limit is read and treated as a hard ceiling on sizing, so the agent is *structurally* unable to exceed a boundary the human set in the Binance app. |
The wallet's daily limit as a sizing input is the part worth pausing on. Most
agent projects treat wallet limits as an error to handle. Here it is one of the
four terms in the position-sizing minimum, on equal footing with Kelly.
### Using the MCP Server
```bash
claude mcp add binance-mcp-server --transport http https://agent.binance.com/mcp/agentic
```
The agent fetches candles under its own authenticated session, writes them to a
file as `{ "BTCUSDT": { "interval": "1m", "closes": [...] } }`, and hands the
path over with `skin scan --mcp-data ./klines.json`. A committed fixture
exercises the path with no session and no network:
```bash
npm run skin -- scan --demo --mcp-data fixtures/klines.example.json
```
Two of those fixture markets have a strike sitting exactly on spot, and the
model prices both at exactly 50%: the cheapest available check that the
lognormal is wired up correctly. At the money, over any horizon, a zero-drift
walk is a coin flip.
Without `--mcp-data` the CLI falls back to Binance's public market-data REST
endpoint, which needs no credentials. Same arithmetic either way. It tries four
hosts in order, including `data-api.binance.vision`, because `api.binance.com`
does not resolve from every country this was developed in.
---
## Skin as an MCP server
The agent keeping this record should not be the only one able to read it. `skin
mcp` speaks MCP over stdio, so any host (Claude Code, Claude Desktop, Cursor,
another desk's agent) can interrogate this one directly instead of trusting a
screenshot of it.
Eight tools. **None of them spends.**
### Running it
**Zero setup, in this repo.** A [`.mcp.json`](.mcp.json) is committed, so a
clone opened in Claude Code offers the server on its own. Approve it at the
prompt and the tools are there.
**One line, from the repo root:**
```bash
npm install
claude mcp add skin -- npx tsx src/cli/index.ts mcp --demo
```
**Any other MCP host** (Claude Desktop, Cursor, an agent you wrote yourself)
launches the same process. Give it an absolute path, because the host's working
directory is not yours:
```json
{
"mcpServers": {
"skin": {
"command": "npx",
"args": [
"tsx",
"/absolute/path/to/skin-in-the-game/src/cli/index.ts",
"mcp",
"--demo"
]
}
}
}
```
Run `npm install` in the clone once and that is the whole install. No API key,
no wallet, no build step: `--demo` serves fixtures that ship in the source, so
the tools answer with the network off.
One wrinkle worth knowing before you blame the config: launched from outside
the repo, `npx` cannot see the local `tsx` and fetches its own copy, which takes
a few seconds and can trip a host's first connect timeout. It is cached after
that, so reconnecting works. Hosts that support a `cwd` field avoid it entirely
by pointing at the clone.
**Pointing it at the real wallet** is one flag less. Drop `--demo` and the
tools read live markets, the live journal and settled positions through `baw`,
which requires `baw auth signin` to have happened. It stays read-only either
way:
```json
"args": ["tsx", "/absolute/path/to/skin-in-the-game/src/cli/index.ts", "mcp"]
```
Sizing defaults for `skin_propose` are passed exactly as the CLI takes them,
after the `mcp` argument:
```bash
npx tsx src/cli/index.ts mcp --limit 12 --budget 6 --per-call 1.5 --kelly 0.25
```
**Knowing it came up.** Every byte of stdout belongs to the JSON-RPC stream, so
the server greets you on stderr, where the host will show it and the protocol
will never read it:
```
SKIN MCP demo mode · listening on stdio
skin_record settled calls, PnL, Brier score
skin_calibration where the numbers break down
skin_slips convictions written before the outcome
skin_positions money committed, not yet resolved
skin_unclaimed settled wins still on-chain
skin_scan every open market, with the working
skin_opinion one market: vol, horizon, probability
skin_propose a sized stake, and the command a human must run
Eight tools, none of which can spend. There is deliberately no
skin_stake: an agent can reason all the way to a position and still
cannot open one.
stdout is the protocol, so nothing more will be printed here.
Waiting for a client. Ctrl-C to stop.
```
Silence after that line is what success looks like.
### What the calling agent gets
| Tool | Answers |
|---|---|
| `skin_record` | Settled calls, hit rate, realized PnL, ROI, Brier score |
| `skin_calibration` | Reliability curve: where the agent's numbers break down |
| `skin_slips` | Every conviction written down before the outcome was known |
| `skin_positions` | Money committed, not yet resolved, with the conviction behind it |
| `skin_unclaimed` | Settled wins still sitting on-chain |
| `skin_scan` | A view on every open market, with the full working |
| `skin_opinion` | One market: spot, volatility, horizon, probability, edge |
| `skin_propose` | A sized stake, and the command a human must run to place it |
Each tool returns a one-line headline a model can act on, then the JSON behind
it, so a host that only shows the first line still shows something true. In
demo mode every headline is tagged `[demo mode: synthetic fixtures, no wallet,
no money moved]`, because the one thing worse than a fake number is a fake
number a second agent repeats as real.
The server also ships MCP `instructions`, which most hosts load into the calling
model's context on connect. They tell it how to *judge* this agent rather than
how to relay it: start at `skin_record`, treat a Brier score above 0.25 as
worse than a coin-flipper who always says 50%, and never tell a user a bet was
placed.
### Ask it these three things, in this order
1. *"How good is Skin's record? Is it worth listening to?"* It answers off
`skin_record` and `skin_calibration`, which are pure functions of settled
positions. The demo record is mediocre and it will say so.
2. *"What does it like right now, and why?"* `skin_scan` for the book,
`skin_opinion` for the model's working on a single market: spot, realized
volatility, horizon, probability, edge.
3. *"Go ahead and place it."* It cannot, and it will tell you why.
That third turn is the demo.
### There is no `skin_stake` tool, and that is the design
An MCP server is a surface an arbitrary model can drive, usually several turns
removed from the person who owns the wallet. On a venue where anyone can create
a market, market titles are attacker-controlled text arriving inside the same
context window as the tool descriptions. Prompt injection there is not
hypothetical.
So the spend boundary sits at the process edge rather than inside a tool
description. `skin_propose` returns a fully sized stake (Kelly fraction,
binding constraint, thesis, token id) and then this:
```json
"execution": {
"placed": false,
"executableHere": false,
"command": "npm run skin -- stake --live --budget 6 --per-call 1.5 --kelly 0.25",
"confirmation": "That command prints the receipt, then waits for the operator
to type the word `stake` before any order is submitted."
}
```
A model can reason the whole way to a position. Only a human at a terminal can
open it. [`test/mcp.test.ts`](test/mcp.test.ts) asserts that invariant twice:
once against the tool names, and once against the server's source, so an edit
that reaches for `client.placeOrder` fails the suite rather than the review.
---
## Why it refuses
Money moves only if a call survives all seven checks. Each one has a single job,
and each produces a named verdict rather than a silent skip.
| # | Check | Rejects |
|---|---|---|
| 1 | **Horizon** | A market resolving sooner than `--min-horizon` (default 2 minutes), which cannot be priced, quoted and confirmed before it settles → `closing-soon` |
| 2 | **Model** | A market whose title cannot be parsed, or with too little price history to measure volatility → `no-price`. Barrier questions ("will BTC *hit* $110k") are refused outright: they resolve on touching the level, and this model prices terminal probability only. |
| 3 | **Edge** | Disagreement with the market smaller than 4 points → `no-edge` |
| 4 | **Bounds** | Conviction outside `[0.02, 0.98]`, where the model's own error exceeds its claimed edge → `conviction-bounds` |
| 5 | **Sizing** | Quarter-Kelly stake below the venue minimum → `below-minimum`. Never rounded up. |
| 6 | **Budget** | Run cap or the wallet's daily limit already spent → `budget-exhausted` |
| 7 | **Re-quote** | The edge was measured against the last *traded* price. The quote returns `averagePrice`, the price this order actually fills at, and the two are not close: a market listed at `0.81` quoted at `0.90` seconds later. Every call is re-tested against the price we are really paying, and abandoned if the edge that justified it has gone. |
Six of the seven run before a quote is requested. The seventh runs with the
quote in hand, immediately before the order goes out, because it is the only one
that can see the real price.
**A scan where nothing clears is a successful scan.** The demo fixtures are
tuned so all the verdicts are reachable, because a gate you cannot see fire is
a gate you cannot trust.
---
## How the model works
No language model touches this path. Everything is computed, so it can be
rechecked.
**Pricing.** Threshold markets are barrier questions with a standard answer.
Assume zero-drift geometric Brownian motion:
```
P(S_T > K) = Φ(d₂) d₂ = [ln(S/K) − σ²T/2] / (σ√T)
```
Zero drift is the honest assumption. Over a 5-minute horizon any drift you
could estimate is swamped by noise, and assuming one is exactly how forecasters
talk themselves into positions. μ = 0 makes the price a martingale and leaves
volatility as the only real input.
**Volatility.** Realized, from log returns of 1-minute klines, annualised.
Under 10 returns it returns `null` and the market is declined, because a model that
always produces an answer is a model that sometimes lies.
**Sizing.** Kelly, for a binary token bought at cost `c`:
```
b = (1 − c) / c
f* = p − (1 − p)·c/(1 − c)
```
`f*` is zero exactly at `p = c`, so agreeing with the market produces a zero
bet with no special-case rule. **Quarter-Kelly** by default: full Kelly is only
optimal if `p` is exactly right, and `p` is a model output. Kelly is brutal
under overestimated edge, and the edge estimate is the least reliable term in
the equation.
**Scoring.** Brier score, `(1/N)·Σ(conviction − actual)²`. It is *strictly
proper*: minimised only by reporting true beliefs. The agent cannot improve it
by sounding confident and cannot improve it by hedging to 50%. It returns
`null`, never `0`, when there is nothing to score, since zero would read as perfect.
Full derivations in [`skill/references/model.md`](skill/references/model.md).
---
## Architecture
```
┌──────────────────────────────┐
Binance MCP ─────▶│ marketdata.ts │
(klines) │ realized volatility │
└──────────────┬───────────────┘
▼
┌──────────────────────────────┐
│ analyst.ts P(S>K) = Φ(d₂)│──▶ conviction
└──────────────┬───────────────┘
▼
baw prediction ───▶┌──────────────────────────────┐
market list │ sizing.ts Kelly + 4 caps │──▶ stake or decline
└──────────────┬───────────────┘
▼
┌──────────────────────────────┐
│ journal.jsonl (append-only) │ ◀── written BEFORE
└──────────────┬───────────────┘ the outcome
▼
baw trade ◀───┌──────────────────────────────┐
place-order │ cli/index.ts │
redeem └──────────────┬───────────────┘
▼
baw position ───▶┌──────────────────────────────┐
settled-history │ record.ts Brier + calib. │──▶ record.json ──▶ web
└──────────────────────────────┘
```
Three layers, and the boundary is enforced rather than aspirational:
**`adapters/`** is everything that touches the outside world, and all failure
lives there; **`engine/`** is pure and total, with no I/O, clock, randomness or
network, which is why its unit tests are worth something; **`cli/`** formats and
confirms, and never computes.
`PredictionClient` is the seam. `LiveClient` shells out to `baw`, `DemoClient`
returns fixtures, and nothing above the seam knows which it holds. That is what
makes the demo rail a real exercise of the code rather than a mock of it.
---
## Repo layout
```
src/
domain/types.ts the vocabulary: Call → StakeVerdict → Position → Record
adapters/
exec.ts execFile process runner (never a shell string)
baw.ts Agentic Wallet client + defensive JSON pickers
marketdata.ts REST / MCP / static kline sources, realized vol
demo.ts 12 settled fixtures + 5 markets, every verdict reachable
engine/
analyst.ts erf, Φ, P(S>K), market-title parsing
sizing.ts Kelly, edge, the four ceilings
record.ts Brier, calibration bins, equity curve
journal.ts append-only JSONL, first-write-wins
scan.ts one scan as data, shared by CLI, dashboard and MCP
cli/ commands, receipts, doctor, ANSI-aware rendering
mcp/server.ts eight read-only tools · no tool can spend
skill/ the Agent OS skill: the policy the agent follows
web/ Vite + React dashboard ("Ledger Noir")
fixtures/ a klines payload for exercising the --mcp-data path
test/ 68 tests: engine, adapter, MCP surface, live gates
```
~3,300 lines of TypeScript, `strict` with `noUncheckedIndexedAccess`.
---
## Commands
| Command | Money | What it does |
|---|---|---|
| `skin record` | — | Realized PnL, Brier, hit rate, ledger, equity, calibration |
| `skin scan` | — | Form opinions on live markets, commit nothing |
| `skin stake` | **spends** | Scan, print receipts, place orders after typed confirmation |
| `skin claim` | **moves** | Redeem settled winnings after typed confirmation |
| `skin positions` | — | What is still open, with time to resolution |
| `skin export` | — | Dump the record as JSON for the dashboard |
| `skin doctor` | — | Probe every live wallet call before the first real stake |
| `skin mcp` | — | Serve the record to other agents over MCP, read-only |
Flags: `--demo --live --yes --json --budget --per-call --kelly --min-order --min-horizon --limit --mcp-data --deep --out`.
Full reference: [`skill/references/commands.md`](skill/references/commands.md).
---
## Safety
This spends real money from a real wallet, so:
- **Typed confirmation.** `stake` and `claim` require the whole word typed.
`y` is rejected. On a non-TTY they refuse outright unless `--yes` was passed,
and `--yes` announces itself in the output.
- **No shell strings.** [`exec.ts`](src/adapters/exec.ts) uses `execFile` with
an argv array. A market title cannot become a command.
- **Market titles are never sent to a model.** They are attacker-controlled
text from a public venue. They are parsed by regex and rendered as data. This
removes the prompt-injection surface rather than trying to filter it.
- **The wallet's daily limit is a sizing input**, not an error to handle. The
agent cannot raise it; that lives in the Binance app.
- **Errors are reported verbatim.** [`BawError`](src/adapters/exec.ts) carries
the CLI's own words. A wallet error is never paraphrased into a guess.
- **Nulls are never defaulted.** A missing price is `null` and produces a
decline. It never silently becomes `0`, which would read as a 100% edge.
- **The journal is append-only, first-write-wins.** Editing it is the one way
to make the Brier score meaningless, which is why re-recording a token is
refused rather than overwritten.
---
## Testing
```bash
npm test # 68 tests
npm run typecheck # tsc --noEmit, strict
```
[`math.test.ts`](test/math.test.ts) covers the mathematics rather than the
plumbing. [`mcp.test.ts`](test/mcp.test.ts) pins the spend boundary twice: no
tool name may contain a state-changing verb, and the server's own source may not
reference `client.placeOrder` or `client.redeem`.
[`adapter.test.ts`](test/adapter.test.ts) runs the client against a fake `baw`
binary, and every case in it is a shape this adapter got wrong against the live
wallet: a `CONNECTED` wallet read as signed out, the general daily limit read in
place of the prediction quota, USDT on the wrong chain counted as bankroll. All
three failed silently. [`scan.test.ts`](test/scan.test.ts) covers the two gates
only a live venue could have taught us, including the market that nearly cost a
real dollar: a BNB coin flip listed at 3%, quoting at 53%.
[CI](.github/workflows/ci.yml) runs all of it on every push, plus a smoke test
that executes every command in demo mode on a clean machine with no wallet, no
network and no journal in `$HOME`. It re-exports `web/public/record.json` and
fails if it differs from the committed copy, so the dashboard payload cannot
quietly drift from the fixtures that produce it.
One test caught a real bug: `parseClaim('BTC Up or Down · 5 min')` returned
`'below'`, because the `below` pattern was tested first and matched the word
*"Down"*. That silently inverted every two-sided short-duration market, which is
most of the tradable universe.
---
## For judges and reviewers
Five minutes, in order:
1. **`npm install && npm run skin -- record --demo`**: the whole thesis in one
screen. The demo agent is barely profitable and its Brier score is mediocre.
That is on purpose.
2. **`npm run skin -- scan --demo`**: markets refused, each with a named
reason.
3. **`npm run skin -- stake --demo`**: the receipt and the typed confirmation.
Try typing `y`; it will not accept it.
4. **`claude mcp add skin -- npx tsx src/cli/index.ts mcp --demo`**: ask your
own agent *"how good is skin's record?"* and watch it answer from settled
positions. Then ask it to place a bet, and watch it discover it cannot.
Setup for any other host is in [Skin as an MCP server](#skin-as-an-mcp-server).
5. **[`src/engine/sizing.ts`](src/engine/sizing.ts)**: ~120 lines, the heart of
it. Four ceilings, the tightest binds, and the binding one is named in the
output.
If you would rather watch than run: **[2-minute demo](https://youtu.be/kcDXys0YT3k)**.
What is different here, stated plainly: the agent's **track record is
adversarially verifiable**. You do not have to trust the dashboard, the README
or the agent. The positions are on-chain and the scoring rule is strictly
proper. If the agent inflated its confidence, the Brier score would get worse,
and the Brier score is the number it is judged on.
---
## Honest limitations
Written before anyone asks.
- **The model is simple by choice.** Zero-drift lognormal on realized vol. It
has no view on order flow, funding, news, or microstructure. It will be
systematically wrong around scheduled events. A better model would improve
the PnL; it would not change what the project demonstrates.
- **Track record length.** Brier over a dozen calls is indicative, not
conclusive. The CLI says so itself: under five scored calls it prints *"too
few to judge"* instead of a verdict.
- **Fees and slippage** are inside the realized PnL because it is computed from
settled positions, but they are not modelled *ahead* of a trade. On $1 stakes
the effect is small; at size it would need to enter the edge threshold.
- **Orders are submitted, not filled.** `place-order` returning success means
accepted. The CLI says this and points at `baw prediction order history
--status FILLED`.
- **Journal integrity is local.** Append-only and first-write-wins, but it is a
file on the operator's machine. Anchoring a hash on-chain would make it
externally verifiable; the settled positions already are.
---
## License
MIT. See [LICENSE](LICENSE).
Built for the **Binance Agent OS Mini Hackathon**, Track A.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues