Reins
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Reinspaper trade a 0.05 BTC long with 5x leverage and a 1% stop-loss"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Reins
Let an AI agent trade without giving it the power to lose everything.
npx @r2rlabs/reins init # adds a paper-trading Reins server to .mcp.jsonSee it at work: an AI agent trading Hyperliquid on paper, run 2.
The source is on GitHub: R2Rlabs/reins. Every limit in this README is a function you can read.
What Reins is
Limits the agent can't argue with. Every order passes checks that live outside the model: position size, leverage, daily loss, which markets it may trade, and a stop-loss on every position. No prompt, clever reasoning or mistake gets around them.
A record you can check. Every decision is logged with the reason the agent gave — including the orders it was refused, the times it chose to wait, and what the exchange did without it. You can compare what the agent said with what actually happened.
Somewhere to rehearse. Paper trading on live Hyperliquid prices with realistic fills and fees, using the same tools and limits as live trading.
Non-custodial. Your money stays in your own Hyperliquid account. Reins trades through an API wallet, which can place orders but cannot withdraw.
Paid for openly. A small builder fee (2 bp) on live orders, shown up front.
What Reins is not
Not a trading bot or a strategy. Reins doesn't decide what to trade and won't make an agent profitable. The demo agent is a test drive, not the product.
Not an exchange, broker or custodian. It holds no funds and matches no orders; Hyperliquid does that.
Not a promise against losses. Limits cap how much can go wrong. They don't make bad trades good, and stops can slip in fast markets.
Not financial advice.
Not something you have to take on trust. Not the agent's word, and not ours: every action is on the record.
Principles
Limits live outside the model.
The record outranks the reasoning.
Paper before real money.
Your keys, your funds.
Why this exists
Several open-source Hyperliquid MCP servers already exist. They are thin API wrappers: you hand them a private key and hope the model behaves. The gap is everything around that — the limits, the sandbox, the audit trail. That is what this repo is.
Related MCP server: Hyperliquid MCP
Status
Early, but real. 391 tests, no network calls in any of them.
Module | What it does |
| The risk engine — position cap, risk per trade, leverage cap, daily loss halt, allowlist, rate limit |
| Hyperliquid REST client — orders, cancels, book, candles, fills, account state, builder code |
| The agent-facing tools, each risk-checked and logged before anything is sent |
| Paper trading — live prices, simulated fills. See below |
| Persists a paper run across restarts |
| The interface live and paper both satisfy |
| Append-only record of every attempt and its stated reason |
| JSON Lines log on disk |
| The |
| stdio entry point |
|
|
| The same nine tools as HTTP routes, for bots that do not speak MCP |
| The limits, client, log and banner both front doors share |
|
|
| Reins' builder fee, and the approval live trading needs before it starts |
|
|
|
|
|
|
| Reads Hyperliquid's LZ4-compressed data files, without a dependency |
| The ApproveBuilderFee action: EIP-712 typed data, signature checks |
| Price and size formatting to Hyperliquid's tick and lot rules |
| L1 action signing, verified against the Python SDK's vectors |
| A Signer backed by a private key |
| The Signer interface plus a stub. See "On signing" below |
| In-memory API stand-in, so agents can be tested without a network |
| Claude trading on paper through Reins, producing a publishable decision log |
All of it works end to end. Going live is now a matter of funding an account and
setting REINS_MODE=live with a key.
Going live
REINS_MODE=live
REINS_NETWORK=testnet # mainnet spends real money
REINS_ACCOUNT_ADDRESS=0x... # your Hyperliquid account
REINS_PRIVATE_KEY=0x... # an API wallet's key, not your account's; omit to stay read-onlyOn mainnet, approve Reins' 2 bp builder fee once, from the account's own
wallet, before the first live session: npx @r2rlabs/reins approve-builder.
Until the account has, live mode refuses to start and says so.
Use an API wallet, not your account's own key. Hyperliquid lets an account
approve API wallets that can trade for it but cannot withdraw from it — create
one on Hyperliquid's API page, signing the approval with your main wallet. Put
its key in REINS_PRIVATE_KEY and your account's address in
REINS_ACCOUNT_ADDRESS: orders are signed by the API wallet, while positions,
fills and stops are read from the account, because under the API wallet's own
address the account looks empty. The startup banner names both. Hyperliquid
prunes API wallets that expire or are replaced; make a fresh one rather than
reusing an old address.
Every account mode works. Hyperliquid starts new accounts in Unified,
where the USDC sits in the spot balance and the perps account reads $0; Reins
asks the account which mode it is in and, for Unified and Portfolio margin,
takes equity from the USDC balance plus the positions' unrealized PnL. Manual
accounts are read from perps as before. (A builder address is different: it
must be in Manual — standard — to earn fees.)
Defaults are deliberately safe: REINS_MODE is paper and REINS_NETWORK is
testnet, so trading real funds takes two explicit changes rather than one
forgotten variable. Live mode without a key stays read-only instead of failing
at the first order, and mainnet prints a loud banner on startup.
The private key is only read in live mode — paper never signs anything, so it is never handed the means to.
See it run
npm run build
npm run demo -- --scripted # free: no API key, live prices, paper fillsWith an Anthropic API key, npm run demo runs Claude as the agent under a hard
spending cap. See examples/demo-agent for cost, limits and how
to publish the result honestly.
The decision log
Every attempted action is appended to a log with the reason the agent gave for it — including the ones the risk limits refused. That is the point. "The agent tried to open six times its position cap at 3am, and here is what it said it was doing" is the most useful thing this system can tell you, and it only exists if refusals are written down.
place_order, close_position and set_stop_loss take a required reason. An optional
field gets omitted; a required one gets answered.
Records are JSON Lines — one self-contained object per line, appended, never rewritten. Appends stay cheap as the file grows, a truncated write damages one line instead of the file, and you can query it with ordinary tools:
jq -r 'select(.risk.allowed == false) | [.time, .risk.code, .reason] | @tsv' decisions.jsonlEach record holds the stated reason, the request, the account context at the
time, the risk verdict, and the outcome. Set REINS_LOG_FILE to persist it;
without it the log lives in memory and dies with the process.
The exchange also acts on its own: a stop-loss fires, a resting order fills
hours after it was placed. Before every tool call Reins compares the account's
fills with the log and appends each one it has not seen as an exchange_fill
record, whose reason starts "Not an agent decision:" and says whether it was a
stop, a resting order, or a fill Reins did not place. Without them the trade
that mattered most, the stop that closed it, would be missing from the record.
The agent reads them back through get_recent_decisions like anything else.
Two things to be clear about
reason is testimony, not ground truth. It is the model's own account of
why it acted, captured because reasoning that is not asked for cannot be
recovered later. A model can rationalise after the fact, and a confident
explanation is not evidence that the explanation is what actually drove the
decision. It is extremely useful for debugging and audit. It is not proof.
A failed log write never fails a filled order. If the log cannot be written, the tool still reports success and attaches a warning. An agent told its order failed will place it again — so a lost log line would become a doubled position, which is far worse than the missing line. The risk engine is the safety mechanism; the log is observability.
Paper trading
REINS_MODE=paper (the default) runs the same tools, the same risk engine and
the same agent against real prices with simulated fills. Nothing is signed
and nothing reaches the exchange.
Paper and live are swapped behind one interface, so an agent cannot tell which one it is talking to — which is the point. A strategy that works on paper runs unchanged on live.
What the simulation models
Real book depth. A large order walks levels and pays real slippage instead of filling entirely at the touch.
Partial fills when the book is too thin inside the limit price.
Taker and maker fees at Hyperliquid's base tier (0.045% / 0.015%), plus any builder fee, so the equity curve is net of what trading actually costs.
Resting orders that only fill when the market trades strictly through them, never merely to them. At your own price you are behind a queue this simulation cannot see; assuming a fill there is the most common way a paper equity curve lies.
Fills between polls. Each call checks resting orders against the one-minute candles printed since they were placed, not just the book at that moment, and fills them at the minute the market first went through. Cancelling an order that already filled is refused, as the exchange would. Before this, the demo agent's breakout bid went unfilled under a dip that lasted two minutes, and it then cancelled an order Hyperliquid would already have filled.
What it does not model
The first five flatter the result; the last two can miss a fill either way. All of them are listed rather than buried:
Latency — fills are priced off the book as it was when the tool was called
Market impact — your order never moves the price or removes liquidity
Funding payments on perps
Slippage past a stop inside a minute already gone — a stop found triggered in a past candle fills at its trigger, or at the candle's open if the market gapped through it. One triggered right now walks the real book.
Mark price — Hyperliquid triggers stops on the mark price; paper uses traded prices, which can differ briefly
The minute an order was placed in — candles are only counted from the first minute that opened after it, so a dip in that same minute is missed
Orders resting longer than about three days — only the latest 5000 one-minute candles are available, so older stretches go unchecked
Treat a paper equity curve as an upper bound on live performance, not an estimate of it.
Running a paper account
{
"env": {
"REINS_MODE": "paper",
"REINS_PAPER_BALANCE": "10000",
"REINS_PAPER_FILE": "./paper-run.json"
}
}Without REINS_PAPER_FILE the account lives in memory and is lost on restart.
Set it for anything you intend to run for more than one session — writes go to a
temp file and are renamed into place, so a crash cannot leave a half-written run
behind.
The MCP server
Nine tools. Every one that can move money passes the risk engine first, and every attempt is written to the decision log.
Tool | Notes |
| Limits, headroom per symbol, remaining loss budget, halted state |
| Signed notional per symbol, account value, today's realised PnL, stop-losses, and which positions have none |
| Best bid/ask, spread, nearest levels |
| Price history in 1m to 1d candles, the average range of one candle, and whether the newest is still forming |
| Sized in USD, not asset units — the same unit as the limits. Optional |
| Puts a stop under the whole of a position, or moves one; held by the exchange |
| By exchange order id, stops included |
| Reduce-only, crosses the spread, works even when halted |
| The agent's own recent actions and reasons, after a restart |
A blocked order comes back as a tool error with the specific reason
(BLOCKED (POSITION_TOO_LARGE): Would put BTC at $61,400, over the $25,000 cap),
so the agent can correct itself rather than retry the same rejected order. There
is deliberately no tool that changes the limits.
Orders are never true market orders — place_order without a price becomes a
marketable limit that crosses the spread by a small buffer, so a thin book can't
fill an agent at an unbounded price.
Stop-losses
An agent that says "I'll get out below 2,630" is only as good as its next wake-up. A stop-loss turns that into an order the exchange holds and executes while the agent is asleep: a reduce-only stop-market trigger that closes the whole position at market once the price reaches it.
set_stop_losssizes the stop to the entire current position and replaces any stop already on that symbol — the new one goes on before the old one comes off, so the position is never bare in between. It refuses a trigger on the wrong side of the market, which would fire at once.place_orderwithstopLossprotects the entry in the same call. An order that fills now gets a stop under the whole position straight after; if that stop fails, the fill still stands and the agent is told plainly that the position is unprotected. A resting order carries its stop with it, and the exchange places it for each part as it fills — so an entry that waits on the book at the cheaper maker fee is never bare once it fills.close_positiontakes the stop off with the position.REINS_REQUIRE_STOP_LOSS=truemakes it a limit rather than a habit: the risk engine refuses any order that adds risk without its ownstopLoss, and anything that adds risk while a position lacks a stop covering all of it (NO_STOP_LOSS). Reducing risk is never blocked.
On the wire a standalone stop is the trigger order Hyperliquid's Python SDK
signs in its tpsl test vector, and signing.test.ts reproduces that vector
byte for byte on both networks. A resting order and its stop go out together in
the normalTpsl grouping, as the SDK's basic_tpsl example sends them. The
stop's worst fill price sits 5% past the trigger — the SDK's own default
slippage — because a stop that refuses to fill protects nothing. Paper mode
fires stops from the same one-minute candles as resting orders, in the order the
market reached them; a stop attached to a fill found in a past candle watches
that same minute too, since which came first cannot be told from a candle.
Both kinds, and moving a stop, were accepted by mainnet from a funded account
in live-check --trade on 2026-09-21.
Running it
npx @r2rlabs/reins init # from a clone: npm run build && node dist/bin/cli.js initinit adds a paper-mode reins server to ./.mcp.json (keeping any other
servers there) with the limits below: $5,000 max position, $500 daily loss,
BTC and ETH, and absolute paths for the decision log and paper account. Change
any of them with flags (--symbols SOL,BTC --max-position 2500), write
elsewhere with --file, or --print the entry instead. It never writes live
mode or a key, and refuses to replace an existing reins entry without
--force. reins init --help lists everything.
Reins charges a builder fee of 2 bp on live mainnet orders, paid to
REINS_BUILDER_ADDRESS in src/builder-fee.ts. init says so when it runs,
and paper results include the fee so they match what live would cost. Before
live trading starts, the account approves the fee once from its own wallet
(npx @r2rlabs/reins approve-builder); until it has, the server refuses to
start in live mode and says how to approve. Testnet orders carry no fee.
Plug a bot in (HTTP)
A bot that does not speak MCP — Python, Go, Rust, anything that can make an HTTP request — gets the same limits and the same decision log:
npx @r2rlabs/reins http --port 8787 # prints a token; paper by defaultcurl -H "Authorization: Bearer $TOKEN" http://127.0.0.1:8787/limits
curl -H "Authorization: Bearer $TOKEN" -H "content-type: application/json" \
-d '{"symbol":"ETH","side":"buy","sizeUsd":1000,"stopLoss":2620,"reason":"Breakout retest"}' \
http://127.0.0.1:8787/ordersRoute | What it does |
| Paper or live, and the routes. The only one that needs no token |
| The limits, headroom, and whether the daily loss has halted trading |
| Exposure, account value, stops, anything unprotected |
| Top of book |
| Price history with the average range |
| The log back out, newest first |
| Place an order. |
| Set or move a stop |
| Cancel a resting order by id |
| Close a position |
Every route runs the same risk engine and writes the same records as the MCP
tools, so a refused order comes back as 400 with the limit that stopped it:
BLOCKED (TRADE_RISK_TOO_LARGE): Stopping out would lose $50.99, over the $30 allowed on one trade.
It is a trading endpoint. It binds 127.0.0.1 unless --host says
otherwise, and every route but /health needs the bearer token. Exposing it
beyond your own machine means choosing your own token and putting TLS in front
of it.
A working example in 40 lines of Python, sizing its order from the stop:
examples/bot.py.
Checking it against the real exchange
Unit tests prove Reins signs what Hyperliquid's own SDK signs. They cannot
prove Hyperliquid accepts it. reins live-check does, on a real account:
REINS_ACCOUNT_ADDRESS=0x… npx @r2rlabs/reins live-check # reads only, free
REINS_ACCOUNT_ADDRESS=0x… REINS_PRIVATE_KEY=0x… \
npx @r2rlabs/reins live-check --trade --size-usd 12 # a few cents in feesThe reads check the account, the market, open orders and whether the account
has approved Reins' builder fee. With --trade it then walks the whole order
path with the smallest position the venue allows: a post-only order with a stop
attached, a cancel, a market entry, a stop placed and moved, the position
closed, and a sweep for anything left open. It stops at the first answer that
is not what Reins expects and prints what came back instead, so a failure names
the step rather than leaving you to guess. The test size is capped at $100, and
the resting order sits 3% away so it cannot fill while the check runs. With the
position open it also re-reads the account value, which catches equity read
wrongly for the account's mode. All 12 steps have passed on mainnet, from a
Manual account and from a Unified one.
Usage
npx @r2rlabs/reins stats # the last 30 days
npx @r2rlabs/reins stats --days 365 --jsonHyperliquid publishes every fill that carried a builder code in a daily file,
once the UTC day has closed. stats reads those files for Reins' builder
address and reports accounts (and new ones by day), trades, volume, fees
earned, maker fills, stops fired and the busiest markets. It also prints the
fees Hyperliquid has credited the builder in total, which comes from the API
rather than the files and is never late: a day's file can arrive a day or more
after the day closes. Reins itself sends nothing home, so this is live mainnet
trading only: paper runs leave no trace.
Approving the fee
npx @r2rlabs/reins approve-builder # from a clone: node dist/bin/cli.js approve-builder
npx @r2rlabs/reins approve-builder --check 0x… # what a wallet has approvedThe approval must be signed by the user's main wallet, so Reins never asks
for that key. approve-builder serves a page on 127.0.0.1 behind a random
path, the user's browser wallet (MetaMask or similar) signs the EIP-712
approval there, and only the signature comes back. Reins checks that it names
Reins' builder, the requested rate and network, a fresh nonce, and that it
recovers to the wallet that says it signed — then sends it to Hyperliquid and
reads the approval back with maxBuilderFee. Signing costs no gas.
--max-fee approves a higher ceiling than the default 2 bp; --network testnet approves on testnet.
The signing was checked against Hyperliquid itself: a throwaway key's approval, sent to testnet, was refused only for having no deposit, with the error naming exactly the throwaway address — so the signature recovered correctly, including when signed at Arbitrum's chain id as a browser wallet would.
Or point an MCP client at it by hand:
{
"mcpServers": {
"reins": {
"command": "node",
"args": ["/absolute/path/to/hyperliquid-agent/dist/bin/serve.js"],
"env": {
"REINS_NETWORK": "testnet",
"REINS_SYMBOLS": "BTC,ETH",
"REINS_MAX_POSITION_USD": "5000",
"REINS_DAILY_LOSS_USD": "500",
"REINS_MAX_LEVERAGE": "3",
"REINS_MAX_ORDERS_PER_MIN": "12",
"REINS_LOG_FILE": "./decisions.jsonl"
}
}
}
}REINS_SYMBOLS, REINS_MAX_POSITION_USD and REINS_DAILY_LOSS_USD are
required — there is no default for "how much of your money may this thing lose".
Until a signer is wired the server runs read-only: reads work, trading throws. It says so on stderr at startup.
On signing
L1 actions are msgpack-hashed into a "phantom agent" and signed EIP-712. Hyperliquid's docs warn against implementing this by hand, and the warning is well earned: msgpack preserves map key order, so key order is part of the hash. Reorder two fields in an action and the signature silently becomes invalid, with nothing in the rejection pointing at ordering as the cause.
src/signing.ts implements the scheme as a set of pure functions holding no key
material, which is what makes it checkable. src/signing.test.ts verifies it
against the published test vectors from the Python SDK's own
signing_test.py — same key, same action, same nonce, and byte-identical
r, s and v on both mainnet and testnet. That is a stronger guarantee than
"one order went through once."
The things that are easy to get wrong, all covered by tests:
msgpack key order is part of the hash — there is a test that reordering two keys changes it
the nonce is 8 big-endian bytes appended to the packed action
a vault address adds a
0x01marker plus 20 bytes; no vault adds0x00sourceis"a"on mainnet,"b"on testnet — the wrong one signs perfectly and the other network refuses itthe EIP-712 domain is fixed at chainId 1337 with a zero verifying contract, regardless of what you are trading on
r and s are emitted as minimal hex with leading zeroes stripped, matching
the Python SDK exactly. This was confirmed end to end: a signed order sent to
testnet came back rejected on a price-band rule, which is order validation
and therefore only runs after authentication has already passed.
PrivateKeySigner never stores the key on the instance and overrides toJSON,
so signing authority cannot leak into a log line or a decision record.
Usage
import { HyperliquidClient, RiskEngine } from "./src/index.js";
const client = new HyperliquidClient({
network: "testnet",
signer,
builder: { address: "0xYOUR_BUILDER_ADDRESS", feeTenthsBps: 10 }, // 1 bp
});
const engine = new RiskEngine({
maxPositionUsd: 25_000,
maxLeverage: 5,
dailyLossLimitUsd: 2_500,
symbolAllowlist: ["BTC", "ETH"],
maxOrdersPerMinute: 12,
});
const snapshot = await client.positionSnapshot();
const state = { ...snapshot, realizedPnlTodayUsd: todaysPnl };
const decision = engine.check({ symbol: "BTC", side: "buy", sizeUsd: 18_400 }, state);
if (!decision.allowed) throw new Error(decision.reason);
const outcome = await client.placeOrder({
symbol: "BTC",
side: "buy",
size: 0.18,
price: 102_000,
});
engine.recordOrder();Construct it without a signer and the client is read-only: market data works,
placeOrder throws.
Development
npm install
npm testDesign
Product screens live in design/ and render as a canvas at the
Artifact linked in the project notes.
Builder code facts, verified against Hyperliquid's docs
Fee cap, perps | 0.1% of fill value |
Fee cap, spot | 1% of fill value |
Fee units | Tenths of a basis point — |
Builder requirement | ≥100 USDC perps account value, |
User approval |
|
Approvals per user | 10 active at a time |
Scope | Both sides of perps; sell side only on spot |
Payout | Claimed through the normal referral reward process |
Safety notes
The agent gets a
get_limitstool so it can see its constraints. It gets no tool that can change them.Reduce-only orders survive a halt. Closing risk is always permitted; opening it is not.
Without a private key the server runs read-only — market data works, nothing can trade.
Available Tools
9 toolscancel_orderCancel an orderB
Cancel a resting order by its exchange order id.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Why you are cancelling, if not obvious. | |
| symbol | Yes | ||
| orderId | Yes | The oid returned when the order was placed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the entire behavioral burden, and it discloses almost nothing beyond the verb. It does not say whether cancellation is irreversible, what happens if the order is already filled or partially filled, whether auth/permissions are required, or what the call returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler — the identifying information comes first. It is perhaps too terse for a mutation tool, but there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a state-changing operation with no annotations, no output schema, and an undocumented required parameter. The description omits failure modes, irreversibility, and required permissions, so an agent lacks enough context to call it confidently in edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% (orderId and reason are documented inline), and the description only re-frames the orderId concept already covered by the schema. The required 'symbol' parameter has no description in either the schema or the description, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (cancel) plus resource (resting order) and the keying mechanism (exchange order id). It distinguishes itself from close_position and set_stop_loss by operating on orders rather than positions, but never names or contrasts a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The qualifier 'resting' implies the tool applies only to live/unfilled orders, which is a mild scoping constraint, but there is no stated when-to-use, no alternative (e.g., close_position for filled exposure), and no guidance on what to do if the order is already gone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_positionClose a positionA
Flatten the whole position in one symbol with a reduce-only order that crosses the spread. This reduces risk, so it is permitted even when the daily loss limit has halted new trading.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Why you are closing this position, in one or two sentences. | |
| symbol | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the order is reduce-only and crosses the spread (i.e. it will take liquidity rather than rest), and that it is exempt from the loss-limit halt. It does not cover failure modes (e.g. no open position), irreversibility, or what the response contains, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core action is front-loaded ahead of the permission nuance. Each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description supplies the essentials an agent needs to act safely: the order type, the price behavior, and the halt exemption. Remaining gaps are edge cases (no position to close, partial-fill behavior) rather than core missing context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'reason' is documented in the schema, but 'symbol' has no description anywhere. The prose clarifies scope ('in one symbol', closing the whole position rather than partially), which adds a little, but it does not compensate for the undocumented symbol parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('flatten the whole position in one symbol') and names the execution mechanism ('a reduce-only order that crosses the spread'), which implicitly distinguishes it from place_order. An agent can tell exactly what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete use condition: this is a risk-reducing action, so it remains permitted even when the daily loss limit has halted new trading. That is real routing guidance, though it never names place_order as the alternative for opening/increasing exposure, so the when-not-to-use side is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bookGet order bookB
Top of book for a symbol: best bid, best ask, spread, and the nearest levels on each side.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | How many levels per side to return. Defaults to 5. | |
| symbol | Yes | Perp symbol, for example BTC or ETH. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden alone. It does disclose the return payload (bid, ask, spread, level depth), which is meaningful value given there is no output schema. However, it says nothing about read-only semantics, latency/streaming, ordering, or rate limits, leaving behavioral traits thin for a wholly un-annotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the primary artifact first and then enumerates contents. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, describing the returned fields is exactly the right compensation and it is done competently. The remaining gap is the absence of any cue about read-only nature or data freshness for a two-parameter market-data call, which keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: 'symbol' is documented as a perp symbol with examples and 'depth' carries its own description plus a default of 5. The phrase 'nearest levels on each side' loosely corroborates the depth parameter but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (order book for a symbol) and enumerates the content: best bid, best ask, spread, nearest levels per side. This clearly distinguishes it from position, limits, and candle getters by subject matter. It stops short of naming a sibling for contrast, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus get_candles or get_limits, and no stated preconditions. Usage is only inferable from the word 'book' and the required symbol parameter. No exclusions or alternatives are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_candlesGet price candlesB
Recent price history for a symbol as candles, oldest first: open, high, low, close and volume for each interval. The newest candle is usually still forming — lastCandleComplete says whether it has closed. averageRange is the mean high-to-low of the complete candles: how far price typically moves in one interval.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | How many of the most recent candles to return. Defaults to 24. | |
| symbol | Yes | Perp symbol, for example BTC or ETH. | |
| interval | No | Length of each candle. Defaults to 1h. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the newest candle is usually still forming and that lastCandleComplete signals closure, which is real behavioral context. It says nothing about permissions, rate limits, or whether data older than 'recent' is reachable, leaving notable gaps for a read tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no waste: the first defines the payload, the second handles staleness, the third defines averageRange. Content is front-loaded and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain return values, and it does so well (OHLCV per interval, oldest-first ordering, lastCandleComplete, averageRange). Minor gaps remain, e.g. behavior for an unknown symbol or how 'recent' relates to the count cap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so symbol, count (default 24, max 200) and interval (default 1h, enum values) are already fully documented in the schema. The description only echoes 'for each interval' and adds no syntax or semantic detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Recent price history for a symbol as candles') and enumerates the returned fields (open, high, low, close, volume), so an agent knows exactly what this fetches. It does not, however, contrast itself with siblings like get_book or get_limits, so the differentiation work is left to the reader.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as get_book (order book depth) or get_limits, nor any prerequisites or exclusions. Context only weakly implies 'fetch market data', which is not enough routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_limitsGet trading limitsA
Report the risk limits this account trades under and how much room is left against each of them. These limits are enforced outside you and no tool changes them — check here before sizing an order rather than discovering a limit by being rejected.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does real work: it discloses that the limits are enforced outside the agent and that no tool modifies them, which tells the agent this is a read-only, externally-governed resource. It stops short of describing return shape or freshness/staleness of the limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded with the operative information (what it reports, then when to call it). No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter read tool with no output schema, the description covers purpose, the key return concept (limits plus remaining headroom), and the critical behavioral fact that limits are externally enforced and immutable. It does not enumerate the specific limit fields returned, which is a minor gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. Schema coverage is 100% and the empty input object needs no further explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report) and resource (risk limits this account trades under) plus the derived value it returns (room left against each). It is clearly distinguishable from siblings like get_positions or get_book, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete workflow guidance: check here before sizing an order rather than discovering a limit by being rejected. That establishes when to use it relative to the order-placement siblings, but it does not name place_order or state any conditions for when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_positionsGet open positionsA
Open positions with their signed notional in USD (negative is short), account value, and realised PnL so far today.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden; it does disclose useful data semantics (negative = short, PnL scoped to today) that the schema cannot convey. However, it never states that this is a read-only operation, nor anything about auth, rate limits, or how positions are scoped to an account.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the key field semantics front-loaded; every clause earns its place and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only lookup with no output schema, the description does the essential work of telling the agent what fields come back and how to interpret them. It stops short of covering account scoping or the read-only nature, which matters given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate — the baseline for a no-parameter tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (open positions) and enumerates the data returned — signed notional in USD, account value, and today's realised PnL — so an agent knows exactly what this tool fetches. It omits an explicit verb and does nothing to differentiate itself from the overlapping get_book sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives such as get_book or get_limits, and no stated prerequisites or context. The agent must infer usage entirely from the resource name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recent_decisionsGet recent decisionsA
Your own recent actions and the reasons you gave for them, most recent first, including any that the risk limits refused. Useful after a restart, when you no longer remember what you already did. Records with tool exchange_fill are what the exchange did without you: a stop-loss firing, or a resting order filling after you placed it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many records to return. Defaults to 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses ordering (most recent first), that risk-limit-refused actions are included, and that exchange_fill records represent exchange-side events rather than agent decisions — a non-obvious distinction that prevents misinterpretation. It stops short of covering auth, rate limits, or truncation beyond the limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what is returned before the usage hint and the exchange_fill caveat. Sentence two is slightly conversational but still earns its place as the usage trigger.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey return semantics, and it does: record contents, ordering, inclusion of refused actions, and the meaning of exchange_fill entries. The only gap is the absence of any mention of volume/pagination behavior beyond the limit cap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema coverage is 100%, with the schema stating the integer type, range, and default of 10. The description adds nothing about the limit parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (your own recent actions) and a distinguishing attribute (the reasons you gave), plus ordering. An agent can immediately tell this apart from state-lookup siblings like get_positions, get_limits, or get_book, which report current state rather than a decision log.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Useful after a restart, when you no longer remember what you already did" gives a concrete triggering context. There is no explicit when-not or named alternative, but the usage context is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
place_orderPlace an orderA
Place an order sized in USD notional. Every order is checked against the risk limits first; if it breaches one it is refused and nothing reaches the exchange. Omit price for a marketable order that crosses the spread. Both the order and your stated reason are written to a permanent log, including when the order is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| tif | No | Time in force. Defaults to Ioc when marketable, Gtc otherwise. | |
| side | Yes | ||
| price | No | Limit price. Omit to cross the spread and fill now. | |
| reason | Yes | Why you are placing this order, in one or two sentences. State the signal or condition you are acting on and why this size. A human will read this later to understand what you were doing, so write what actually drove the decision rather than a generic summary. | |
| symbol | Yes | Perp symbol, for example BTC. | |
| sizeUsd | Yes | Notional size in USD, not asset units. | |
| stopLoss | No | Trigger price of a stop-loss for this order: below the entry for a buy, above it for a sell. An order that fills now gets a stop under the whole position; a resting order takes its stop with it, placed as it fills, so the position is never unprotected. How far the stop sits decides what the trade risks (distance to stop x size), which get_limits reports as maxTradeRiskUsd: a wider stop needs a smaller order. | |
| reduceOnly | No | True if this order may only shrink an existing position. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that every order is risk-checked first, that a breach causes outright refusal with nothing reaching the exchange, and that both the order and the reason are written to a permanent (irreversible) log even when refused. It does not cover auth/permissions or what happens to partial fills, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose and followed by the most decision-relevant behaviors (risk refusal, marketable order, permanent logging). Slightly redundant with the schema's own wording on price, but no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no output schema, the description covers the critical unknowns: the risk gate, the refusal path, and the logging side effect. The main gap is that, with no output schema, it says nothing about what a successful or rejected call returns, and it does not clarify stopLoss-vs-set_stop_loss interaction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 88%, so the schema already documents almost every parameter in depth (price, tif, stopLoss, reason, sizeUsd). The description's 'sized in USD notional' and 'omit price for a marketable order' restate facts already present in the property descriptions rather than adding new meaning, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb plus resource ('Place an order') and immediately narrows the semantics with 'sized in USD notional', which distinguishes it from the close_position/cancel_order siblings. An agent can identify exactly what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives in-tool conditional guidance ('Omit price for a marketable order that crosses the spread'), which is genuinely useful. However, it never addresses how this tool relates to overlapping siblings such as set_stop_loss (despite offering a stopLoss parameter) or close_position, so the when-to-use-this-vs-alternatives question is only partly answered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_stop_lossSet a stop-lossA
Protect an open position with a stop-loss held on the exchange: once the price reaches triggerPrice, the whole position is closed at market, whether or not you are running at the time. Replaces any stop already on that symbol. Below the market for a long, above it for a short. Logged with your reason, like an order.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Why this level: what would have to happen for your view to be wrong, in one or two sentences. A human will read this later. | |
| symbol | Yes | Perp symbol with an open position, for example ETH. | |
| triggerPrice | Yes | Price at which the position is closed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: the stop is held exchange-side, fires whether or not the user is running, closes at market (slippage implication), is directional per long/short, replaces existing stops, and is logged with a reason. This is exactly the behavioral context an agent needs for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with what the tool does and the exchange-held guarantee. Every clause carries information: mechanism, replacement behavior, directional rule, and audit logging. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers mechanism, side semantics, replacement, and logging, and there is no output schema to explain. It does not say what happens on failure (e.g., missing position, invalid trigger relative to market), which is a minor gap for a mutation tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds directional semantics for triggerPrice ('below the market for a long, above it for a short') that the schema does not capture, plus the framing of reason as a human-readable rationale. That goes beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (set a stop-loss) and immediately explains the mechanism: an exchange-held stop that closes the whole position at market when triggerPrice is hit. This clearly differentiates it from place_order and close_position, both of which appear as siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the use case ('protect an open position') and adds a critical replacement caveat ('Replaces any stop already on that symbol'), which prevents accidental overwrites. It does not explicitly contrast with place_order/close_position, so an agent must infer that a market order is not the right tool for protective stops.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.6- First observed
cancel_order - First observed
close_position - First observed
get_book - First observed
get_candles - First observed
get_limits - First observed
get_positions - First observed
get_recent_decisions - First observed
place_order - First observed
set_stop_loss
TDQS
Scored across 9 tools
Each tool has a clearly distinct purpose: market data (get_book, get_candles), account state (get_limits, get_positions), order actions (place_order, cancel_order, close_position, set_stop_loss), and audit (get_recent_decisions). While place_order, set_stop_loss, and close_position all involve orders, their descriptions clarify the specific action and constraints, so an agent can easily select the right one.
All tool names use consistent snake_case with a verb_noun structure (get_, place_, set_, cancel_, close_). No deviations or mixed conventions, making the naming predictable.
The server provides 9 tools covering market data, account information, order management, risk limits, and an audit log. This is a well-scoped set for a trading agent—neither too thin nor too heavy—with each tool earning its place.
Core trading workflows are covered: checking limits, viewing positions, placing/canceling orders, setting stop-losses, closing positions, and reviewing recent decisions. However, there is no direct way to list currently open/resting orders or query a specific order's status, which are notable gaps for order management, though agents may partially work around this via the decisions log.
Related MCP Connectors
Trade across 22+ exchanges and brokers from any MCP-capable AI agent, no install required.
No-KYC managed MCP for AI agents: sandboxed TypeScript trading SDK, isolated sub-accounts, futures.
Trade 16 crypto exchanges + MetaTrader 5 from your AI assistant via one MCP connection.
Live prices, perps, prediction markets and a paper trading desk over one MCP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to securely trade on Hyperliquid perpetual exchange, including order placement, position management, market data retrieval, and vault operations via natural language.31 PyPI21MIT
- AlicenseAqualityCmaintenanceEnables natural language control of Hyperliquid perpetual futures, including querying positions, prices, orderbook, and executing trades like market and limit orders, all from MCP-compatible clients.13MIT
- AlicenseAqualityDmaintenanceEnables AI agents to interact with Hyperliquid perpetual futures exchange for market analysis, account management, and risk-managed trading.51MIT
- AlicenseAqualityAmaintenanceProvides AI agents read-only, keyless access to Hyperliquid market data, funding rates, account risk, and HyperEVM token transfers through MCP tools, with caching and rate limiting to protect upstream APIs.121MIT