lob-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@lob-mcpSubmit a limit order to buy 10 AAPL at $150 and show me the resulting book."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
lob-mcp — a limit order book exchange as an MCP server
A deterministic limit order book matching engine exposed to language models as a set of MCP tools. The point is that a model can run real experiments against an exchange — build a book, cross the spread, provoke a rejection, roll the venue back and try the other branch — entirely through tool calls, with nobody editing code between steps.
Not a real venue: no money, no connectivity, no fees, no latency model.
What MCP is, in one section
The Model Context Protocol is a JSON-RPC 2.0 protocol that lets a model's host application talk to an external process that supplies capabilities.
The server (this repo) is a process that exposes capabilities and answers requests. It never initiates work.
The client (Claude Code, MCP Inspector, your own script) launches the server and drives it.
Under the stdio transport, the client spawns the server as a subprocess and speaks JSON-RPC over its stdin and stdout. No ports, no auth, no CORS, and the process lives exactly as long as the client session. That is why it is the default for local tools — and why stdout is the wire: a stray
print()corrupts the protocol. Every log line in this server goes to stderr.The initialize handshake is the first exchange: the client sends its protocol version and capabilities, the server replies with its own plus its name and instructions, and the client confirms with an
initializednotification. Only afterwards may either side send anything else.
A server can expose three kinds of thing, and the distinction is about who decides:
Primitive | Controlled by | In this project |
Tool | the model — a verb it chooses to invoke, possibly with side effects |
|
Resource | the application — a noun addressed by URI that the host may attach to context; read-only by contract |
|
Prompt | the user — a template a person invokes, e.g. a slash command |
|
The book is exposed both as a resource and as a tool, on purpose.
book://AAPL is "attach the current state so the model starts the turn
knowing it"; get_book(symbol, depth) is "the model decided it needs to
look, at a depth it chose, right now". Same data, different control.
Related MCP server: MetaTrader MCP Server
Quick start
git clone https://github.com/pranjalsharma-6/mcp.git && cd mcp
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest -q # 65 tests
python -m evals.demo # a full scripted session
python -m evals.run_eval --mode referenceAdd it to Claude Code
claude mcp add lob -- python -m mcp_server.serveror, in .mcp.json / claude_desktop_config.json:
{
"mcpServers": {
"lob": {
"command": "python",
"args": ["-m", "mcp_server.server"],
"cwd": "/absolute/path/to/mcp",
"env": { "PYTHONPATH": "/absolute/path/to/mcp", "LOB_SEED": "0" }
}
}
}Read-only: add "LOB_READ_ONLY": "1" to env, or pass --read-only.
Debug with MCP Inspector
npx @modelcontextprotocol/inspector python -m mcp_server.serverInspector opens a browser UI that performs the handshake and lets you list and call tools by hand, with the raw JSON-RPC visible. Use it before wiring the server into a model — it separates "my tool is broken" from "the model called it wrong", which are very different bugs.
Debugging notes. The server logs to stderr at INFO; if the client shows
nothing at all, run the command by hand and check it starts. If the handshake
succeeds but every call fails, you almost certainly have a PYTHONPATH
problem, since the client's working directory is not yours. If the handshake
itself fails, suspect something writing to stdout.
Tool surface
Read (always registered)
Tool | Signature | Notes |
|
| Aggregated levels, best first. Depth capped at 20; truncation is announced in the output. |
|
| Status, filled/remaining, trade ids. |
|
| Page forward with |
|
| Symbols and their tick/lot rules, spreads, volume, rejection tally. |
Mutating (registered only when not read-only)
Tool | Signature | Notes |
|
| Idempotent on |
|
| Idempotent: re-cancelling returns |
|
| Reports whether queue priority was kept or lost. |
|
| Destructive. Clears book, tape and the id ledger. |
|
| Named starting books. No args lists them. |
|
| Returns a snapshot id. |
|
| No args lists available snapshots. |
Resources
docs://matching-rules · book://{symbol} · tape://{symbol}
Prompts
explain_order(order_ref) — trace one order's life from the tape.
stress_book(symbol) — systematic edge cases, report anomalies. Withheld in
read-only mode, since it instructs the model to submit orders.
Agent-facing API design notes
The surface is designed for a language model, not a browser. That changes things.
Output is bounded, and truncation announces itself
get_book returns top-N levels, capped at 20; get_trades caps at 100. Every
truncated response ends with a line saying what was withheld and which
argument reveals more:
[showing depth 3: 2 more bid level(s), 2 more ask level(s) not shown. raise `depth` (max 20) to see them]Without that line, a model shown 3 of 12 levels will state, confidently and wrongly, that the book is 3 deep. Silent truncation is worse than a small window.
Every result carries text and structure
Tools return both a compact text rendering and structuredContent. The text
is what the model skims; the JSON is what code indexes. Columns are aligned
because alignment is what lets a model compare adjacent levels without doing
arithmetic:
BOOK AAPL seq=10 bid=99.99 ask=100.01 spread=0.02
BIDS | ASKS
-------------------------------------------------------
99.99 500 (1) | 100.01 300 (1)
99.98 1200 (1) | 100.02 900 (1)Rejections are data, not exceptions
{
"status": "REJECTED",
"reason": "TICK_SIZE_VIOLATION",
"message": "price 100.005 has more decimal places than AAPL supports (price_scale=2)",
"field": "price",
"suggestion": "round to 2 decimals, e.g. 100.00",
"suggested_value": "100.00"
}isError is deliberately not set. isError means "the tool
malfunctioned" — a client may retry it, log it as a fault or hide it. A
tick-size violation is not a malfunction; it is the correct answer to an
incorrect order, and the model has to read it and adapt.
suggested_value exists because of a finding from the evals: recovering from
a rejection used to require float(suggestion.split()[-1]), pulling a number
out of an English sentence. The prose is for the reader; the value is for the
retry.
Two failure classes, kept distinct
A schema violation (negative qty, side="SIDEWAYS") never reaches the
engine — the MCP layer rejects it and it surfaces as isError. A venue
rejection is a well-formed call the exchange declines, and comes back as an
ordinary result. Keeping these apart is what lets a model tell "I called the
tool wrong" from "the venue said no".
Anything expressible in the JSON Schema lives there, because the model reads the schema before calling — a constraint stated in the schema costs zero tool calls, the same constraint discovered through a rejection costs a round trip.
Idempotency, and the failure mode it prevents
submit_order takes a caller-supplied client_order_id. The server
stores the result of the first call bearing each id; a repeat returns that
stored result marked DUPLICATE and creates nothing.
This is not the HTTP retry story. In HTTP, retries come from network layers you control. Here the retry comes from the model, and models retry on ambiguity, not just on error: the result was truncated, the turn was interrupted, or the plan step "submit the buy" got re-read after context was compacted. Without an id, that retry is a second real order and you have bought 200 instead of 100 — and the model is now reasoning about a book that does not exist, so every conclusion afterwards is wrong, silently.
The id must be model-supplied. A server-generated id cannot deduplicate a call whose response the caller never saw.
Two consequences worth knowing:
A rejected submit also consumes its id. Retrying the identical wrong call is therefore always safe, and a corrected order is forced to be a genuinely new order.
Reusing an id for a different order silently returns the old result. That is the sharp edge of the design; the tool description says so explicitly and
test_a_different_order_under_a_reused_id_is_not_submittedpins it.
cancel_order takes the other route to idempotency: cancelling an
already-cancelled order returns NOOP, not a rejection, so a model can reach
a known state without first working out whether its earlier cancel landed.
Read-only mode unregisters rather than refuses
With LOB_READ_ONLY=1 the mutating tools are not registered. They are
absent from tools/list — 11 tools become 4.
Refusing at call time is the weaker design:
A registered tool is in the model's context. It plans with it, calls it, reads the refusal, and retries with different arguments — because refusals read like "you did that wrong", not "this is impossible". You pay tokens and turns litigating something that was never negotiable.
An unregistered tool cannot be planned with at all. The model sees a read-only venue at handshake time and builds a read-only plan from its first move. Constraints are cheapest when visible before planning.
Authorisation by absence of a code path beats authorisation by an
if. There is no branch to get wrong and no future refactor that moves the check after the mutation.
The cost: flipping the flag needs a restart, since clients cache tools/list
after the handshake. For a per-process stdio server that is the right trade.
Tool descriptions are prompts
Every mutating tool's description ends with an explicit Side effects: block,
enforced by a test. A model plans by reading descriptions, and a tool whose
description does not say it mutates shared state will be called
speculatively, mid-plan, to "check" something.
The descriptions also carry the non-obvious semantics. modify_order spends a
paragraph on queue priority — that a price change or size increase re-queues
the order at the back of its level while a size reduction keeps its place —
because a model will otherwise treat modify as free.
Determinism
Nothing reads a clock or an unseeded random source. Events are ordered by a
monotonic integer seq, order ids come from a counter, and money is integer
ticks internally with decimals only at the render boundary — float prices
produce fills off by 1e-13 and comparisons that flip with operand order, which
would quietly destroy reproducibility. take_snapshot / restore_snapshot
round-trip the whole mutable state, so two experiment arms can start from an
identical book rather than a reconstructed one.
test_the_same_call_sequence_replays_identically asserts it.
Layout
engine/ pure Python matching engine. No MCP imports.
types.py Order, Trade, Side, TIF, SymbolSpec, tick arithmetic
book.py price levels, FIFO queues
engine.py matching loop, validation, snapshot/restore
errors.py structured rejections
mcp_server/
server.py registration, READ_ONLY gate, transport bootstrap
tools_read.py always registered
tools_write.py registered only when not read-only
resources.py book, tape, matching-rules doc
prompts.py explain_order, stress_book
render.py the LLM-facing formatting layer
schemas.py pydantic models -> published JSON Schema
config.py get_engine(session_id) indirection
tests/ 65 tests, driven through the MCP client SDK
evals/ 4 tasks, graders, reference solutions, demo
docs/ matching rules, captured demo sessionThe engine never imports MCP, and the tool layer never imports the transport.
build_server() assembles capabilities; main() picks stdio. Adding
streamable-HTTP is a change to main() plus handing get_engine() a real
session id instead of a constant — no tool changes. That indirection exists
from day one because the thing HTTP actually breaks is state scope: stdio
means one engine per process, HTTP means many clients sharing one.
Tests
pytest -q # 65 passedTests drive the server through the real MCP client SDK, not by calling
tool functions, so schema validation, result serialisation and the handshake
are all exercised. test_stdio.py additionally runs the server as a real
subprocess over stdio — the only way to catch a stray write to stdout.
Two bugs the suite caught while being written:
Retrying a submit that was rejected the first time crashed, because the duplicate path assumed a stored order existed. Precisely the case the design promises is safe.
Float representation error (
99.99 - 0.02 == 99.97000000000001) was rejected as sub-tick. Now absorbed, while a genuine sub-tick price like100.005still rejects.
Evals
python -m evals.run_eval --mode reference # 4/4
python -m evals.run_eval --mode model # needs ANTHROPIC_API_KEYFour tasks with programmatic graders that read the tool-call log and engine
state, never the model's prose as evidence of what it did. idempotent_retry
fails a run where every submit used a distinct id even if the final volume is
correct — getting the right answer by luck is not passing.
Reference mode passes 4/4. The model-in-the-loop run has not been
performed — the environment this was built in has no API key — and that is
the mode that actually answers whether the descriptions are unambiguous, since
a scripted solution never reads them. See
evals/FINDINGS.md for the two ambiguities that writing
the reference solutions exposed, both since fixed, and for the description most
likely to fail a real run.
Demo
docs/demo_session.txt is a captured 13-step session:
learn the venue, build a book, get rejected twice and recover, snapshot, sweep
the offers, retry the same id without double-filling, roll back. Regenerate it
with python -m evals.demo.
It is a scripted session, not a recording of Claude driving the exchange —
same tool calls, same real responses, but the driver is a script. A recording
of a model needs --mode model with an API key.
Available Tools
11 toolscancel_orderCancel orderADestructiveIdempotent
Remove an order's remaining quantity from the book. Already-filled quantity is not undone -- a cancel only affects what has not traded yet.
Identify the order by EITHER order_id OR client_order_id.
Safe to retry: cancelling an order that is already cancelled or filled returns status NOOP, not a rejection, so you can reach a known state without first working out whether your earlier cancel landed.
Side effects: MUTATES the book. Removes resting liquidity, which widens the spread or empties a level. Never creates trades.
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | No | ||
| client_order_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive and idempotent behavior, but the description adds valuable context beyond those hints: it specifies that the book is mutated, resting liquidity is removed, the spread may widen or a level may empty, and no trades are ever created. It also discloses the NOOP retry semantics in concrete terms, giving the agent a precise model of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into three tight, purposeful sections: core operation semantics, parameter identification, and retry/side-effect behavior. Every sentence earns its place, with no filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter mutation tool with no output schema, the description covers all essential invocation knowledge: what the tool affects, how to identify the target, retry safety, and market impact. There are no significant gaps that would prevent an agent from calling it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by stating that the order is identified by EITHER order_id OR client_order_id. This is the core semantic needed to invoke correctly and clarifies that the parameters are alternative identifiers. It doesn't detail formats or precedence, but the parameter names are self-explanatory and the either/or guidance is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Remove an order's remaining quantity from the book') and clearly differentiates this from siblings like submit_order and modify_order by emphasizing it only affects untraded quantity. It also clarifies what a cancel does not do (undo filled quantity), leaving no ambiguity about the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: to remove remaining resting quantity, and when it is safe to retry, noting that already-cancelled/filled orders return NOOP rather than rejection. It doesn't explicitly name alternative tools or exclusions, but the combination of purpose clarity and sibling names makes the appropriate usage obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bookGet order bookARead-onlyIdempotent
Show the top price levels of one symbol's order book, best price first.
Returns aggregated levels, not individual orders: each row is a price with the total resting quantity there and how many orders are queued at it. Use get_order to inspect a specific order.
Output is deliberately bounded. If levels exist beyond depth the response says so explicitly -- do not assume the book ends where the listing ends.
Prices come back as decimal STRINGS ('100.01'), exactly as the venue quotes them. Pass them straight back to submit_order or modify_order unchanged -- those tools accept a string price, and doing so avoids introducing floating-point error.
Side effects: none. Reading the book never changes it.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | How many price levels per side to return. Capped at 20. | |
| symbol | Yes | Instrument symbol, e.g. 'AAPL'. Call get_stats to list what is tradable. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already state readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds meaningful behavioral detail beyond that: returns aggregated levels rather than individual orders, output is bounded with explicit indication if levels exist beyond depth, prices are returned as decimal strings, and side effects are explicitly stated as none.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then each paragraph adds essential behavioral, formatting, or side-effect information. There is no filler; every sentence earns its place, and the structure is logical from purpose to return semantics to safety.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description carefully explains the return format (aggregated rows, price/quantity/order count, decimal strings) and the bounded nature of results. It also clarifies side effects and reinforces correct downstream usage with submit_order/modify_order, making it fully actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The schema already documents depth's range and default, and symbol's meaning. The description adds useful output-related context, like price string formatting and depth behavior, but it does not materially change how an agent interprets the parameters themselves beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Show the top price levels of one symbol's order book, best price first.' It clearly distinguishes this from inspecting a specific order via get_order, and the scope is explicit (one symbol, aggregated levels).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly directs the agent to 'Use get_order to inspect a specific order,' providing a clear alternative and differentiating the use cases. The symbol parameter also guides the agent to call get_stats to discover tradable symbols, so selection context is well covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_orderGet orderARead-onlyIdempotent
Look up one order's current state: status, how much filled, how much remains, whether it is still resting on the book, and the ids of the trades it took part in.
Identify the order by EITHER order_id (the venue's id, like 'O7') OR client_order_id (the id you supplied when submitting). Supply at least one.
This is the tool to use after a submit whose result you did not see -- it tells you what actually happened without changing anything.
Side effects: none.
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | No | ||
| client_order_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnly, idempotent, and non-destructive hints, and the description reinforces this with 'without changing anything' and 'Side effects: none.' It also adds useful behavioral context about identifying an order by either venue or client id, going beyond the structured annotation data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the most important information: what the tool returns. Each sentence serves a purpose: return fields, parameter identification, usage scenario, and side-effect statement. No waste or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for a read-only lookup tool. It explains what data is returned, how to identify the target order, when to use it, and confirms it has no side effects. Despite lacking an output schema, the description enumerates the return fields, so an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden of explaining parameters. It clearly defines order_id as the venue's id with an example, defines client_order_id as the id supplied when submitting, and states that at least one must be supplied. This fully compensates for the empty schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb 'Look up' and a clear resource: 'one order's current state,' then enumerates the exact fields returned. It distinguishes itself from sibling tools like get_trades and get_stats by focusing on a single order's state rather than trade lists or market statistics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: after a submit whose result was not seen, and notes it reports what happened without changing anything. It doesn't explicitly name alternatives or list when-not-to-use conditions, but the context is clear enough for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statsGet venue statsARead-onlyIdempotent
Venue-wide summary: which symbols exist and their trading rules, best bid/ask and spread per symbol, resting depth, traded volume, and a tally of how many orders were rejected for each reason.
Call this first in a fresh session -- it is the cheapest way to learn what is tradable and what the tick and lot sizes are before you submit anything.
Side effects: none.
| Name | Required | Description | Default |
|---|---|---|---|
| symbol | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only, idempotent, and non-destructive. The description adds 'Side effects: none' and a 'cheapest' cost hint, consistent with the annotations, but it does not discuss auth, rate limits, or unexpected behavior. With annotations covering safely, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the core capability, add a direct usage recommendation, and close with a side-effect note. Every sentence serves a purpose with no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description details the return contents and the recommended invocation timing, which is useful since there is no output schema. However, it leaves the only input parameter (symbol) unexplained, creating a notable gap for a tool with minimal parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'symbol' parameter has 0% schema description coverage, and the description never mentions that parameter or explains whether it filters the summary or whether null returns all symbols. The description fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence defines the resource as a venue-wide summary and enumerates the exact contents (symbols, trading rules, quotes, depth, volume, rejection counts). This clearly distinguishes it from order, trade, and book-specific siblings like get_order and get_book.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second paragraph explicitly instructs to call it first in a fresh session and frames it as the cheapest way to learn tradability and tick/lot sizes before submitting anything. It gives clear context for when to use it, though it does not name specific alternative tools for alternate needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_tradesGet trade tapeARead-onlyIdempotent
Read the trade tape: executions in the order they happened.
Page through it with since_seq: pass 0 for the beginning, then pass the seq of the last trade you saw to get only what is new. The response tells you when it truncated and which since_seq to use next.
Trades print at the MAKER's resting price, not the aggressor's limit. An order that crosses with a generous limit still pays only the prices already on the book.
Side effects: none.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| symbol | No | ||
| since_seq | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly, idempotent, non-destructive), the description discloses non-obvious behavior: trades print at the maker's resting price, not the aggressor's limit, and paging handles truncation with a next since_seq. It also explicitly states 'Side effects: none', which alignment with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with purpose, and organized into short focused paragraphs: what the tool returns, how paging works, and a critical pricing nuance. Each sentence contributes useful, non-redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description provides purpose, paging protocol, maker-price semantics, and side-effect disclosure. It does not enumerate the response fields or clarify limit/symbol behavior, but an agent can still invoke the tool correctly and interpret the core pagination response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since_seq receives excellent semantic treatment: it is a cursor, 0 means the beginning, and the last seen seq continues the stream. However, limit and symbol are not described, and with 0% schema_description_coverage the description must compensate for all parameters; it does not fully do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Read') and resource ('trade tape') and defines executions as appearing in the order they happened. It is distinguishable from sibling tools by the concept of a trade tape, though it does not explicitly point out how it differs from get_book or get_stats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage instructions for paging: start with since_seq=0, continue with the last seen seq, and expect truncation signaling. It does not explicitly compare itself to alternative tools or state when not to use it, but the intended context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_scenarioLoad scenarioADestructive
Build a known starting book in one call, by replaying a fixed list of orders through the normal submit path. Use this instead of hand-placing ten orders when you want a reproducible starting point.
Call with no arguments to list the available scenarios and what each one is for.
Scenarios are deterministic: the same name on a freshly reset book always produces the same levels and the same sequence numbers.
Side effects: MUTATES the book. Resets the target symbol first, then submits the scenario's orders. Any existing orders on that symbol are destroyed.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| symbol | No | AAPL |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only state readOnly=false, idempotent=false, destructive=true, but the description goes further: it mutates the book, resets the target symbol first, destroys existing orders, is deterministic on a freshly reset book, and replays orders through the normal submit path. This meaningfully enriches the safety profile without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: core action, use case, no-args discovery, determinism, then side effects. Every sentence adds useful information and none merely restates the tool name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with no output schema, the description covers the critical side effects, determinism, and how to list scenarios. It does not specify the return format or exact scenario names, but the no-args listing path mitigates that gap. Slightly more explicit parameter detail would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter burden. It implies that 'name' selects a scenario, that 'symbol' is the target book, and that no arguments lists scenarios, but it does not explicitly define the parameter roles or value formats. This is partial compensation rather than full clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: load/build a known starting book by replaying a fixed order list through the normal submit path. It is clearly distinguishable from siblings like reset_book and submit_order because it is scenario-driven. The no-args listing behavior is also included.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this instead of hand-placing ten orders when a reproducible starting point is needed, and tells the agent to call with no arguments to discover available scenarios. It does not name exact alternative sibling tools, but the usage context is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
modify_orderModify orderADestructive
Change a resting order's price and/or quantity in place.
QUEUE PRIORITY -- this is the part that surprises people. Orders at the same price are filled in arrival order. Changing the PRICE, or INCREASING the quantity, re-queues the order at the BACK of its price level: you lose your place to everyone already there. Only REDUCING the quantity keeps your position. The response states which happened.
A modify that moves the price across the spread will match immediately, exactly as a new order would.
Side effects: MUTATES the book. May create trades. May lose queue priority irreversibly -- restoring the old price does NOT restore the old position.
| Name | Required | Description | Default |
|---|---|---|---|
| new_qty | No | ||
| order_id | No | ||
| new_price | No | ||
| client_order_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the destructiveHint annotation by explaining queue-priority loss, when position is preserved, immediate matching across the spread, book mutation, trade creation, and irreversibility. This is exactly the kind of behavioral nuance an agent needs to predict side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place. The core action is front-loaded, the surprising queue-priority behavior is flagged clearly, and the side-effect warning is compact but complete. The formatting guides attention to the most consequential facts.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers the critical behavioral outcomes including what the response indicates. It lacks explicit statement that order_id selects the target order and whether both fields can be omitted, but overall it is thorough enough for correct invocation in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description enriches the meaning of new_price and new_qty with queue-priority and matching consequences, which is valuable given 0% schema coverage. However, it does not explicitly explain order_id or client_order_id, leaving partial burden on the agent despite the self-explanatory parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Change a resting order's price and/or quantity in place." This clearly distinguishes it from sibling tools like submit_order and cancel_order while precisely stating what operation is performed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this tool is for modifying an already-resting order rather than submitting a new one, and it explains the important cross-spread behavior. It does not explicitly name alternatives like cancel-then-resubmit or state when-not-to-use it, but the usage context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_bookReset bookADestructiveIdempotent
Wipe state and start clean. Removes all resting orders, clears the trade tape, forgets every client_order_id, and restarts the sequence counter.
Pass symbol to clear one instrument only; omit it to clear everything and reseed the engine. Use this between independent experiments so results are comparable.
Side effects: DESTRUCTIVE and irreversible. Everything not captured in a snapshot is gone, including the tape you may still want to read. Take a snapshot first if you might need the old state.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| symbol | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by detailing exactly what is destroyed, emphasizing irreversibility, warning about the trade tape being lost, and advising to take a snapshot first. This aligns with destructiveHint=true and adds meaningful operational context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core action, and uses a clear side-effects warning section. Every sentence contributes either to understanding the operation, its parameters, or its risks.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive reset tool, the description covers the main behavior, scoping, side effects, and safety guidance well. The only notable gap is the undocumented seed parameter, which prevents full completeness for an agent trying to invoke the tool with all possible options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description gives clear semantics for the symbol parameter: passing it clears one instrument, omitting it clears everything and reseeds the engine. However, the seed parameter is never explained, despite the input schema providing only its name and default value, leaving a significant gap in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific, action-oriented language ('Wipe state and start clean') and enumerates exactly what gets reset: resting orders, trade tape, client_order_ids, and the sequence counter. This clearly distinguishes it from sibling tools like get_book or take_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: 'between independent experiments so results are comparable.' It also explains the symbol-scoped vs. global-reset distinction. It does not name alternative tools or explicitly say when not to use it, but the usage context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_snapshotRestore snapshotADestructiveIdempotent
Roll the entire engine back to a previous snapshot. Book, resting orders, trade tape and sequence counter all return to exactly what they were.
Call with no snapshot_id to list what is available.
Side effects: DESTRUCTIVE to current state -- everything that happened after the snapshot is discarded, including trades. Snapshot ids are never reused, so the snapshot you restore from stays valid and you can restore it repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
| snapshot_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructive and idempotent, but the description adds important behavioral detail: everything after the snapshot is discarded including trades, and snapshot ids are never reused so restores remain valid and repeatable. This goes well beyond the annotations and provides genuine transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: action and scope first, then the listing mode, then side effects. Every sentence contributes meaningful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description adequately covers input modes, destructive consequences, and repeatability, which is especially important given the absence of an output schema. A minor gap is that it does not describe the return format for the listing mode, but this is not critical for selecting or invoking the tool safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides snapshot_id as an optional string with no description, and schema coverage is 0%. The description compensates by explaining that omitting the parameter lists available snapshots, while providing it restores to that snapshot. This clarifies the parameter's semantics, though it does not detail error behavior or id format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('roll the entire engine back to a previous snapshot') and specifies the exact scope (book, resting orders, trade tape, sequence counter). This distinguishes it from narrower sibling tools like reset_book or load_scenario.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to call with no snapshot_id to list available snapshots, which is a concrete usage guideline. However, it does not explicitly discuss when to prefer this over alternatives like load_scenario or reset_book, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_orderSubmit orderADestructiveIdempotent
Submit a new order to the matching engine. It matches immediately against resting liquidity on the opposite side, and whatever does not trade either rests on the book or is cancelled back, depending on tif.
IDEMPOTENCY -- read this before retrying anything. You choose client_order_id. The server remembers the result of the first call carrying each id. If you call again with the same id, you get that ORIGINAL result back, marked duplicate, and no second order is created and nothing extra is matched. So: if a submit's outcome was lost or you are unsure whether it went through, RETRY WITH THE SAME id -- that is safe and it tells you what happened. Use a NEW id only for a genuinely different order. Reusing an id for a different order will NOT submit it; you will silently get the old result.
Rejections come back as data, not errors: a reason code, the offending field, a suggestion in words, and suggested_value -- a value you can drop straight into that field and retry, with no parsing. Read it and retry with a FRESH client_order_id (the rejected one is spent).
Side effects: MUTATES the book. May create trades, may consume other orders' resting quantity, advances the sequence number, and appends to the trade tape. Not reversible except via restore_snapshot or reset_book.
| Name | Required | Description | Default |
|---|---|---|---|
| qty | Yes | Order quantity in units. Must be a positive multiple of the symbol's lot size. | |
| tif | No | GTC: rest until filled or cancelled. IOC: trade what is available right now, cancel the remainder. FOK: trade the full quantity immediately or do nothing at all. POST_ONLY: never take liquidity; rejected if the price would cross the spread. | GTC |
| side | Yes | BUY lifts offers / rests on the bid. SELL hits bids / rests on the ask. | |
| price | No | Limit price, e.g. 100.01. Required for LIMIT, forbidden for MARKET. Must be a multiple of the symbol's tick size. Accepts a number or a decimal string, so a price read from get_book (which returns prices as strings) can be passed straight back here unchanged -- a string is in fact safer, since it cannot pick up binary floating-point error. | |
| symbol | Yes | Instrument symbol, e.g. 'AAPL'. Call get_stats to list what is tradable. | |
| order_type | Yes | LIMIT requires `price` and will rest on the book if it does not fully trade. MARKET must omit `price`, takes whatever the book offers, and never rests. | |
| client_order_id | Yes | An id YOU choose for this order, unique per order. If a call with this id was already processed, the server returns the ORIGINAL result and creates nothing new. Reuse the same id when retrying a call whose outcome you did not see; use a fresh id for a genuinely new order. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly discloses idempotency, rejection-as-data behavior, and side effects: it 'MUTATES the book,' may create trades, consume resting quantity, advance the sequence number, and append to the trade tape, with reversal only via restore_snapshot/reset_book. This goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but organized into labeled sections: purpose, idempotency, rejections, and side effects. Every section earns its place; no filler or repetition of the schema is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does a good job covering idempotency, rejections, and side effects. However, it does not describe the successful response shape beyond saying the original result is returned; knowing whether it includes order id, status, or fills would help. Still, the essential invocation guidance is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds valuable semantics beyond the schema: retry with the same client_order_id, use a fresh id after rejection, and pass price as a string to avoid floating-point error. This is meaningful added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Submit a new order to the matching engine.' It then explains the matching behavior and side effects, making it clear this is the creation tool distinct from get_order, cancel_order, and modify_order.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly frames this as the tool for submitting new orders and gives strong retry guidance, but it never explicitly says when to use a sibling like modify_order or cancel_order instead. Usage context is clear; alternative routing is only implied by the word 'new'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_snapshotTake snapshotA
Freeze the complete engine state -- book, orders, tape, sequence counter, client_order_id ledger -- and get back a snapshot id.
This is how you run a controlled experiment: snapshot, try something, restore, try the alternative, and compare against an identical starting point rather than a reconstructed one.
Side effects: does not change the book. Adds to the server's snapshot store, which lives in memory and dies with the process.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses concrete side effects beyond whats the annotations convey: it does not change the book, it adds to the server's snapshot store, and that store is in memory and dies with the process. This gives the agent an accurate picture of persistence, mutation, and lifecycle without contradicting the readOnly/destructive/idempotent hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: core function first, then the controlled-experiment workflow, then side effects. Every sentence carries useful information and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a snapshot operation with a single optional parameter and no output schema, the description is nearly complete: it enumerates the captured state, says the return is a snapshot id, explains side effects, and embeds the tool in a restore/comparison workflow. The only substantive omission is the meaning of the label parameter, which is minor because the parameter is optional.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description never mentions the optional label parameter, and the schema provides only the title 'Label' with a null default and no description. With schema description coverage at 0%, the description was responsible for explaining this parameter but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence uses a specific verb, 'Freeze', with a clearly identified resource: the complete engine state. It enumerates exactly what is captured (book, orders, tape, sequence counter, client_order_id ledger) and states the return value, a snapshot id, which cleanly separates it from restore_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second paragraph explicitly explains when to use this tool: snapshot, try something, restore, try the alternative, compare against an identical starting point. This gives strong contextual guidance and implies restore_snapshot as the counterpart, but it does not name the alternative tool explicitly or state when snapshot should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
cancel_order - First observed
get_book - First observed
get_order - First observed
get_stats - First observed
get_trades - First observed
load_scenario - First observed
modify_order - First observed
reset_book - First observed
restore_snapshot - First observed
submit_order - First observed
take_snapshot
TDQS
Scored across 11 tools
Each tool maps to one distinct action or query: order lifecycle, market data, and state management are cleanly separated, and even the reset/snapshot tools are differentiated by their intended use (wipe vs. reproducible starting point vs. rollback). No two tools appear to do the same thing.
All tools use a consistent snake_case verb_noun pattern: get_* for read-only queries, submit/cancel/modify_order for order mutations, and reset/load/take/restore for state management. This makes the set predictable and easy for an agent to select from.
11 tools is a well-scoped count for a limit-order-book control server: no bloat, and every tool serves a distinct operational need covering order entry, cancellation, modification, data reads, and state management.
The set covers the full order lifecycle (submit/get/cancel/modify), comprehensive market-data reads (book, trades, stats), and thorough state management (reset, scenario loading, snapshot/restore). There are no obvious dead ends: actions can be queried, reversed, or reproduced.
Maintenance
Related MCP Connectors
Trade across 22+ exchanges and brokers from any MCP-capable AI agent, no install required.
MCP server for OpenMM — exposes market data, account, trading, and strategy tools to AI agents
Live prices, perps, prediction markets and a paper trading desk over one MCP.
MCP server exposing the Backtest360 engine API as tools for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to interact with Interactive Brokers through 48 tools for market data, orders, account management, and more, via the MCP protocol.MIT
- FlicenseNot gradedqualityDmaintenanceEnables LLMs to trade on MetaTrader 5 via REST API or MCP tools, supporting market/pending orders, position management, and account info retrieval.-
- FlicenseNot gradedqualityCmaintenanceEnables agents to access live crypto market data and execute risk-controlled paper trades through standardized MCP tool calls.-
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to automate simulated stock trading on the Tonghuashun trading system, including checking account balances and positions, viewing orders and deals, placing market orders, and canceling pending orders via the MCP protocol.1MIT