abacus
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@abacusCompute implied vol for a 1-month call, spot 100, strike 102, price 2.5."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
abacus
An MCP server that exposes option, portfolio, execution and backtest analytics as tools.
pip install abacus-mcp
abacus tools # the whole surface, in one screen
abacus stdio # what an MCP client launchesNot on the package index yet. The lines above are what they will be. Until the first release lands, install from source — the server depends on five sibling libraries, so they go first and then nothing has to be resolved by name:
pip install \ "git+https://github.com/DaniyalMlk/moneyness.git" \ "git+https://github.com/DaniyalMlk/shortfall.git" \ "git+https://github.com/DaniyalMlk/tenor.git" \ "git+https://github.com/DaniyalMlk/slippage.git" \ "git+https://github.com/DaniyalMlk/holdout.git" pip install "git+https://github.com/DaniyalMlk/abacus.git"
abacus toolsandabacus stdiothen work exactly as above.
Once released, nothing needs installing at all if you have
uv — uvx abacus-mcp stdio fetches it, runs it
and caches it. Registering it with a client is
one entry naming that command.
Documentation — the conventions worth reading before the first call, every tool with its arguments, and the numbers this project claims with the test that checks each one.
Language models are unreliable at arithmetic, so the useful thing a tool
boundary can do is move the numbers somewhere trustworthy. abacus puts them in
libraries whose numerical cores are validated against closed forms,
high-precision references and published results, and exposes those cores over the
Model Context Protocol. Nothing here wraps a third-party pricing API: the numbers
are computed by these libraries and tested where they live.
It targets MCP revision 2026-07-28 and has five runtime dependencies, all
pure Python: moneyness for the
option mathematics, shortfall for
the portfolio risk estimators, tenor
for curves and bonds, slippage for
execution cost and holdout for
backtest validation. The last two install as slippage-tca and
holdout-backtest, because the short names on the index belong to unrelated
projects; their import names are unchanged.
Running it
pip install abacus-mcp
abacus stdio # what an MCP client launches
abacus http # Streamable HTTP on 127.0.0.1:8000/mcp
abacus tools # print the tool surface and exit
abacus conform # run the conformance suite against this serverThe distribution is abacus-mcp; the package you import is abacus. The short
name was already taken on the index by an unrelated project, and renaming the
package to match would have changed every import for the sake of a registry
collision.
Related MCP server: genpark-black-scholes-merton-greeks-engine-skill
Registering it with a client
Nothing has to be installed first if you have uv:
uvx fetches the server, runs it, and caches it for next time.
{
"mcpServers": {
"abacus": {
"command": "uvx",
"args": ["abacus-mcp", "stdio"]
}
}
}Where that file lives depends on the client:
Client | Configuration |
Claude Desktop |
|
Claude Code |
|
Cursor |
|
VS Code |
|
Zed |
|
With the package installed into an environment rather than run through uvx,
point the client at the installed script instead. An absolute path is worth the
noise: a client launched from a desktop session rarely has the same PATH as
your shell, and a bare abacus that works in a terminal and not in the client is
the most common way this goes wrong.
{
"mcpServers": {
"abacus": {
"command": "/path/to/venv/bin/abacus",
"args": ["stdio"]
}
}
}To check the server is healthy before wiring a client to it, run the conformance suite against the launch command you are about to configure:
abacus conform --stdio "uvx abacus-mcp stdio"It exits non-zero on a failure, so it works as a gate rather than a report.
The tools
Tool | What it does |
| Price under generalised Black-Scholes-Merton, with forward, intrinsic and time value |
| Delta, gamma, vega, theta, rho, the higher-order Greeks, and sensitivities to strike and carry |
| Both of the above in one call |
| Both sides of a strike and the parity residual |
| The price range attainable by some non-negative volatility |
| Recover volatility from a price, with convergence evidence |
| Fit a raw SVI slice, with fit quality and a butterfly check |
| Fit a surface, with both no-arbitrage conditions checked |
| Dupire local volatilities, including where the identity has no answer |
| American price on a lattice, with convergence reporting |
| Bjerksund-Stensland 2002, labelled as an approximation |
| The early-exercise boundary as a series |
| Assemble option and underlying positions into a book, returning a handle |
| Read a book back from its handle |
| Add or remove legs, or move the market, returning a new handle |
| Total value and net sensitivities, with the per-leg breakdown |
| Reprice a book across a grid of spot and volatility shifts |
| Covariance from a returns matrix, with shrinkage, diagnostics and a reusable handle |
| Value at risk and expected shortfall by seven methods, each naming itself, including a tail fitted to the exceedances and a fitted copula |
| Euler risk contributions, concentration, effective bets |
| Weights that equalise risk contributions, with the convergence evidence |
| Deepest drawdown, time underwater, ulcer index, Calmar and Sortino |
| Whether a value-at-risk forecast worked: coverage, clustering, traffic light |
| A GARCH(1,1) fit, and the forecast series the validator scores |
| A curve from deposits, futures and par swaps, with a reusable handle |
| Discount factors, zero rates and forwards at whatever dates you ask for |
| Price, yield, duration and convexity, from a yield and from a curve |
| Key rate durations, curve shape risk, and the tradeable hedge |
| Z-spread, I-spread, and option-adjusted spread off a calibrated lattice |
| What a position earns to a horizon, split into carry and roll-down |
| What an order cost, split into delay, trading, opportunity and explicit |
| The Almgren-Chriss trajectory, its cost, and the half-life's elasticities |
| Expected cost against cost risk, one schedule per risk aversion |
| A Sharpe ratio corrected for how many things were tried, on the effective count |
| How long a record must be before a ratio that size means anything |
| How many independent bets a correlated set of trials really is |
| How often the in-sample winner lands in the bottom half out of sample |
| Whether the best candidate beats the benchmark by more than the search would |
Conventions, which are also stated in the server's instructions and in every
schema description: volatilities and rates are decimal fractions, so 20% is
0.2 — and so is a 20% return; time is a year fraction, so thirty days is about
0.082; log-moneyness is measured on the forward as log(strike / forward); a
value at risk is a positive loss over one period of whatever frequency the
returns have, so periodsPerYear is required rather than assumed.
Portfolio risk
# estimate once, then ask several questions of the estimate
abacus tools | grep -A2 estimate_return_momentsFour decisions in this part of the surface are worth knowing before using it.
Every result names its method. A one-day 99% value at risk of 2.2% under a normal assumption and 2.0% from the sample are the same quantity estimated two ways, and nothing about either number says which. So the method is on the result, along with the observation count and whatever diagnostics that method has: the degrees of freedom, the moments a correction used, the effective sample behind a historical tail.
A returns handle carries the second moments, not the returns. A handle has to
fit in a message a model carries through its context, and is capped at 8192
characters for that reason. Measured: a year of daily returns on four assets
encodes to about 6000 characters and five years on ten assets to about 69,000, so
a handle carrying the matrix would refuse almost every portfolio worth asking
about. A mean vector and a covariance matrix are n + n² numbers whatever the
history length.
What that buys is real — send a matrix once, then reweight, decompose and rebalance across as many calls as you like. What it costs is that the historical estimators and every drawdown statistic read the path, which a second-moment summary has discarded. Those need the matrix again, and say so rather than answering from what they have. The Cornish-Fisher correction is in the same position for a subtler reason: it needs the skewness and excess kurtosis of the portfolio, which depend on the weights.
Some questions have no answer for some portfolios, and get none. Cornish-Fisher is refused outright when the estimated moments put its corrected quantile outside the region where it increases with the probability — outside it the mapping is not a quantile function and nothing read off it is a quantile of anything. Concentration and effective bets come back null for a portfolio with a negative risk contribution, because they read the shares as a distribution and a negative share is not one; the contributions themselves are unaffected and still returned.
Weights are not normalised silently. A sum of 0.98 is either a two percent cash position or a typo, and scaling it quietly turns the second into a plausible answer. The sum is reported on every result and a sum far from one is refused with the total named.
Above about 99.5%, ask for the fitted tail. portfolio_tail_risk's sixth
method, extreme-value, fits a generalised Pareto to the exceedances over a high
threshold and extrapolates past the largest observation. It is the only method on
this surface that can answer a far-tail question at all: the historical ones cannot
report a loss larger than the worst observed, and the parametric ones report a
shape fitted to the body, where almost all of the likelihood lives.
The fit comes back with the figure rather than behind it — the shape with its
standard error, the exceedance count, the lowest confidence the fit says anything
about, whether the answer is beyond every loss in the sample, and the mean excess
curve, which is linear above a generalised Pareto threshold with slope
shape / (1 - shape) and is therefore both how to choose the threshold and a
second reading of the shape.
Three refusals rather than plausible numbers. A confidence below the threshold's own exceedance probability is outside the fit, and the error names the lowest legal one, because reading the empirical quantile there instead would be a different estimator answering under this one's name. A fitted shape at or above one has no finite mean, so the expected shortfall comes back null while the value at risk still stands. And a threshold leaving fewer than ten exceedances is refused with both counts named.
When assets fall together, ask for the copula. The other six methods on that
tool either tie the joint distribution to a covariance matrix or read the joint
tail straight off the sample. The first forces the probability of two assets being
beyond their own q quantile together, divided by q, to zero as q falls — at
any correlation below one under a normal, and to one number for every pair under a
multivariate t. The second cannot report a joint event worse than the worst one
observed. So the question a tail-risk tool is most often asked had no method here
that could answer it.
copula fits the dependence to the ranks and each marginal separately, then
simulates, and reports the Gaussian-copula figure beside its own from the same
normal draws — so the gap between them is the assumption rather than an argument.
The fitted degrees of freedom and the likelihood ratio against the Gaussian
special case come back with the figure, and the note reads as weak evidence when
the ratio is weak, because a number quoted without that reads as strong.
Two things about it are worth knowing before reading the output, and both are counterintuitive enough that the tool says them.
The gap is widest for a book that looks diversified. A book already correlated at 0.9 gets nearly the same answer from either copula, because both move it together and the portfolio behaves as a single asset whose own marginal tail no copula can change. A book correlated at 0.08 is where the covariance matrix is most reassuring and most wrong.
And the two measures disagree about the direction. At 95% confidence the fitted copula's value at risk has been measured 4.0% below the Gaussian copula's while its expected shortfall is 6.0% above. Tail dependence moves probability mass from the near tail to the far tail and the total is one, so a quantile close to the body has less beyond it, while the mean of what is beyond is larger. Compare the two methods on expected shortfall; reading the 95% value at risk alone says the assumption made the book safer.
On a Student-t factor with four degrees of freedom — tail index 0.25 exactly — the
fit reads 0.2275 above the worst 5% of 2,000 observations with a standard error of
0.12. Biased low, and known to be: a Student-t approaches its limiting tail slowly
and a threshold that far inside the body is still being told about the body. Above
the worst 20% the same fit reads 0.1176 with a standard error of 0.056, which is
the bias-variance trade the tailFraction argument exposes rather than decides.
Curves and bonds
Three things to know before using this part of the surface.
No convention has a default, and that is the point. A bond priced on the wrong day count basis is wrong by a few basis points — exactly the size of the spread anyone is trying to measure — so it is not approximately right, it answers a different question and looks entirely normal doing it. Every basis, frequency and rolling rule is an argument. The one exception is the rolling rule, which has a market convention; it is read out of the underlying library's own default rather than chosen here, and every result reports the rule it used, because it changes the cashflow dates and therefore the price.
The curve handle carries the curve, unlike the returns handle. The difference is measured rather than stylistic. A returns matrix grows with the length of the history — five years of daily data on ten assets encodes to about 69,000 characters, against a handle limit of 8192 — so that handle carries second moments and gives up the path. A bootstrapped curve is its pillars: two numbers per instrument, 264 characters at five pillars and 764 at sixty. So this handle carries the pillars and the quotes behind them, and every curve tool works from it exactly as from a fresh bootstrap, including instrument risk, which is a question about the quotes.
A bootstrap says whether it worked. A curve that fails to reprice the instruments it was built from is not slightly wrong; it means nothing, and its pillar values look ordinary either way. The result carries the sweep count, the per-instrument solve residuals, and the repricing check. The same applies to the short-rate lattice: it reports whether it reprices the curve it was calibrated to, because a tree that does not is not a model of that curve and every spread read off it is wrong.
A zero rate at the curve's reference date comes back as null rather than zero. The discount factor there is one whatever the rate is, so no rate is implied, and a zero would read as a rate rather than as the absence of one.
Execution
Two questions sit either side of a trade, and both are about the gap between the price a model priced and the price that happened.
abacus tools | grep -A2 decompose_implementation_shortfallAfter the fact. decompose_implementation_shortfall takes an order, its
fills and three prices, and splits what it cost four ways. The total is the least
useful number in the result: a trade that cost 87 basis points tells you to feel
bad, while the same trade split into 20 of delay, 51 of trading, 15 of
opportunity and 1 of commission tells you which thing to change. Delay is the
price moving before the order reached the market, and is fixed by shortening that
gap. Trading is the order's own footprint, and is fixed by spreading it out.
Opportunity is the part that never got done. Commission is a contract.
Fill timestamps are not required. slippage.Order carries them for its volume
work, but the decomposition reads only quantities, prices and commissions and
gives an identical breakdown at one-minute and six-hour fill spacings — so asking
for them would be asking for data to be invented.
Whether delay is charged on the quantity ordered or the quantity executed is a house convention, and it matters more than it looks. On the worked order above the order basis gives a delay of 200 and an opportunity of 150; the executed basis gives 180 and 170. The total is 869 either way. It is a reattribution that survives a check on the headline, which is exactly how two desks end up agreeing on the cost and disagreeing on the cause, so every result names the basis it used.
Before the fact. optimal_execution_schedule lays the order out against
impact and price risk. At zero risk aversion the answer is a straight line — an
equal slice each period, which is TWAP — and raising the aversion front-loads the
schedule, paying more impact to spend less time exposed. execution_cost_frontier
does the same across a range of aversions, because a single optimal schedule
answers a question the caller has already had to answer.
Two notes on the model. A fixed cost per share is paid whatever the order is
traded in, so it moves the expected cost by exactly itself and does not move a
single trade. Permanent impact is the one the textbook says drops out of the
schedule, and in continuous time it does; in discrete time it enters through
eta - gamma*tau/2 and the leftover is of the order of the period length —
measured here, it shortens the half-life by 2.5% at tau = 1, 0.25% at
tau = 0.1 and 0.025% at tau = 0.01.
A risk-neutral schedule's half-life is infinite, and it comes back as null
rather than as a number. Infinity is not JSON: a strict parser rejects the
whole message over it, so one unbounded quantity would take every number beside
it down.
Backtest validation
Every other tool here answers a question about a price. These answer a question about a claim: somebody says this strategy earns a Sharpe of 1.5, and the honest reply depends on facts about the search that found it, none of which are in the number.
abacus tools | grep -A2 deflated_sharpe_ratioThe trial count that matters is the effective one. Two hundred variations of
one moving-average rule are not two hundred independent bets, and deflating as
though they were over-penalises the result. On the forty one-factor trials in
tests/test_validation.py the eigenvalue method puts the effective count at
3.0 against a raw 40, and the deflated Sharpe at 0.8244 against 0.7196 — ten
points of probability the raw count throws away. On forty independent trials
both counts are 40.0 and both figures are 0.5591, identical to the digit. The
effective count is free where it is not needed and substantial where it is, so
it is what the deflation uses; both figures come back, and so do all three
estimates of the count, because they disagree — 2.49, 3.00 and 1.08 on the same
data.
A Sharpe ratio here is per period, and its standard error is always beside it. The number people quote is annualised and these formulas take the unannualised one, which is the quietest way to get a wrong answer out of this group. The standard error is usually the answer anyway: an annualised 1.0 over 252 observations carries an annualised standard error of 1.00, and over 30 observations of 2.90.
The bootstrap tools are seeded and say so. superior_predictive_ability
resamples, so an unseeded call would return a different p-value each time, which
would make the idempotent annotation a lie and make two runs look like a change
in the data. The seed defaults to a fixed value and comes back in the result, so
a figure can be reproduced from the result alone.
Refusals, because everything in this group will compute. A deflated Sharpe
from four observations is a number. An overfitting probability from two
strategies is a number, drawn from a two-point distribution — with k
strategies the logit takes at most k distinct values, measured at 2 for two
strategies and 3 for three — and it prints to four decimal places exactly like a
real one. Skewness and kurtosis are refused when no distribution could have
them: every distribution satisfies kurtosis >= 1 + skewness^2, so a skewness
of -1.5 needs an excess kurtosis of at least 0.25.
Nothing here nominates a benchmark for you. superior_predictive_ability
needs one and model_confidence_set does not, which is the difference that
decides which to reach for. Given a field of candidates with no incumbent among
them, picking the sample-best as the benchmark and testing the rest against it
chooses the benchmark with the same data the test runs on, so under the null it
is the luckiest column present and every comparison is biased towards finding
nothing. The confidence set asks instead which models cannot be told apart from
the best, and returns them all.
Its answer is usually larger than anyone expects. On thirty crossover rules over
ten years of a market with a genuine drift in it — a sweep whose best rule clears
the deflated Sharpe test — 29 of the 30 survive at the 10% level, and the
surviving set spans 9.3% of annualised mean return. The size of the set is the
result, not a shortcoming of it. A smaller alpha gives a larger set, because
this is a confidence region rather than a hypothesis test.
The volatility tool now estimates the tail, and hands back the multiplier.
conditional_volatility used to return a volatility and tell the caller to
multiply it by "the quantile of whatever distribution you are assuming". For the
distribution worth assuming that is a footgun: the standardised Student-t
quantile is the raw one times sqrt((v-2)/v), and the raw one is 41% larger at
four degrees of freedom — so a caller who reaches for it widens every forecast
and undershoots the breach count, which looks conservative rather than wrong.
quantileMultiplier is now in the result, backed out of the fitted risk so the
two cannot drift apart, with the mean removed so it is a quantile of the
innovation and of nothing else.
99% forecasts on regime-switching data | breaches per 2000 | nominal |
constant volatility | ~51 | 20 |
GARCH, normal innovations | 28.2 | 20 |
GARCH, estimated tail | 22.45 | 20 |
A horizon figure is simulated, not scaled. The tool takes a horizon and has
always reported the aggregate volatility over it. Turning that into a quantile is
the step the library refuses one period out, because the sum of the horizon's
innovations is not a member of the family they were drawn from — so paths now
runs the recursion forward instead and returns horizonRisk with a Monte Carlo
error beside it. The substitution it replaces is wrong by 2.5, 11.0, 8.2 and 0.6
standard errors on four samples: real on average, not decisive on any one series,
because how far a horizon quantile departs from a scaled one depends on where the
fit sits relative to its long-run level.
Both square-root-of-time ratios come back, on the volatility and on the quantile, because they are different quantities and can sit on opposite sides of one. A stochastic variance path makes the accumulated return leptokurtic and pushes the quantile above the scaled figure; aggregating fat innovations pulls the total towards normality and pushes it below. A caller handed only the first would read a horizon as conservative when it is not.
About 70% of the excess a Gaussian fit leaves behind, and the note says what is
still there rather than claiming the problem is solved: a series whose volatility
jumps between regimes does not have identically distributed standardised
residuals, so one tail index for the whole sample is closer than the normal's and
still an approximation. The innovation is tested for rather than assumed —
fatTail carries the likelihood ratio — because a series with thin innovations
should not have its quantile widened for no reason anybody asked for. On such a
series the two agree to within half a breach in twenty.
The skill
skill/SKILL.md is the overview the individual schemas cannot
be: which tool to reach for, the conventions the whole surface shares, and three
worked transcripts — a book and its Greeks, a portfolio's risk, and whether a
backtest is evidence of anything.
The transcripts are not prose about calls. They are fenced transcript blocks
holding the arguments verbatim, with $name placeholders for the handles
threaded between steps, and tests/test_skill.py parses them out and runs every
one against a live server. An argument renamed in a schema, a required field
added, a handle that stopped round-tripping: each shows up as a failing test
rather than as an example somebody copies and cannot make work.
The same file checks that every tool the skill names is registered and every registered tool is named — a rename breaks the first, a new phase breaks the second — and enforces a budget on the tool listing, which is 115,162 characters today against a ceiling of 140,000, with no single tool over 12,000. The listing is loaded before any work happens and grows with every group added, so the ceiling exists to make crossing it a decision rather than a drift.
The documentation site
docs/build.py renders
the site from the live registry — every
tool entry is the object a client receives from tools/list, so a renamed
argument changes the page on the next build and there is no version of it that
is confidently wrong. The prose around it is written: the conventions, the
validated-numbers commentary, and a page on the libraries underneath and when to
import them directly instead.
python docs/build.py # writes site/, standard library onlyThe validated numbers page is the part worth knowing about. Every specific
figure this project prints was measured rather than guessed, and
docs/claims.py names, for each one, the test that measures
it. tests/test_docs.py then asserts that each of those tests exists, that the
figure appears in its source, and that all of them pass. A number that changes in
the code and not on the page fails a build rather than going on reading as
authoritative.
Checking a server against the specification
The conformance suite drives a server through the wire format and reports on requirements of the specification — not of this code — so it is meaningful pointed somewhere else:
abacus conform # this server, in process
abacus conform --stdio "abacus stdio" # a launched command
abacus conform --http http://127.0.0.1:8000/mcp # a running endpoint
abacus conform --http ... --json # machine-readableIt exits non-zero on a failure, so it works as a gate rather than only as a report, and CI runs it over both transports and against the built wheel.
Checks cover discovery, resultType and server identity on every result, cache
hints on list results, version negotiation, the error-code allocation rules, the
distinction between a protocol error and a tool execution error, and — on
Streamable HTTP — header validation, the Base64 sentinel, and the verbs and
headers this revision retired. A requirement that cannot be tested against a
given server reports skip with the reason; counting an untested requirement as
satisfied would make the whole suite worthless.
The suite is itself tested by injecting one defect at a time into this server —
the handshake still implemented, resultType missing, a retired error code, an
execution failure escalated to a JSON-RPC error — and requiring that the check
written for that requirement is the one that turns red. Passing a healthy server
proves very little; failing a broken one on the right check is the evidence.
Design
The protocol core is written against the specification
Revision 2026-07-28 removed the initialize handshake and made MCP stateless.
There is no session: every request carries its own protocol version, client
identity and capabilities in _meta, and a server may not infer any of them
from the connection a request arrived on, because two requests on one stdio pipe
may belong to unrelated conversations.
Building the core directly against that model keeps it explicit, keeps the
runtime dependency list at a single entry, and makes the specification's own
requirements — the error-code allocation policy, the $ref rules, the header
validation — things the code enforces rather than things it assumes someone else
did.
Failures are sorted by who can act on them
The specification draws a line between two kinds of failure, and the line is about audience.
A protocol error is a JSON-RPC error. It says the request was not actionable at all: an unknown tool, a malformed body, an unsupported protocol version. A model cannot usually recover from one, because the fault is in the plumbing.
A tool execution error comes back as an ordinary result with isError set.
It says the call arrived intact and the server declined to answer it, and it is
addressed to the model: which field, what was wrong with it, what would be
accepted instead.
So calling price_european_option with a misspelled field gets back a result
rather than an error, naming both the mistake and the properties that are
accepted:
price_european_option was called with invalid arguments:
/time: is required but was not given
/vol: is required but was not given
/volatility: is not a recognised property; accepted properties are:
carry, rate, spot, strike, time, type, volEvery violation is reported at once, so a caller with three problems learns about three problems instead of fixing them one round trip at a time.
Nonsense is refused before it is priced
A schema can say vol is a positive number. It cannot say that vol: 20 is
twenty percent written the wrong way — and pricing it as two thousand percent
returns a number that is arithmetically correct and completely wrong, with
nothing downstream any the wiser. Inputs beyond plausible bounds are refused
with the conversion spelled out:
vol=20 is out of range; it looks like a percentage. This field is a
decimal fraction, so 20% is 0.2. Values above 5 are rejected as implausible.The same applies to a rate given as 5 and an expiry given as a count of days.
Quantities that genuinely do not exist are not invented either. At expiry the
price is still well defined as a limit and is returned, but d1 and d2 are
omitted rather than filled with an infinity that would read downstream as a real
number, and the Greeks are refused outright because the payoff is kinked at the
strike and the derivative does not exist there. A local volatility at a point
where the Dupire identity has no answer comes back marked inadmissible, naming
the two quantities that failed, rather than as a nan that is not portable JSON
and reads as a number to anything parsing it loosely.
A surface is not returned as a grid of numbers
The obvious representation of a volatility surface is a dense matrix, and it is the one a language model reads worst: an unlabelled nested array gives no clue which axis is maturity and which is strike, and a transposed reading still looks plausible. A surface therefore comes back three ways at once, in descending order of how much trust each deserves:
The five SVI parameters per slice — exact, sufficient to rebuild the slice, and small enough to survive a conversation without truncation.
Named scalars — the at-the-money volatility, the wing slopes against Lee's bound, the fit residuals.
A labelled grid, only when asked for, where every row carries its maturity and every cell its strike as an explicit key, so a row cannot be read as a column.
Both no-arbitrage conditions are reported as flags with the offending location attached, rather than left to be inferred. They are scanned between the quoted maturities as well as at them, because quoted slices are usually fitted to be admissible and interpolation is where the condition quietly stops holding.
A numerical answer carries how it was produced
An American price is the output of a method, not a formula, so its error is invisible in the number. A lattice price at 64 layers and the same price at 4096 layers are different numbers, and a tool returning a bare float invites a caller to treat a discretisation artefact as a market fact. So every American valuation carries the method, the resolution, the early-exercise premium, and on request a ladder of prices as the grid doubles, with the last change quoted as an error estimate — described as such, because nothing here proves a bound.
Two European prices are reported beside it, analytic and on-lattice, because the
premium is the lattice-internal difference: the discretisation error is common
to both legs and largely cancels, which makes it the better estimate. The cost
is that price - europeanPrice does not reproduce it, so both are given rather
than leaving a reader to find the discrepancy and distrust all three.
The boundary read-out carries a similar caveat. On a binomial lattice it alternates between two values one node apart, because consecutive layers sample interleaved node grids of opposite parity — the true boundary is monotone, and the wobble is the discretisation rather than the option. The tool says so and points at the trinomial lattice, which has no such parity.
A book is carried by its handle, not stored behind one
A position book is state, and this revision has nowhere to keep it. The obvious implementation — a dictionary on the server, keyed by a random string handed back to the caller — fails three ways that have nothing to do with taste. It reintroduces the session the revision removed, and with no session scope the book is reachable by whoever presents the key. It grows without bound, because nothing in the protocol tells a server that a caller has finished. And it breaks across processes, since a handle minted by one worker is unknown to the next.
So the handle contains the book: canonical JSON, compressed, and authenticated with a truncated HMAC-SHA256 that is checked in constant time before anything is decompressed. The server stores nothing and verifies everything, which disposes of all three problems at once.
The trade is stated rather than buried. The payload is signed, not encrypted, so the caller can read it — acceptable because it is the caller's own book, and a reason nothing else may be put in there. It cannot be altered, because an edited handle fails its MAC. It is bounded in size, so a book too large to encode is refused when it is opened rather than minting something that fails on use. And the key lives for the life of the process, which is why a handle that fails its check is reported as unrecognised rather than expired: a caller seeing that after a working call has learned something true about the server.
Books are therefore immutable. Amending one mints a new handle and leaves the old one working until it expires, and the result says so — the server holds no record of either and could not revoke the old one if it claimed to.
An aggregate says which legs it covers
A leg that has expired, or that carries no volatility, has no derivative: the payoff is kinked and there is nothing to differentiate. Contributing a zero for it would report a book as flat in precisely the case where part of it has no delta at all, and nothing downstream could tell.
Such legs are priced into the total — the price is a limit and exists — but
excluded from the aggregate Greeks, which then come back with complete: false,
the indices of the excluded legs, and a sentence saying the aggregate covers the
remaining legs only.
Scenario grids are labelled for the reason surfaces are. A spot-versus-vol grid is square often enough that a transposed reading still looks plausible, so every row states its volatility shift and every cell states the spot it was priced at.
The reference client implements this revision and nothing else
There is no initialize, no session header, no fallback to an older shape, and
no accommodation for a server that answers the way servers used to. That is the
point: MCP moved, most published guidance still describes the stateful form, and
a client that quietly tolerated it would make a non-conforming server look fine.
It is also deliberately thin on judgement. It checks only the JSON-RPC framing
that must hold for a response to be matched to its request, and returns
everything else untouched — because a client that raised on a missing
resultType could not report on one.
Header validation is a security control, not a formality
Streamable HTTP mirrors selected body fields into headers so intermediaries can
route without parsing the body. That creates two sources of truth for one fact.
If a load balancer routes on Mcp-Name: read_file while the server executes a
params.name of delete_everything, every control in front of the server was
applied to a request that is not the one that ran.
So a disagreement between header and body is refused with HeaderMismatch and
400 — and the comparison is made on decoded values, since a server that
skipped the check for Base64-encoded headers would have handed an attacker the
way around it.
The transport also declines the older shape of itself: 405 for the GET and
DELETE that used to open a stream and end a session, Mcp-Session-Id ignored
rather than echoed, and 404 with a JSON-RPC body for an unknown method, which
is what lets a client tell a modern server that lacks the method from a legacy
endpoint that was never there.
Development
pip install -e ".[dev]"
pytest # 629 tests
mypy --strict
ruff check .Continuous integration runs the suite on Python 3.10 through 3.13, type-checks and lints, runs the conformance suite over both transports, and installs the wheel and the sdist into separate clean environments to confirm each one's entry point answers a real request and passes conformance — a distribution that imports but cannot serve is not a working server, and the two artefacts are built by different code paths.
The declared dependencies are ordinary bounded version ranges. One of them used to
be a git+https direct reference, which resolves perfectly well locally
and is refused outright when a distribution carrying it is uploaded to an index —
so the server was unpublishable while every check was green. tests/test_metadata.py
now fails if such a requirement reappears. Until the library has a release on the
index, CI builds it from its repository into a local wheelhouse and lets pip
resolve the declared range against that; the resolution path is the one an index
install takes, and only the source of the file differs.
Releasing
The version lives in pyproject.toml, is mirrored by abacus.__version__, and a
test asserts they agree. Pushing v<version> builds both artefacts, checks the
metadata the way the index will, installs each into a clean environment and makes
it pass conformance, and then publishes — using the index's trusted-publishing
flow, so there is no upload token in this repository or in its secrets.
Prices and Greeks are checked against the library exactly and, independently, against finite differences of the prices the pricing tool itself returns. Those come from different formulae, so agreement is evidence the wiring is right rather than merely self-consistent.
Licence
MIT.
This server cannot be deployed
Maintenance
Related MCP Connectors
Portfolio risk analytics — VaR, Monte Carlo, optimization, options Greeks, stress testing.
Option analytics over the SYNTH sample or a permitted chain: Greeks, positioning, payoffs, plots.
FlashAlpha MCP — wraps the FlashAlpha options-exposure analytics API
HPSILab Quant finance MCP for US stocks, ETFs, options, Monte Carlo, backtesting, and risk analysis.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides real-time options analytics, pricing with Greeks, Monte Carlo simulations, volatility analysis, strategy backtesting, and risk metrics using actual market data from Yahoo Finance and Polygon.io.1-
- AlicenseNot gradedqualityBmaintenanceEnables users to compute European option prices and sensitivity metrics, run Monte Carlo GBM simulations, calculate historical/parametric VaR/CVaR, evaluate bond duration and convexity, and interpolate Nelson-Siegel yield curves through MCP.7MIT
- AlicenseNot gradedqualityBmaintenanceEnables analytical pricing of European options and calculation of first- and second-order Greeks including Delta, Gamma, Vega, Theta, and Rho through MCP. It also supports related quantitative finance analytics such as Monte Carlo simulations, VaR/CVaR, bond duration, and yield curve interpolation.7MIT
- AlicenseNot gradedqualityBmaintenanceEnables quantitative finance and risk analysis through MCP, including geometric Brownian motion Monte Carlo simulations, Black-Scholes Greeks, VaR/CVaR, bond duration/convexity, and Nelson-Siegel yield curve interpolation.7MIT