Skip to main content
Glama

backtest_strategy

Perform a full strategy backtest over a historical period (Walk-forward analysis). Use this for testing general rules or long-term performance.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fYesForecast horizon
qYesQuery length
stepNoSimulation step in bars
topKNoNumber of nearest neighbours
endTsNoEnd timestamp for simulation
feePctNoPer-side fee percentage
symbolYesTicker symbol
maxBarsNoMaximum recent bars to load for the backtest
minProbNoDirectional probability threshold
startTsNoStart timestamp for simulation
intervalYesTimeframe
token_idNoOptional Manus access token. Paid tools use tokenized service access, not a monthly subscription: when token_id is omitted the server returns payment_required with a Solana Pay invoice, and after payment you retry with the same token while the server uses Manus token/resolve to recover pending access.
directionNoAllowed direction: long, short, or both
minAvgSimNoMinimum average similarity required to trade
onlySignalsNoReturn only non-neutral decisions
slippagePctNoPer-side slippage percentage
includeStatsNo
embeddingModeNoPattern embedding mode for ANN retrieval

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description should disclose behavioral traits like side effects, data mutation, permissions, or computational costs. It only describes the action and purpose, missing any transparency about what happens during the backtest or whether it produces side effects. The lack of disclosure is significant for a complex tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences: the first states the core action and method, the second provides usage guidance. It is front-loaded with the essential purpose and contains no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This tool has 18 parameters and no output schema or annotations, yet the description gives only a high-level purpose. It does not explain walk-forward analysis, parameter interactions, or return values, leaving the agent under-informed for a complex operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (94%), so the input schema already documents parameter meanings thoroughly. The description adds no additional parameter insights, but the baseline of 3 applies because the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Perform') and names a concrete resource ('full strategy backtest') with a defined methodology ('Walk-forward analysis'). It clearly distinguishes this from siblings by focusing on historical backtesting for general rules and long-term performance, rather than single decisions or pattern searches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this for testing general rules or long-term performance' provides clear context for when to invoke this tool. However, it does not explicitly state when not to use it or mention alternatives, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.2/5.0
Disambiguation2/5

Several tools inhabit overlapping territory: find_market_analogs, pattern_search, search_by_sketch, and get_candle_market_snapshot all relate to historical pattern matching, while get_trading_decision, get_trader_decision_v2, and get_live_polymarket_trade_decision all produce trade-oriented decisions. The descriptions add context, but an agent could still easily pick the wrong tool for a given request.

Naming Consistency3/5

The tools are consistently snake_case and mostly readable, but the naming conventions are mixed: many tools use get_<noun>, while others start with verbs like backtest, detect, find, forecast. Minor irregularities such as pattern_search and the v2 suffix in get_trader_decision_v2 also reduce predictability.

Tool Count4/5

Fifteen tools is within a reasonable size, and the server covers a broad domain: pattern search, regime detection, backtesting, track records, private datasets, live Polymarket decisions, and documentation. The count is not excessive, but some tools are functionally redundant enough that the set could be tightened.

Completeness4/5

The tool surface covers the main evidence workflow well: discovering patterns, analyzing analogs, backtesting strategies, checking track records, and producing trading decisions. Minor gaps remain around private dataset management and there is no separate low-level raw candle query tool, but most core user journeys are supported.

Resources