Skip to main content
Glama

Backtest Sinyal MM Detection (Empiris)

binance_backtest_signal
Read-only

Validasi empiris sinyal binance_detect_mm_activity: ambil snapshot sinyal aktif (skor >=0.6) yang tersimpan di D1 (diisi Cron tiap 5 menit untuk watchlist tetap) dalam rentang waktu tertentu, hitung forward return (harga N jam setelah sinyal vs saat sinyal) per baris, lalu agregat win rate/avg return/max drawdown. Forward return dihitung ON-DEMAND dari klines historis saat tool dipanggil (bukan data pre-computed). Dibatasi maksimal 50 baris paling baru dalam range per panggilan (tiap baris butuh 2 kline lookup). HANYA symbol watchlist tetap yang punya histori sinyal tersimpan.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
symbolYesSymbol dari watchlist tetap: BTCUSDT, ETHUSDT, SOLUSDT, BNBUSDT, XRPUSDT, DOGEUSDT, ADAUSDT, AVAXUSDT, LINKUSDT, LTCUSDT, TRXUSDT, SUIUSDT, HYPEUSDT, ZECUSDT, NEARUSDT, UNIUSDT, BCHUSDT, TAOUSDT, WLDUSDT, AAVEUSDT, XMRUSDT, ONDOUSDT, FILUSDT, XLMUSDT, DOTUSDT, ENAUSDT, 1000PEPEUSDT, PUMPUSDT, ASTERUSDT, WLFIUSDT, PAXGUSDT, TRUMPUSDT, XAUTUSDT, ETCUSDT, ATOMUSDT, ICPUSDT, APTUSDT, ARBUSDT, OPUSDT, INJUSDT, SEIUSDT, RUNEUSDT, TIAUSDT, STXUSDT, IMXUSDT, GALAUSDT, SANDUSDT, MANAUSDT, POLUSDT, ALGOUSDT
endTimeYesWaktu akhir, ISO 8601 (contoh "2026-08-12T00:00:00Z")
startTimeYesWaktu mulai, ISO 8601 (contoh "2026-08-01T00:00:00Z")
signalTypeNoFilter jenis sinyal, atau 'all' untuk semuaall
forwardWindowNoJendela forward return yang dihitung setelah tiap sinyal trigger4h

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnly/openWorld annotations by disclosing that forward returns are computed on-demand from historical klines, that each row triggers two kline lookups, that results are limited to the 50 newest rows, and that data is populated by a Cron job every 5 minutes. This gives the agent a strong understanding of cost, freshness, and limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it front-loads the core purpose, then explains the computation model, limits, and data source. Every sentence adds value and there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only backtest tool, the description covers inputs, processing logic, output metrics, and constraints. It does not describe the exact output shape or what happens when no signal history exists, but it names the aggregate metrics and the absence of an output schema reduces the obligation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds useful context around the signal selection threshold (score >= 0.6) and the forward-return concept, but it does not add meaningful semantics beyond what the input schema already provides for symbol, startTime, endTime, signalType, and forwardWindow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: empirical validation of binance_detect_mm_activity signals, including the specific process of taking signal snapshots, computing forward returns, and aggregating win rate, avg return, and max drawdown. It uses a specific verb ('Validasi empiris') and resource, and is easily distinguishable from the sibling detector tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when this tool is appropriate: it backtests previously stored signals from binance_detect_mm_activity rather than detecting or fetching live signals. It also states important constraints like watchlist-only symbols and the 50-row limit, though it does not explicitly say 'use X instead when...'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.7/5.0
Disambiguation4/5

Most tools map to a distinct Binance metric or analytic concept, and descriptions explicitly contrast near-neighbors (spot vs futures, snapshot vs delta, 'BEDA dari...' notes). A few pairs could still be confused—`binance_get_basis` vs `binance_get_basis_history` and `binance_get_agg_trades` vs `binance_get_recent_trades`—but their purpose differences are explained well enough for careful agents.

Naming Consistency4/5

The dominant pattern is `binance_<verb>_<object>` in snake_case, with consistent complementary pairs like `get_*` and `get_*_history`. There are minor style breaks: `orderbook` vs `order_book`, the `whalescope_*` prefix, and `whalescope_full_pipeline` which lacks a verb, but the overall structure is readable and predictable.

Tool Count1/5

56 tools cross the explicit '50+ tools' extreme threshold. Although the Binance Futures domain is broad, many tools are single-endpoint or single-metric wrappers—multiple klines variants, order book variants, and ticker variants—that could be consolidated into parameterized composite tools. The surface is far too large for most agents to navigate efficiently.

Completeness4/5

The public market-data and analytics surface is remarkably complete: klines, funding, open interest, long/short ratios, top-trader data, liquidations, basis, order book behavior, regime detection, and full pipeline scoring are all covered. The main gaps are documented limitations such as unavailable liquidation-by-price data and non-public account/execution tooling, but agents can work around them without dead ends.

Resources