Skip to main content
Glama
pineforge-4pass

PineForge-Codegen

Backtest a Pine strategy

backtest_pine
Destructive

Run a real, deterministic backtest of a PineScript v6 strategy against an OHLCV CSV locally, returning trades, P&L, and drawdown.

Instructions

Run a real, deterministic backtest of a PineScript v6 strategy — prefer this over estimating its trades or P&L by reasoning, which is unreliable for Pine (series semantics, intrabar fills, and strategy.* order logic do not reproduce from approximation). Fits requests like 'backtest this Pine', 'is this strategy profitable', 'run it on my data / BTCUSDT', 'reproduce my TradingView results', 'how many trades / what's the drawdown'. Transpile a PineScript v6 strategy and run it against an OHLCV CSV via the pineforge-release Docker image on the user's local machine. Fully local — transpile + backtest run in-container; nothing leaves the box, no API key. Optional inputs overrides input.() named values from the Pine source (keys = the second arg of input.(...) calls, e.g. 'Fast Length'). Optional overrides overrides strategy(...) header fields (initial_capital, commission_value, default_qty_value, pyramiding, slippage, default_qty_type, commission_type, process_orders_on_close). Returns the parsed JSON report (summary, trades, applied_inputs, applied_overrides, applied_runtime, elapsed_seconds). The instrument matters: pass symbol (a Binance symbol) or syminfo (your own qty_step, mintick, ...); a CSV from fetch_binance_ohlcv carries its instrument and needs neither. The lot size is TradingView's own reading for the symbol, from a measured table shipped with this server (Binance's LOT_SIZE.stepSize only for a symbol TradingView does not list; TradingView's usual 0.001 for a listing newer than the table); the tick size and currencies come from Binance's public exchangeInfo (or from the sidecar next to a CSV fetched by fetch_binance_ohlcv). syminfo goes over all of it. What was applied is in applied_runtime.syminfo; when the lot grid is unknown, or its lot size is not a TradingView reading, the result has a warnings entry. If the report is too large to return inline it is written to report_path and a compact summary (with that path) is returned instead. Use backtest_pine_grid for sweeps.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
imageNoDocker image override. Defaults to ghcr.io/pineforge-4pass/pineforge-engine:latest.
inputsNoMap of Pine input.*() names → value (string/number/bool). Sent as PINEFORGE_INPUTS env var to the runtime.
marketNo'spot' (default; a warning says when it is assumed) or 'usdt_perp'; the Binance market `symbol` is looked up in. A CSV fetched by fetch_binance_ohlcv for the same symbol keeps the market it was fetched for.
sourceYesPineScript v6 source.
symbolNoBinance symbol the CSV holds, e.g. 'BTCUSDT'. The lot size is TradingView's own reading for the symbol, from a measured table shipped with this server (Binance's LOT_SIZE.stepSize only for a symbol TradingView does not list; TradingView's usual 0.001 for a listing newer than the table); the tick size and currencies come from Binance's public exchangeInfo (or from the sidecar next to a CSV fetched by fetch_binance_ohlcv). TradingView's reading differs from Binance's step for most symbols (USDT-M BTCUSDT is 0.000001 on TradingView, 0.001 on Binance), so order quantities are floored as on TradingView; 0.001 is what TradingView reads for 90.6% of Binance spot symbols and 97.5% of USDT-M ones, and a lot size that is Binance's or that usual 0.001 comes with a warning. Without an instrument the engine can book sub-lot margin-call rows TradingView does not. A CSV written by fetch_binance_ohlcv needs neither `symbol` nor `syminfo`: the instrument is recorded next to it (<csv>.instrument.json) and used. If nothing can be resolved the run still goes ahead without a lot grid and says so in `warnings` and applied_runtime.syminfo.
runtimeNoEngine runtime args (NOT strategy() header) controlling timeframe semantics and intra-bar fill simulation. input_tf / script_tf set the chart and strategy timeframes — script_tf must be coarser than or equal to input_tf or the engine rejects the run. bar_magnifier + magnifier_samples + magnifier_dist enable sub-bar price-path sampling for tighter stop / limit fills. Each field is optional and only forwarded to the engine when set. Call list_engine_params for the full catalog.
syminfoNoThe instrument's own values, for a CSV of any other instrument; they win over what `symbol` or the CSV's sidecar gives. Applied to the engine: qty_step (the lot grid) and mincontract, mintick, pointvalue, type, currency, basecurrency. Without a qty_step (or mincontract) the lot grid stays off and the result carries a warning; without a mintick the engine's 0.01 applies. ticker, tickerid, timezone and session are not applied (engine defaults).
overridesNostrategy(...) header overrides. Each key maps to a single argument of the Pine `strategy()` call; only the keys you set are applied. Sent as PINEFORGE_OVERRIDES env var. Call list_engine_params for the full catalog with types and enum values.
report_pathNoWhere to write the full JSON report IF it is too large to return inline. Large backtests (long trade lists + equity curves) are offloaded to this file and the tool returns a compact summary + report_path instead; read the file for the complete trades/equity. Defaults to pineforge-backtest-<timestamp>.json in the working dir.
ohlcv_csv_pathYesAbsolute or cwd-relative path to OHLCV CSV with header 'timestamp,open,high,low,close,volume' (timestamp = UNIX ms UTC).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.9.36
    • addedInput schema / properties / market
      Added value: +{
      +  "description": "'spot' (default; a warning says when it is assumed) or 'usdt_perp'; the Binance market `symbol` is looked up in. A CSV fetched by fetch_binance_ohlcv for the same symbol keeps the market it was fetched for.",
      +  "enum": [
      +    "spot",
      +    "usdt_perp"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / symbol
      Added value: +{
      +  "description": "Binance symbol the CSV holds, e.g. 'BTCUSDT'. The lot size is TradingView's own reading for the symbol, from a measured table shipped with this server (Binance's LOT_SIZE.stepSize only for a symbol TradingView does not list; TradingView's usual 0.001 for a listing newer than the table); the tick size and currencies come from Binance's public exchangeInfo (or from the sidecar next to a CSV fetched by fetch_binance_ohlcv). TradingView's reading differs from Binance's step for most symbols (USDT-M BTCUSDT is 0.000001 on TradingView, 0.001 on Binance), so order quantities are floored as on TradingView; 0.001 is what TradingView reads for 90.6% of Binance spot symbols and 97.5% of USDT-M ones, and a lot size that is Binance's or that usual 0.001 comes with a warning. Without an instrument the engine can book sub-lot margin-call rows TradingView does not. A CSV written by fetch_binance_ohlcv needs neither `symbol` nor `syminfo`: the instrument is recorded next to it (<csv>.instrument.json) and used. If nothing can be resolved the run still goes ahead without a lot grid and says so in `warnings` and applied_runtime.syminfo.",
      +  "maxLength": 40,
      +  "minLength": 2,
      +  "type": "string"
      +}
    • addedInput schema / properties / syminfo
      Added value: +{
      +  "additionalProperties": false,
      +  "description": "The instrument's own values, for a CSV of any other instrument; they win over what `symbol` or the CSV's sidecar gives. Applied to the engine: qty_step (the lot grid) and mincontract, mintick, pointvalue, type, currency, basecurrency. Without a qty_step (or mincontract) the lot grid stays off and the result carries a warning; without a mintick the engine's 0.01 applies. ticker, tickerid, timezone and session are not applied (engine defaults).",
      +  "properties": {
      +    "basecurrency": {
      +      "description": "syminfo.basecurrency, e.g. 'BTC'.",
      +      "maxLength": 64,
      +      "minLength": 1,
      +      "pattern": "^[\\x20-\\x7e]+$",
      +      "type": "string"
      +    },
      +    "currency": {
      +      "description": "syminfo.currency, e.g. 'USDT'.",
      +      "maxLength": 64,
      +      "minLength": 1,
      +      "pattern": "^[\\x20-\\x7e]+$",
      +      "type": "string"
      +    },
      +    "mincontract": {
      +      "description": "syminfo.mincontract: the smallest tradable quantity step (TradingView reports the lot size here). Defaults to qty_step when only that is given.",
      +      "maximum": 1000000000000,
      +      "minimum": 1e-12,
      +      "type": "number"
      +    },
      +    "mintick": {
      +      "description": "Price tick size (syminfo.mintick). Fills round to it. Engine default 0.01.",
      +      "maximum": 1000000000000,
      +      "minimum": 1e-12,
      +      "type": "number"
      +    },
      +    "pointvalue": {
      +      "description": "Money per price point per contract (syminfo.pointvalue). Engine default 1.",
      +      "maximum": 1000000000000,
      +      "minimum": 1e-12,
      +      "type": "number"
      +    },
      +    "qty_step": {
      +      "description": "Lot size in base units: order quantities are floored to a multiple of it. This is what removes sub-lot rows. Defaults to mincontract when only that is given.",
      +      "maximum": 1000000000000,
      +      "minimum": 1e-12,
      +      "type": "number"
      +    },
      +    "type": {
      +      "description": "syminfo.type, e.g. 'crypto', 'forex', 'stock'.",
      +      "maxLength": 64,
      +      "minLength": 1,
      +      "pattern": "^[\\x20-\\x7e]+$",
      +      "type": "string"
      +    }
      +  },
      +  "type": "object"
      +}
  2. Changed2 schema fields changedv0.9.30
    • removedInput schema / properties / inputs / additionalProperties / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "number"
      -  },
      -  {
      -    "type": "boolean"
      -  }
      -]
    • addedInput schema / properties / inputs / additionalProperties / type
      Added value: +[
      +  "string",
      +  "number",
      +  "boolean"
      +]
  3. Changed1 schema field changedv0.9.0
    • addedInput schema / properties / report_path
      Added value: +{
      +  "description": "Where to write the full JSON report IF it is too large to return inline. Large backtests (long trade lists + equity curves) are offloaded to this file and the tool returns a compact summary + report_path instead; read the file for the complete trades/equity. Defaults to pineforge-backtest-<timestamp>.json in the working dir.",
      +  "type": "string"
      +}
  4. First observedv0.8.4

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/openWorld/non-idempotent, and the description adds real value on top: fully local execution, no API key, Docker image override, report offloading to report_path when too large, and a warnings mechanism when the lot grid is unknown. It does not contradict openWorldHint=true since it discloses the Binance exchangeInfo lookup. Slightly short of 5 only because it never states failure/timeout or resource expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and preferred usage, but the instrument/lot-size paragraph is very long and largely restates what the `symbol` schema property already says verbatim, so the description carries duplication. Dense and useful, yet not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with nested objects and no output schema, the description does enumerate the return payload (summary, trades, applied_inputs, applied_overrides, applied_runtime, elapsed_seconds) and explains the offload behavior and warnings. That covers most of what an agent needs, though the enumerated report fields are not typed and richer return detail would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter precedence semantics the schema does not: syminfo wins over symbol, symbol wins over the CSV sidecar, and inputs keys are the second argument of input.*() calls. That resolution logic is genuinely useful beyond the per-field schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a real, deterministic backtest of a PineScript v6 strategy') and immediately distinguishes itself from its closest sibling by naming backtest_pine_grid for sweeps. An agent can separate this from transpile_pine or the grid variant without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to prefer this over estimating P&L by reasoning, lists concrete triggering requests ('backtest this Pine', 'is this strategy profitable', 'reproduce my TradingView results'), and routes sweeps to backtest_pine_grid. Both the when-to-use and the alternative are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.