FPL Strategy MCP
This server is a rules-aware Fantasy Premier League decision engine that recommends, analyzes, and backtests FPL strategy through MCP tools.
Generate legal hold, transfer, and chip recommendations with reasons, risk, short/long-term outlook, and price economics.
Plan lineups: choose formation, starting XI, bench order, captain, vice-captain, and projected Bench Boost impact.
Search the official player pool by name, team, position, price, or availability.
Inspect per-player forecast signals: expected points, price-change risk, minutes/role security, news, ownership, and uncertainty.
Rank and score legal one-transfer moves under custom weight overrides.
Return strategy details and catalogs: champion policy, benchmarks, research sources, limitations, and tunable fields.
Backtest strategies against historical gameweek files or point-in-time scenarios, including research candidates like patient_chips and external_hgb.
Provides Fantasy Premier League decision-making tools, using official FPL data to evaluate squads, transfers, chips, and player signals for hold, buy, and sell recommendations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@FPL Strategy MCPShould I use my wildcard this week? Here are my squad and signals."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
FPL Strategy MCP
FPL Strategy MCP is a local, rules-aware Fantasy Premier League decision engine. Give it your current 15-player squad, the players you can buy, prices/selling values, free transfers, chips, and current signals. It returns legal hold/transfer/chip options with the reason, risk, short-term outlook, long-term outlook, and price economics.
It is an action policy, not a promise that one player will score the most points. The shipped champion is deliberately conservative: it uses the validated free-transfer anchor unless a learned action clears the temporal guardrail. See the research paper draft for the evidence and limitations.
Install
macOS or Linux:
curl -fsSL https://raw.githubusercontent.com/bsovs/fpl-strategy-mcp/main/install.sh | sh -s -- --clients allWindows PowerShell:
$env:FPL_STRATEGY_CLIENTS="all"; irm https://raw.githubusercontent.com/bsovs/fpl-strategy-mcp/main/install.ps1 | iexThe installers download the latest release binary, register it with the selected
clients, and run a fast health check. Use --clients claude,
--clients claude-code, --clients codex, or --clients none to narrow the
setup. Existing Claude JSON and Codex TOML are backed up before they are
changed. Each GitHub release also publishes SHA-256 checksums. To pin a
version, set FPL_STRATEGY_VERSION=0.1.7 before running the installer.
Related MCP server: Fantasy Premier League MCP Server
Connect a client
The default command is a stdio MCP server:
fpl-strategy-mcpThe server writes one readiness line to stderr so MCP protocol stdout stays clean. To inspect the installation later:
fpl-strategy-mcp status
fpl-strategy-mcp status --json
fpl-strategy-mcp status --deepThe default status check is fast and only verifies installed assets and client
configuration. --deep additionally loads the bundled model.
To register an already-installed binary:
fpl-strategy-mcp setup --clients allIn Claude Desktop, add the installed command to claude_desktop_config.json under mcpServers. Use the absolute path to the binary:
{
"mcpServers": {
"fpl-strategy": {
"command": "/absolute/path/to/fpl-strategy-mcp"
}
}
}On macOS the file is ~/Library/Application Support/Claude/claude_desktop_config.json; on Windows it is %APPDATA%\Claude\claude_desktop_config.json. Fully restart Claude after editing it. The same stdio command works with Claude Code and Codex; the installer can register both automatically.
For ChatGPT or Claude web, start the optional remote transport:
FPL_MCP_BEARER_TOKEN="choose-a-long-random-token" \
fpl-strategy-mcp --transport streamable-http --host 127.0.0.1 --port 8000Expose http://127.0.0.1:8000/mcp through an HTTPS tunnel or authenticated reverse proxy, then add that HTTPS MCP URL as a custom connector/remote MCP server. Do not expose an unauthenticated listener. OpenAI’s API can call remote MCP servers through the Responses API; Claude web also expects a reachable remote connector. The default stdio mode remains the safer local option.
For HTTP mode, http://127.0.0.1:8000/health returns a JSON readiness report
and is protected by FPL_MCP_BEARER_TOKEN when that variable is set.
MCP tools
The server exposes these tools. fpl_recommend_moves is the primary decision
tool; the others make the forecasts, weights, player universe, and evaluation
loop inspectable and tunable.
Tool | Purpose |
| Return the recommended hold, transfer, or model-backed chip action under FPL legality, prices, short/long forecasts, projected XI/bench effects, uncertainty, news, and league context. |
| Choose the legal formation, starting XI, bench order, captain, and vice-captain. Reports the projected four-player Bench Boost increment; normally only the XI scores. |
| Search the cached official pool by name, team, position, price, or availability. Useful for inspecting candidates or constructing a smaller request payload. |
| Inspect each player’s short/long expected points, future price signals, minutes/role, risk, news/social context, ownership leverage, and uncertainty before making a decision. |
| Rank legal one-transfer moves under explicit |
| Return the shipped champion, benchmark summaries, research sources, and limitations. |
| Return available strategies, action kinds, default weights/gates, signal components, and the fields that can be tuned. |
| Compare candidate weights on supplied point-in-time scenarios, or replay the legal simulator over Vaastav-format historical GW files. |
For a normal recommendation, provide gameweek and the exact 15-player
current_squad. You can provide buyable_players and trained signals, or
set auto_official_signals: true and let the server load the current official
bootstrap pool and transparent fallback signals. Per-request weight changes go
under weight_overrides (also accepted as weights); they do not modify the
bundled model or the shipped champion.
For loss-aware decisions, include each owned player’s current price,
selling_price, and original purchase_price when available. The latter is
optional; without it, the recommendation can still score the move but cannot
explain or learn the cost of realizing a loss.
Input
Call fpl_recommend_moves with gameweek and current_squad, optionally
adding buyable_players, config, weight_overrides, and either point-in-time
signals or auto_official_signals: true. If buyable_players or signals
are omitted, the same official fallback is used automatically. Add
league_context when rank/leader information should influence risk. Call fpl_strategy_info or
fpl_strategy_catalog to inspect the served strategy and tuning contract.
fpl_backtest_strategy accepts either:
scenariosorscenarios_path: point-in-time states with realizedaction_outcomessuch asholdandtransfer:out_id>in_id; orhistory_rootandseason: a full legal replay using Vaastav-formatseason/gws/gw*.csvfiles. Use separate development and held-out seasons and starting-squad modes when tuning.
The action learner explicitly carries short/long fixture-window deltas, recent form, value, role security, price-change risk, and any unrealized loss on the player being sold. That lets it distinguish a tactical three-gameweek punt from a season-long core hold and learn when a declining player is worth selling at a loss because the forward upgrade is stronger.
See examples/ for client configuration and a protocol smoke test. The official FPL bootstrap endpoint is used only when requested: https://fantasy.premierleague.com/api/bootstrap-static/.
Development
python -m venv .venv
. .venv/bin/activate # Windows: .venv\Scripts\Activate.ps1
pip install -e ".[dev,remote]"
python -m unittest discover -s tests -qThe model and bootstrap snapshot are bundled under assets/. Historical training and evaluation artifacts are documented in the paper rather than required to run the server.
Train the local historical forecast layer
The research trainer keeps a complete final season untouched. The current protocol uses 2016/17–2023/24 for development, 2024/25 for model selection, and 2025/26 as the final test season. Download the raw Vaastav archive into a local data directory, then run:
python - <<'PY'
from fpl_lab.history import download_seasons
download_seasons(
["2016-17", "2017-18", "2018-19", "2019-20", "2020-21", "2021-22",
"2022-23", "2023-24", "2024-25", "2025-26"],
"data/vaastav",
)
PY
PYTHONPATH=src python scripts/train_player_models.py \
--history-root data/vaastav \
--validation-season 2024-25 \
--evaluation-season 2025-26 \
--output-dir runs/player-modelsThe trainer builds 200+ numeric point-in-time features plus categorical context:
lagged form and volatility, minutes/start security, ownership and transfer
momentum, price movement, team/opponent and prior-matchup form, schedule shape,
cross-season player history and breakout signals, and optional timestamped
news/social context. It writes the selected player
forecast model, a future-price model, evaluation predictions, metrics, and a
data audit under runs/player-models/. News/social columns remain zero unless
an auditable ContextStore is supplied; modern articles are never backfilled
into old seasons. These forecast artifacts are research inputs and are not
promoted to the shipped action-policy champion until the legal season
simulator shows a robust strategy-level improvement.
The downloader also fetches one players_raw.csv roster snapshot per season.
This restores the missing position/team fields in the oldest Vaastav GW files;
the resulting metadata_*_imputed flags are retained in the audit because a
season-level roster snapshot is not a point-in-time transfer history.
To test decision value on the held-out season, bridge the forecast CSV into the legal simulator:
PYTHONPATH=src python scripts/backtest_forecast_strategy.py \
--history-root data/vaastav \
--forecast-csv runs/player-models/evaluation-predictions.csv \
--season 2025-26 \
--previous-season 2024-25This compares the legacy signals and expanded forecasts under the same rules-aware free-transfer policy. It is a strategy smoke test. For the walk-forward action-value layer, run:
PYTHONPATH=src python scripts/train_action_policy.py \
--history-root data/vaastav \
--validation-season 2024-25 \
--evaluation-season 2025-26 \
--output-dir runs/action-policy \
--max-states 4 \
--candidate-width 6 \
--horizon-gameweeks 3 \
--training-chip-depth 5 \
--test-chip-depth 15 \
--forecast-model neural \
--label-policy points_onlyWhen backdated archives are available, add --news-context PATH and/or
--social-context PATH to that command.
This generates legal counterfactual action labels, fits the action ensemble, and evaluates neural actions against the free-transfer anchor over points, value, template, and randomized opening squads. The forecast bridge is walk-forward: it fits point forecasts, direct 3/8-gameweek totals, a next-gameweek price-change model, and a separate expected-minutes model using only earlier seasons. The current career-feature baseline produced 1,476 development and 186 validation examples. After fixing two temporal-grain defects—calendar-window horizon labels and double-gameweek lag aggregation—the clean untouched 2025/26 replay scored 2,076.25 points on average across four opening families, with a best opening of 2,156. The free-transfer anchor averaged 2,008.5 and peaked at 2,073; the anchored cocktail averaged 2,020.75. These are improvements over the anchor in this replay, but the best result is still 257 points below the 2,413 research target. This remains a research artifact, not a promoted champion.
The latest corrected-rules replay is saved under
runs/action-policy-official-snapshot-v4-season-chips/. It uses development
seasons 2016/17–2023/24, validation on 2024/25, and a completely untouched
2025/26 test. It fixes a major simulator defect by using the season-specific
2025/26 eight-token chip inventory and half-season chip gates. The corrected
2025/26 neural policy averaged 1,969.5 points across the four opening families
(best 2,053), versus 1,983.75 for the no-hit free-transfer anchor (best 2,129).
The cocktail averaged 1,997.5 (best 2,069). The corrected run is now the
authoritative corrected-rules baseline; it is below the 2,413 target and is
not a promoted champion. The later external-style candidate is reported
separately below. The earlier v3 hit-aware result was an ablation under the
old one-copy chip inventory and must not be used as the final 2025/26 score.
The public fpl-luck-or-skill challenger
reports 2,431 points from a patient, no-hit, use-it-or-lose-it TC/BB policy.
The local simulator now contains that policy as patient_chips, but the
neural/context replication scored 1,928.25 on average on the untouched
2025/26 openings (best 1,955). The external number is therefore a useful
benchmark and hypothesis, not a locally verified result; see the data-quality
audit for the exact reproducibility limitation and artifact path.
External-style forecast and forecast-optimized opening
The repo now includes a Mac-compatible external_hgb forecast family. It
reproduces the public challenger's minutes-plus-conditional-points design with
53 leakage-safe features: player form, minutes/start security, xG/xA, price,
ownership, transfer momentum, true fixture team, opponent/team form, and
previous-season production. It uses histogram gradient boosting locally, so it
does not require the external LightGBM/OpenMP runtime.
Run the strict walk-forward candidate with:
PYTHONPATH=src python scripts/train_action_policy.py \
--history-root /path/to/fpl-history \
--validation-season 2024-25 \
--evaluation-season 2025-26 \
--forecast-model external_hgb \
--starting-modes forecast \
--label-policy points_only \
--output-dir runs/action-policy-external-hgb-forecast-v1The resulting artifact is runs/action-policy-external-hgb-forecast-v1/metrics.json.
On the untouched 2025/26 test it scored:
Policy | Points | Transfers | Hits |
Neural action policy | 2,160 | 62 | 0 |
Free-transfer anchor | 2,338 | 37 | 0 |
Cocktail | 2,442 | 53 | 0 |
Patient chips candidate | 2,486 | 37 | 0 |
The patient candidate clears the 2,413 target in this strict held-out replay.
Its 2024/25 validation score was 2,357, so the opening rule was checked on a
prior season before the final test was read. This is the strongest current
research candidate. It is now exposed through the installed MCP's
fpl_backtest_strategy tool; it is not silently made the live default because
the published score uses one forecast-optimized opening family rather than a
distribution of random starts.
Run the same candidate through the MCP season simulator with:
{
"history_root": "/path/to/data/vaastav",
"season": "2025-26",
"previous_season": "2024-25",
"forecast_model": "external_hgb",
"candidates": [
{
"name": "published_research_candidate",
"policy": "patient_chips",
"initial_squad_modes": ["forecast"]
}
]
}The binary rebuilds the forecast walk-forward from all seasons before the test season, so this path is part of the MCP rather than a results-only artifact. The exact historical elite-manager alternative archive remains incomplete.
The context files are optional. Each event must carry a publication timestamp;
archived events also carry the snapshot observed_at timestamp. The live
official API exposes current cumulative minutes/starts, current
chance_of_playing_this_round/chance_of_playing_next_round, news,
scout_risks, price projections, and set-piece order fields. Its player
history exposes realized minutes and starts, not historical probability
snapshots; use scripts/fetch_fplcache_context.py to reconstruct those
point-in-time beliefs. The feature builder cuts context off at the simulated
gameweek deadline (90 minutes before the first fixture), not at kickoff.
Supported structured event types
include injury, availability, suspension, rotation,
lineup_predicted, lineup_confirmed, lineup_benched, set_piece, and
transfer. Store the analyzed sentiment, source, reliability, player/team
entity, and expiry alongside the text so the backtest can audit what was known.
See docs/context-data-contract.md.
The validated official archive replay sampled 1,722 snapshots through 1 August 2026, produced 9,366 interval/news events, improved validation action RMSE from 9.016 to 8.670, but reduced the held-out 2025/26 neural policy to 2,007.5 mean points. It is therefore an inspectable context ablation, not part of the shipped champion until the action layer gates and calibrates these signals.
For a model-family ablation, add --forecast-model ridge. The expanded ridge
run reached 2,189 points in an earlier replay, but that result used the
pre-fix temporal grain and is not comparable to the clean result above. The
direct-horizon neural and price-aware variants are retained as inspectable
research outputs; they are not evidence of a winning strategy by themselves.
The optional --starting-modes ... forecast mode adds a legal
forecast-optimized opening squad. Under the older neural/ridge bridge it
reached 2,162 points on 2025/26 and was not promoted. The newer
external_hgb replay above is a separate, validation-approved forecast
candidate and should not be conflated with that older ablation.
The hit-aware replay is saved under
runs/action-policy-official-snapshot-v3-hit-aware/. It generated 165 paid-hit
and 453 multi-transfer counterfactual actions. On untouched 2025/26, its
neural policy averaged 1,986 points (best opening 2,114), versus 1,983.75
(best 2,129) for the no-hit anchor. This is a useful coverage fix and a
negative strategy result under the pre-chip-fix simulator. The corrected-rule
follow-up is v4 above.
Observed elite-manager benchmark
The repo also includes a provenance-tracked aggregate benchmark from a public
archive of 24,041 complete 2025/26 manager seasons under
data/elite_managers/. It includes transfer gain, transfer count, hits,
captain agreement, chip usage, and transfer timing by manager rank band. To
inspect the observed behavior profile locally:
PYTHONPATH=src python scripts/analyze_elite_managers.py \
--autopsy data/elite_managers/autopsy_all.csv \
--output runs/elite-manager-benchmark/summary.jsonTo inspect the pre-deadline conditions around observed transfers, including recent points/minutes, price, ownership, transfer momentum, and a separate forward-outcome audit:
PYTHONPATH=src python scripts/analyze_observed_manager_decisions.py \
--history-root /path/to/fpl-history \
--output runs/elite-manager-benchmark/decision-audit.jsonThe top-100 band has a median 325-point net transfer gain, 62 transfers, 4
hits, 0 unused chips, 71.1% captain agreement, and 17 hours' median timing;
the top-10k band has a median 302-point transfer gain, 63 transfers, 4 hits,
0 unused chips, and 79.0% captain agreement. These are descriptive
benchmarks, not causal labels. The shipped loader deliberately excludes final
rank and final points from behavior features. The aggregate table cannot yet
identify exact weekly alternatives. The detailed 2025/26 weekly archive is now present under
data/elite_managers/season_winners_2025-26/ and is held out from training.
Pre-2025/26 weekly manager archives are still needed for leakage-safe direct
imitation or inverse-decision modeling. The local 2025/26 decision audit found
that top-100 managers bought players with a higher prior three-gameweek point
rate (13.15 versus 12.61 for players sold), slightly lower prior three-
gameweek minutes (216.9 versus 223.7), lower prior ownership, and stronger
positive transfer momentum. The following-three-gameweek points are retained
only as a quarantined outcome audit, never as model features.
To fit descriptive, leakage-audited behavior heads for transfer/hold, bundles, paid hits, chips, and incoming-versus-outgoing player signals:
PYTHONPATH=src python scripts/reverse_engineer_observed_actions.py \
--history-root /path/to/fpl-history \
--output runs/elite-manager-benchmark/action-heads.jsonSee observed action heads for the measured
behavior metrics and the data needed before these heads can train on older
seasons. The patient_chips simulator challenger is also documented there;
it is not the promoted policy.
Data coverage and missing signals
The archive is not missing the basic FPL history: it contains 247,896 raw player-fixture rows across ten seasons (2016/17 through 2025/26), including gameweek points, minutes, starts, form, ownership, transfers, prices, team scores, opponents, and fixture timing. The model turns this into 286 point-in-time features and keeps the 2025/26 season completely out of fitting and model selection.
The important gaps are contextual rather than raw player rows. The public
snapshot archive now supplies official news/availability and set-piece
intervals, and data/elite_managers/ supplies an aggregate observed-manager
benchmark. We still do not have each elite manager's exact point-in-time
alternative set wired into action labels, nor a complete timestamped
expected-minutes history, press-conference/predicted-lineup/social stream, or
richer historical fixture-strength feed. The pipeline has a leakage-safe
expected-minutes model and a full official-context ablation, but neither is
currently a validated strategy improvement. The oldest gameweek
files also need season-level roster snapshots to fill team/position metadata;
those rows are flagged and are not treated as point-in-time transfer history.
News/social signals require an explicitly supplied timestamped context archive
in the historical trainer.
The temporal audit found and fixed 293 three-gameweek label mismatches caused
by skipping blank calendar gameweeks, plus inconsistent lag values in 416
double-gameweek player groups. The current run uses calendar-window labels
and one aggregated player/gameweek grain for lags and horizon/price models.
Details and the remediation plan are in docs/data-quality-audit.md.
License
MIT. This is an independent research tool and is not affiliated with the Premier League or Fantasy Premier League.
Available Tools
8 toolsfpl_backtest_strategyA
Backtest and compare tunable strategies. Use scenarios/scenarios_path for point-in-time replay with realized action_outcomes, or history_root plus season for the legal Vaastav-format season simulator. To reproduce the published 2,486-point research candidate, use forecast_model=external_hgb, policy=patient_chips, and initial_squad_modes=[forecast].
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| policy | No | ||
| season | No | ||
| objective | No | net_points | |
| scenarios | No | ||
| candidates | No | ||
| model_path | No | ||
| base_config | No | ||
| end_gameweek | No | ||
| history_root | No | Root of Vaastav-format season folders when running the full simulator. | |
| forecast_model | No | Walk-forward forecast family used by the published research candidate. Requires history_root and at least one prior season. | |
| scenarios_path | No | ||
| start_gameweek | No | ||
| cocktail_config | No | ||
| include_details | No | ||
| previous_season | No | ||
| max_transfer_depth | No | ||
| chip_transfer_depth | No | ||
| initial_squad_modes | No | ||
| transfer_beam_width | No | ||
| transfer_candidate_width | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It discloses two execution modes and the exact parameter combination needed to replicate the research candidate, which is meaningful. However, it does not describe what the backtest returns, whether it is read-only, what data prerequisites exist, or what failure modes might occur, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences earn their place: the first states the purpose and the two mode routes, the second gives a compact reproducible recipe. It is front-loaded with the core action and uses no filler or repetition, which is exemplary for a tool with 21 parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 21-parameter tool with no output schema and no annotations, the description is not complete enough. It explains two modes and a recipe, but an agent still does not know what the backtest returns, whether one of the two modes is required, which combinations of optional parameters are valid, or what the results look like. The description is a helpful starting point, not a complete invocation guide.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 10% schema description coverage, the description must compensate, and it partially does: it adds meaning to scenarios/scenarios_path, history_root, season, forecast_model, policy, and initial_squad_modes. But 21 parameters exist, and many (seed, objective, candidates, depths, beam widths, cocktail_config, etc.) get no semantic help from either the schema or the description, so compensation is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Backtest and compare tunable strategies.' It clearly identifies the tool's core function and distinguishes it from sibling tools like fpl_recommend_moves or fpl_strategy_info by virtue of the backtesting action, but it never explicitly names or contrasts those siblings, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing between two modes: use scenarios/scenarios_path for point-in-time replay versus history_root plus season for the Vaastav-format simulator. It also names a precise parameter recipe to reproduce the published 2,486-point candidate. It lacks explicit 'when not to use' guidance against sibling tools, but the intra-tool guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_forecast_signalsA
Return the point-in-time signal table for the supplied squad and buyable pool, including short/long expected points, future price signals, minutes, risk, news/social context, ownership leverage, and uncertainty. Use it to inspect the underlying inputs before asking for a decision.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| config | No | ||
| signals | No | ||
| gameweek | Yes | ||
| current_squad | Yes | ||
| bootstrap_path | No | ||
| fetch_official | No | ||
| buyable_players | No | Optional; omit to include the full official pool. | |
| auto_official_signals | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden, and it does frame the tool as a non-mutating inspection operation ('Return', 'inspect'). However, it does not disclose fetch/caching behavior implied by parameters like fetch_official and auto_official_signals, nor whether signals are computed on demand.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two purposeful sentences: the first states the return value and content, the second states the intended usage. It is front-loaded with the core operation and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 9 parameters, nested objects, no output schema, and no annotations, yet the description only covers the high-level purpose. An agent would still be unable to determine how to set config, signals, bootstrap_path, fetch_official, or auto_official_signals, or what the returned signal table's shape is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 11%, so the description must compensate, but it only clarifies current_squad and buyable_players ('supplied squad and buyable pool'). Parameters such as config, signals, bootstrap_path, fetch_official, auto_official_signals, and limit receive no semantic explanation in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Return') and a specific resource ('point-in-time signal table for the supplied squad and buyable pool'), and enumerates the signal content. The closing phrase 'before asking for a decision' differentiates it from decision-oriented siblings like fpl_recommend_moves and fpl_score_moves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it to inspect the underlying inputs before asking for a decision' provides clear situational guidance and implies it should precede decision tools rather than replace them. It does not explicitly name alternatives or state when not to use it, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_lineup_planA
Choose the legal current-week formation, starting XI, bench order, captain and vice-captain from the supplied 15-player squad. It also reports the projected bench points that Bench Boost would add; normally only the XI scores.
| Name | Required | Description | Default |
|---|---|---|---|
| signals | No | ||
| gameweek | Yes | ||
| current_squad | Yes | ||
| bootstrap_path | No | ||
| fetch_official | No | ||
| buyable_players | No | Optional; omit to use the official bootstrap universe. | |
| auto_official_signals | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses non-obvious behavior: it only chooses/plans rather than executes, and it reports projected bench points for Bench Boost while clarifying that normally only the XI scores. This meaningfully helps an agent understand the tool's output semantics, though it does not mention side effects or data-fetching behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the main action front-loaded and the Bench Boost nuance placed second. Every clause adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 7 parameters, no output schema, and no annotations, yet the description leaves most parameter behaviors unexplained. It covers the expected outputs well but omits how optional inputs like fetch_official, signals, and auto_official_signals affect the result, making it insufficient for predictable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 14%, so the description must compensate, but it only loosely implies gameweek and current_squad. It does not explain signals, bootstrap_path, fetch_official, auto_official_signals, or the relationship between buyable_players and the official universe. An agent would struggle to select correct values for most optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Choose') and resource ('legal current-week formation, starting XI, bench order, captain and vice-captain') from a 15-player squad, and even adds the return nuance about Bench Boost points. This is clearly distinct from sibling tools about moves, strategy, search, scoring, or backtesting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the tool for current-week lineup planning from a supplied squad, but it never names sibling alternatives or gives conditions for when to use this tool instead of fpl_recommend_moves or fpl_score_moves. There is no exclusion or when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_recommend_movesA
Return legal FPL transfer moves for the current gameweek using the frozen champion strategy by default. Give a full 15-player current squad. buyable_players is optional: when omitted, the server loads the full current official player pool from the cached bootstrap snapshot/API. Give point-in-time signals or enable auto_official_signals, bank/free transfers, unused chips, and optional mini-league standings context. The default champion ranks hold and legal free transfers; explicit hybrid challengers can rank Wildcard, Free Hit, Bench Boost, and Triple Captain when model-backed. Use weight_overrides to tune the transparent decision layer per request.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| config | No | ||
| signals | No | ||
| weights | No | Alias for weight_overrides. | |
| gameweek | Yes | ||
| strategy | No | champion | |
| chips_used | No | Alternative to chips_available: chips already used this season. | |
| model_path | No | ||
| current_squad | Yes | ||
| bootstrap_path | No | ||
| fetch_official | No | When auto_official_signals is true, fetch the current official bootstrap instead of using the local snapshot. | |
| league_context | No | ||
| buyable_players | No | Optional. Omit to load every player from the official bootstrap pool. | |
| chips_available | No | Unused chips at this deadline. If omitted, all four are assumed available. | |
| weight_overrides | No | Per-request DecisionConfig overrides, e.g. short_weight, long_weight, lineup_weight, price_weight, ownership_weight, risk_aversion, rank_mode. | |
| auto_official_signals | No | If true, omitted signals are built from a local bootstrap snapshot or the official API. The server also falls back automatically when signals are omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It explains that buyable_players can be omitted and the server loads the official pool, that auto_official_signals fetches signals if not provided, and that hybrid strategies are 'model-backed'. It also mentions a 'transparent decision layer' and weight_overrides. It does not explicitly state read-only nature or side effects, but the nature of a recommendation tool implies safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized but front-loaded with the main purpose. It flows logically: primary action, required inputs, optional parameters, strategy details, and tuning. Every sentence adds value without excessive verbosity, though it could be slightly more compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 16 parameters and nested objects, the description covers the most critical aspects but is not fully complete. It lacks an output schema, so it does not describe the response format beyond 'legal FPL transfer moves'. It also omits details on how errors or edge cases are handled, and some parameters remain unexplained. For a complex tool, more explicit guidance is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 44%, so the description must compensate. It adds meaning to key parameters: buyable_players (optional, loads official pool), auto_official_signals (fetch signals if omitted), chips_available (unused chips), weight_overrides (tune decision layer). However, it does not explain several parameters like limit, config, model_path, bootstrap_path, or league_context, leaving gaps for those.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Return legal FPL transfer moves for the current gameweek' with a specific resource (FPL transfers) and a default strategy. It distinguishes itself from siblings by focusing on recommendations, not scoring or lineup planning, and mentions specific strategy options like 'champion' and 'hybrid'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: to get transfer recommendations. It provides context on required inputs (full 15-player squad) and optional parameters (buyable_players, auto_official_signals, chips, etc.). However, it does not explicitly compare against sibling tools like fpl_score_moves or fpl_lineup_plan, nor state when NOT to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_score_movesA
Rank legal one-transfer moves under explicit transparent weights. This is the tuning and explanation tool for short/long points, price economics, ownership leverage, availability, news risk, uncertainty, hit cost, and risk aversion; use fpl_recommend_moves for the final action including chips and the learned policy.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| config | No | ||
| signals | No | ||
| gameweek | Yes | ||
| current_squad | Yes | ||
| bootstrap_path | No | ||
| fetch_official | No | ||
| buyable_players | No | Optional; omit to load the full official pool. | |
| weight_overrides | No | ||
| auto_official_signals | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavioral disclosure. It indicates this is a read-only ranking tool (no side effects mentioned) and clarifies it uses transparent weights. However, it does not disclose the output format, pagination behavior, or what happens on invalid inputs (e.g., missing required params). While not misleading, it leaves notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero fluff. The primary purpose is stated in the first sentence, and the second sentence adds the key contrast with fpl_recommend_moves plus a list of relevant factors. It is appropriately sized and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 10 parameters, nested objects, and no output schema, yet the description covers only purpose and a high-level comparison. It does not explain required parameters (gameweek, current_squad), how to use config, signals, weight_overrides, limit, fetch_official, or auto_official_signals, nor what the ranking output looks like. The description is too sparse for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 10% (only buyable_players has a description). The description mentions several evaluation factors (short/long points, price economics, ownership leverage, etc.) that likely correspond to config or weight_overrides, but it does not explain how to structure these parameters or what values are expected. It provides some high-level semantic context but fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Rank legal one-transfer moves under explicit transparent weights.' It uses a specific verb (rank) and a distinct resource (one-transfer moves), and explicitly differentiates from sibling fpl_recommend_moves by saying it is for tuning/explanation rather than final action. This leaves no ambiguity about its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool versus alternatives: 'use fpl_recommend_moves for the final action including chips and the learned policy.' It also lists the factors it evaluates (short/long points, price economics, etc.), giving clear context for when this tool is appropriate. No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_search_playersA
Search the full cached official FPL player pool by name, position, team, price, and availability. Use this to construct or inspect buyable candidates; it returns current price, official expected points, form, ownership, status, and transfer signals.
| Name | Required | Description | Default |
|---|---|---|---|
| team | No | Team ID or team name fragment. | |
| limit | No | ||
| query | No | Name, player ID, or team text. | |
| position | No | ||
| max_price | No | ||
| min_price | No | ||
| available_only | No | ||
| bootstrap_path | No | ||
| fetch_official | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that results include current price, expected points, form, ownership, status, and transfer signals, and that the pool is cached. However, it does not explain the behavioral impact of non-obvious options such as fetch_official or bootstrap_path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the first identifies what the tool searches and by which criteria, and the second states the intended use and returned data. Every sentence contributes useful guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a search tool with no required parameters, but there are meaningful gaps: no output schema, no annotation coverage, and undocumented semantics for bootstrap_path and fetch_official. The description tells the agent what it will get back but not all the operational details needed to invoke it in more advanced modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 22%, so the description must compensate. It does map several filter dimensions to parameters (name, position, team, price, availability), but it does not clarify the meaning of fetch_official or bootstrap_path, and several parameters still lack any semantic guidance beyond defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Search') and a specific resource ('full cached official FPL player pool'), and lists the main search dimensions. The description clearly separates this from its sibling tools, which are about recommendations, strategies, lineups, forecasts, and backtesting rather than searching players.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: 'Use this to construct or inspect buyable candidates.' It does not explicitly name alternatives or say when not to use it, but the stated use case is actionable and distinct enough from the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_strategy_catalogA
Return all available strategy profiles, action kinds, default decision/cocktail configs, signal components, and the fields that can be tuned per request or in a backtest.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden. It says 'Return', which indicates a read operation, but it does not explicitly state that it has no side effects or require any authentication. Given that it is a catalog tool with no parameters, the ambiguity is low, but the description does not go beyond the bare action to disclose any edge cases or data semantics. It is adequate but not richly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence that front-loads the action ('Return all available') and enumerates the catalog contents. It is concise and avoids fluff, though the list of items makes the sentence a bit long. Overall it is well-structured for the task.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and no output schema, the description covers all it needs: it enumerates the types of information returned (strategy profiles, action kinds, configs, signal components, tunable fields). Nothing an agent would need to call it successfully is missing, though it could optionally mention whether the output is a schema or an array, but that is likely self-evident. It is complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty (0 parameters). Per the rubric, a baseline of 4 applies for 0 params. The description does not need to explain parameters, and it correctly focuses on the output content, which is the meaningful part. No additional parameter context is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and a clear resource ('all available strategy profiles, action kinds, default decision/cocktail configs, signal components, and the fields that can be tuned'). It explicitly says 'all available', which distinguishes it from sibling tools like fpl_strategy_info (likely a single strategy lookup) and action-oriented tools like fpl_recommend_moves. The purpose is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's role as a catalog/overview apparent, but it does not explicitly state when to use it versus alternatives, nor does it provide any 'when not to use' guidance. The usage is implied by the word 'catalog' and the sibling names (e.g., fpl_recommend_moves for actions), but no explicit routing is offered. This matches the 'implied usage' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fpl_strategy_infoA
Return the selected strategy, benchmark status, research sources, and known limitations.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It indicates a read-only action by using 'Return' and lists 'known limitations' as an output, but it does not disclose specifics such as data sources, staleness, or what the benchmark status is compared against.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that lists exactly what is returned, with no filler or repetition. It is front-loaded and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless info tool with no output schema, the description names all four return categories and gives a reasonable expectation of the payload. It could add context about how 'selected' is determined, but nothing is missing that would prevent an agent from invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete and there is nothing for the description to clarify. The baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and names the resource: the selected strategy, benchmark status, research sources, and known limitations. It is clear what the tool produces, though it does not explicitly differentiate itself from the sibling fpl_recommend_moves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus fpl_recommend_moves, and no conditions or exclusions are mentioned. The intended use is only implied by the tool name and the noun phrase 'selected strategy,' but the description never states the decision context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.8- Added
fpl_backtest_strategy - Added
fpl_forecast_signals - Added
fpl_lineup_plan - Changed
fpl_recommend_moves6 fields changed- changed
Input schema / properties / auto_official_signals / defaultPrevious value: -falseNew value: +true - changed
Input schema / properties / auto_official_signals / descriptionPrevious value: -"If true, signals may be omitted and are built from a local bootstrap snapshot or the official API."New value: +"If true, omitted signals are built from a local bootstrap snapshot or the official API. The server also falls back automatically when signals are omitted." - added
Input schema / properties / buyable_players / descriptionAdded value: +"Optional. Omit to load every player from the official bootstrap pool." - added
Input schema / properties / weight_overridesAdded value: +{ + "description": "Per-request DecisionConfig overrides, e.g. short_weight, long_weight, lineup_weight, price_weight, ownership_weight, risk_aversion, rank_mode.", + "type": "object" +} - added
Input schema / properties / weightsAdded value: +{ + "description": "Alias for weight_overrides.", + "type": "object" +} - changed
Input schema / requiredPrevious value: -[ - "gameweek", - "current_squad", - "buyable_players", - "config" -]New value: +[ + "gameweek", + "current_squad" +]
- Added
fpl_score_moves - Added
fpl_search_players - Added
fpl_strategy_catalog
2 tool updates
v0.1.4- First observed
fpl_recommend_moves - First observed
fpl_strategy_info
TDQS
Scored across 8 tools
Most tools have clearly distinct responsibilities: search, forecast, lineup, backtest, catalog, and info are well-separated. The main ambiguity is between fpl_recommend_moves and fpl_score_moves, though the descriptions do clarify that one is for final policy-driven actions and the other for transparent weight tuning of one-transfer moves.
The consistent fpl_ prefix and snake_case help readability, and most names follow a verb_noun pattern like fpl_recommend_moves, fpl_search_players, fpl_score_moves. However, fpl_lineup_plan, fpl_strategy_catalog, and fpl_strategy_info deviate from that pattern, mixing noun-led and verb-led naming.
Eight tools is a well-scoped count for an FPL strategy server, covering search, signals, decision-making, lineup planning, and backtesting without bloat. Each tool appears to serve a distinct part of the analysis-to-action workflow.
The server covers the main FPL workflow: player discovery, signal inspection, move scoring, final recommendations, lineup selection, and strategy backtesting. Minor gaps exist, such as no explicit tool for creating or persisting custom strategies or for managing chip state directly, but these are largely addressable through existing tools.
Maintenance
Related MCP Connectors
FPL data, points predictions, and transfer, captain and chip advice for any team.
141Fantasy Premier League expected points, team ratings, captaincy and mini-league win chances.
Fantasy Premier League tools: rate a team, captains, player points, fixtures, injuries. No login.
1Fantasy analysis for your ESPN, Yahoo, and Sleeper leagues. Reads your leagues, never changes them.
Related MCP Servers
- AlicenseAqualityDmaintenanceAI-powered Fantasy Premier League assistant — scored captain picks, transfer suggestions, differentials, fixture outlook, price predictions, live points, and a full manager hub that auto-detects your squad, bank balance, and free transfers.1312 PyPI11MIT
- AlicenseAqualityCmaintenanceEnables AI assistants to analyze Fantasy Premier League data, providing tools for player search, fixture analysis, manager comparisons, and strategy prompts for transfer planning and lineup selection.1924 PyPI1MIT
- FlicenseNot gradedqualityCmaintenanceEnables Fantasy Premier League squad management with custom tools for player search, fixture outlook, tier classification, hit math, and chip timing, all using public FPL data without requiring login credentials.1-
- FlicenseNot gradedqualityCmaintenanceProvides tools and resources to interact with the Fantasy Premier League API, enabling player analysis, fixture insights, and team management through natural language.-