Skip to main content
Glama

flecs-mcp

An MCP server that lets MCP clients and AI agents inspect a running FLECS application through FLECS' built-in REST API. Agents can query entities, read components, explain queries, and inspect systems and performance statistics. With explicit opt-in, they can also make small, reversible changes to the world.

  • Built on FastMCP 4 with async httpx.

  • Read-only by default. Mutation tools are only registered when you enable them.

  • No generic HTTP passthrough: every tool maps to one verified FLECS REST operation.

  • Configured entirely through environment variables. Runs over stdio (default) or streamable HTTP.

Contents

Related MCP server: unity-agent-gateway

Why it was created

Coding agents such as Claude Code can read your source code, but not the world that code produces at runtime. In an ECS, bugs usually live in runtime state: which components an entity actually has, what values they hold, and which systems match it. Without runtime access, you end up pasting debugger output into the chat.

flecs-mcp gives the agent that runtime view, so it can reason over the code and the running world together:

               Claude Code
                    │
         ┌──────────┴──────────┐
    Source code           Running world
    (C/C++, build)        (flecs-mcp → FLECS REST)
         └──────────┬──────────┘
               AI reasoning

What you can do with it

  • Inspect the live world. "Show entities with Position and Velocity and their values" becomes the FLECS query Position, Velocity. "Health but no Position" becomes Health, !Position. This makes the MCP effectively an ECS debugger interface for the AI.

  • Debug systems. Given "The player isn't moving, investigate", the agent inspects the player and the systems that match it (flecs_get_entity with matches). It then compares the player with an entity that does move, and forms a hypothesis from real state rather than from source code alone.

  • Close the development loop. Given "Implement enemy movement and verify it", the agent edits and builds the code, runs the game, records enemy positions, queries again later, and fixes the code if nothing moved.

  • Check expected state. For example, check that a spawned enemy has Position, Health and AIState. This helps with emergent behavior that is awkward to unit-test.

  • Explore the architecture. See how many entities use each component (flecs_list_components), which systems exist and what they match (flecs_list_queries, flecs_run_named_query), and how FLECS plans a query (flecs_explain_query).

  • Investigate performance. Ask "Which systems take the most frame time?" (flecs_get_pipeline_stats, flecs_get_world_stats), then cross-check against match counts, the source code, or other MCP servers such as a metrics backend.

  • Experiment (requires FLECS_REST_ALLOW_MUTATIONS=true). Create entities, set component values or pause a system, then watch what happens. This makes the running world a sandbox for experiments, or an AI-driven editing console. Deleting entities and running scripts are not available, by design.

Requirements

Installation

git clone <this repository> flecs-mcp
cd flecs-mcp
uv sync

Configuration

All configuration comes from environment variables. Empty values count as unset.

Variable

Default

Description

FLECS_REST_URL

(required)

Base URL of the FLECS REST API, e.g. http://localhost:27750. A path prefix is allowed (for a reverse proxy); a query string is not.

FLECS_REST_TIMEOUT

5

Request timeout in seconds (> 0).

FLECS_REST_VERIFY_TLS

true

Verify TLS certificates for https:// URLs. Accepts true/false, 1/0, yes/no, on/off.

FLECS_REST_ALLOW_MUTATIONS

false

Register the MUTATION tools.

MCP_LOG_LEVEL

INFO

DEBUG, INFO, WARNING, ERROR or CRITICAL. Logs go to stderr.

MCP_TRANSPORT

stdio

stdio, or http for streamable HTTP.

MCP_HOST

127.0.0.1

Bind address for MCP_TRANSPORT=http.

MCP_PORT

8000

Bind port for MCP_TRANSPORT=http. The endpoint is http://MCP_HOST:MCP_PORT/mcp.

export FLECS_REST_URL=http://localhost:27750

.env.example lists every variable. The server doesn't read .env files itself, but uv can load one for you:

cp .env.example .env
uv run --env-file .env flecs-mcp

The server stops with exit code 2 and a clear message if a value is missing or invalid.

FLECS has no authentication. If you put it behind a reverse proxy that uses HTTP basic auth, include the credentials in the URL (https://user:password@host/flecs). They are sent as an Authorization header and are redacted from every log line and error message.

Running

FLECS_REST_URL=http://localhost:27750 uv run flecs-mcp

uv run python -m flecs_mcp is equivalent. With the default stdio transport, the process talks MCP over stdin/stdout, so you normally let your MCP client start it (see below). The server starts even when FLECS isn't running yet: connectivity is only checked when a tool is called.

To serve MCP over HTTP instead:

FLECS_REST_URL=http://localhost:27750 MCP_TRANSPORT=http uv run flecs-mcp
# MCP endpoint: http://127.0.0.1:8000/mcp

The HTTP endpoint has no authentication. Keep it on 127.0.0.1 unless it sits on a trusted network.

MCP client configuration

Replace /path/to/flecs-mcp with the absolute path of this repository.

Claude Code

claude mcp add --transport stdio --env FLECS_REST_URL=http://localhost:27750 \
  flecs -- uv --directory /path/to/flecs-mcp run flecs-mcp

Add --scope project to store the server in the project's .mcp.json so your team shares it. The equivalent .mcp.json entry:

{
  "mcpServers": {
    "flecs": {
      "type": "stdio",
      "command": "uv",
      "args": ["--directory", "/path/to/flecs-mcp", "run", "flecs-mcp"],
      "env": {
        "FLECS_REST_URL": "http://localhost:27750"
      }
    }
  }
}

To connect to a server started with MCP_TRANSPORT=http:

claude mcp add --transport http flecs http://127.0.0.1:8000/mcp

Other MCP clients (Claude Desktop, Cursor, VS Code, ...)

Most clients accept the same mcpServers shape (command, args, env) in their own configuration file, for example claude_desktop_config.json for Claude Desktop. Use the JSON above and add "FLECS_REST_ALLOW_MUTATIONS": "true" to env if you want the mutation tools.

Docker

The image is optional; uv remains the primary workflow.

docker build -t flecs-mcp .
{
  "mcpServers": {
    "flecs": {
      "command": "docker",
      "args": ["run", "-i", "--rm", "-e", "FLECS_REST_URL", "flecs-mcp"],
      "env": { "FLECS_REST_URL": "http://host.docker.internal:27750" }
    }
  }
}

Inside a container, localhost is the container itself. Use host.docker.internal (Docker Desktop) or the host's address to reach a FLECS application running on the host.

Available tools

Entity paths use FLECS dotted notation: Sun.Earth, flecs.core.World. That's the parent and name fields of a query result joined with .. You can also use a numeric id such as #523. Names containing a dot are escaped as FLECS prints them (main\.flecs). Component ids use full paths (planets.Mass) or pairs ((flecs.core.ChildOf, Sun)).

Every tool returns structured JSON, with an output schema where the shape is known. FLECS payloads are passed through unchanged. Tool descriptions start with [READ] or [MUTATION] and carry MCP annotations (readOnlyHint, destructiveHint, idempotentHint).

READ tools (always available)

flecs_get_world_info

Checks the connection and describes the world. Call it first.

  • Arguments: none.

  • Returns: {rest_url, build_info, world_summary, notes}.

    • build_info: FLECS version, compiler, enabled addons and build flags.

    • world_summary: entity, table, component and query counts, fps, frame count, uptime. It is null when the stats module isn't imported, and notes says why.

  • Example: "Which FLECS version is the game running, and how many entities does it have?"

flecs_get_entity

Returns one entity with its tags, relationship pairs and component values.

  • Arguments:

    • entity (required)

    • values (default true)

    • inherited: prefab (IsA) components.

    • type_info: component schemas.

    • entity_id: include the numeric id.

    • doc: flecs.doc names and descriptions.

    • matches: queries, systems and observers that match this entity.

  • Returns: FLECS entity JSON: {parent, name, tags, pairs, components, id?, type_info?, inherited?, matches?}. A component value is null when the component has no reflection data.

  • Example: {"entity": "Sun.Earth", "matches": true} answers "which systems process Earth?"

flecs_get_component

Returns the value of one component.

  • Arguments: entity, component.

  • Returns: {entity, component, value}, e.g. value = {"x": 10, "y": 20}.

  • Example: {"entity": "Sun.Earth", "component": "planets.Mass"}.

flecs_query

Runs a query in the FLECS query language. FLECS parses and evaluates it; this server only passes it through.

  • Arguments:

    • query (required)

    • limit (1–1000, default 100) and offset, for paging.

    • table: return all components of each match instead of only the query fields.

    • values, fields, entity_ids, inherited, type_info, doc.

  • Returns: {page: {offset, limit, returned, may_have_more}, results: [...], type_info?}.

    • Each result has parent, name, and fields (values, ids, sources, is_set per term), or all of the entity's components with table=true.

    • If may_have_more is true, call again with offset += limit.

    • Invalid queries return the FLECS parser error, including the error position.

  • Examples:

    • Position, Velocity: entities with both components.

    • (ChildOf, Sun): children of Sun.

    • !(flecs.core.ChildOf, *): root entities.

    • SpaceShip, $this ~= "Uss": name contains "Uss".

    • Position, ?Mass(up): Mass is optional and may come from an ancestor.

flecs_run_named_query

Evaluates an existing named query, system or observer query.

  • Arguments:

    • name (required), e.g. game.systems.Move.

    • variables, e.g. parent:Sun or x:e1,y:e2.

    • The same paging and serialization options as flecs_query.

  • Returns: same shape as flecs_query.

  • Example: "Which entities does the Move system currently process?"

flecs_explain_query

Shows how FLECS parses and plans a query, without returning results.

  • Arguments:

    • query (required)

    • profile: also measure evaluation time and match counts; evaluates the query repeatedly for about 1 ms.

  • Returns:

    • query_info: resolved terms, operators, sources and traversal.

    • field_info: type and schema per field.

    • query_plan: plain text.

    • query_profile: only when profile is true.

  • Example: debugging a query that unexpectedly matches nothing.

flecs_get_type_info

Returns the reflection schema of a component type.

  • Arguments: component, the dotted path of the component entity.

  • Returns: {component, has_reflection, schema}, e.g. schema = {"x": ["float"], "y": ["float"]}.

  • Example: check the value format before calling flecs_set_component.

flecs_list_components

Lists component, tag and pair ids, with entity counts, storage size, lifecycle hooks and traits.

  • Arguments: name_contains (case-insensitive filter), limit, offset.

  • Returns: {total, offset, limit, components: [{name, entity_count, entity_size, tables, type?, traits, sparse?, memory?}]}.

  • Example: {"name_contains": "game."}: which game components exist, and how many entities use each one?

flecs_list_queries

Lists named queries, systems and observers with their evaluation statistics. FLECS evaluates every query to produce this list.

  • Arguments: name_contains, kind (Query, System, Observer), include_plans (default false), limit, offset.

  • Returns: {total, offset, limit, queries: [{name, kind, expr, results, count, eval_count, eval_time, eval_mode, cache_kind, ...}]}.

  • Example: {"kind": "System"}: list every system and how many entities it matches.

flecs_get_world_stats

Returns world performance metrics. Requires the FLECS stats module.

  • Arguments:

    • period: 1s, 1m, 1h, 1d or 1w.

    • history: return all 60 samples (oldest first) instead of only the latest.

  • Returns: {period, history, metrics: {"performance.fps": {avg, min, max, brief}, "entities.count": {...}, ...}}. Times are in seconds.

  • Example: "What is the frame time over the last minute?" → {"period": "1m"}.

flecs_get_pipeline_stats

Returns per-system timing. Requires the FLECS stats module.

  • Arguments:

    • period

    • pipeline, e.g. flecs.pipeline.BuiltinPipeline: systems in execution order, with sync points. Omit it for all systems.

    • history

  • Returns: {period, pipeline, history, entries: [{name, disabled, time_spent, matched_entity_count?, matched_table_count?} | sync point]}.

  • Example: "Which system takes the most time per frame?"

MUTATION tools (opt-in)

These are only registered when FLECS_REST_ALLOW_MUTATIONS=true. They change the running application.

flecs_create_entity

Creates an entity, including any missing parents. If the path already exists, the existing entity is returned.

  • Arguments: entity.

  • Returns: {entity, id}.

  • Example: {"entity": "Sun.Venus"}.

flecs_set_component

Adds a component, tag or pair and, optionally, sets its value. FLECS starts from the current value, so members you leave out keep their values.

  • Arguments: entity, component, value (any JSON; omit it to only add).

  • Returns: {entity, component, status}, where status is set or added.

  • Example: {"entity": "Sun.Venus", "component": "game.Position", "value": {"x": 5}}.

flecs_remove_component

Removes a component, tag or pair. Its value is lost.

  • Arguments: entity, component.

  • Returns: {entity, component, status: "removed"}.

flecs_set_enabled

Enables or disables an entity (which adds or removes flecs.core.Disabled; a disabled system stops running), or one of its components. A component can only be toggled if it has the CanToggle trait.

  • Arguments: entity, enabled, component (optional).

  • Returns: {entity, component, status}, where status is enabled or disabled.

  • Example: {"entity": "game.systems.Move", "enabled": false} pauses the Move system.

Resources

Resources are used only for data that doesn't change while the application runs. Live world state is served by tools.

URI

Content

flecs://build-info

FLECS build information: version, compiler, addons, flags.

flecs://type-info/{component}

Reflection schema of a component, e.g. flecs://type-info/planets.Mass.

FLECS requirements

Enable the REST API in the application. It listens on port 27750 by default:

// C
ECS_IMPORT(world, FlecsStats);         // optional: statistics tools
ecs_singleton_set(world, EcsRest, {0}); // REST server on port 27750
while (ecs_progress(world, 0)) { }
// C++
world.import<flecs::stats>();  // optional: statistics tools
world.set<flecs::Rest>({});
while (world.progress()) { }
// or: world.app().enable_stats().enable_rest().run();

C# (world.Set<flecs.EcsRest>(default)) and Rust (world.set(flecs::rest::Rest::default())) work the same way. Keep in mind:

  • The main loop must be running. FLECS answers REST requests from inside ecs_progress(). An application that is paused, stopped at a breakpoint or stuck in a long frame doesn't answer, and requests time out.

  • Reflection (the meta addon, e.g. ecs_struct / flecs::meta registration) is needed to see component values and to set them. Components without reflection show null values.

  • Stats module: flecs_get_world_stats, flecs_get_pipeline_stats and the world_summary part of flecs_get_world_info require the stats addon (FlecsStats).

  • To use a different port or bind address, set EcsRest.port / EcsRest.ipaddr.

FLECS REST API assumptions

The implementation targets the FLECS v4 REST API and was verified against the FLECS source (src/addons/rest.c, src/addons/http/http.c, docs/FlecsRemoteApi.md) at commit 9e874bc. It was also exercised end to end against a live FLECS 4.1.6 application. All protocol details are isolated in client.py.

Topic

Behavior relied on

Endpoints used

GET /entity, /component, /type_info, /query, /components, /queries, /stats/world, /stats/pipeline; PUT /entity, /component, /toggle; DELETE /component.

Entity paths

URLs use / separators, while FLECS prints paths with .. The client converts dotted paths, percent-encodes each element, and escapes a literal / in a name as \/.

Parameters

Booleans are sent as the literal strings true/false. Every key and value is fully percent-encoded, with spaces as %20: FLECS splits parameters on ?&= before decoding. Released FLECS (v4.1.6 and earlier) decodes only %XX, so a form-encoded + would arrive as a literal +. Only FLECS after v4.1.6 also decodes + as a space.

Query errors

Queries are sent with try=true, so FLECS doesn't log agent mistakes to the application console. FLECS then reports parse errors as {"error": ...} with HTTP 200, and the client treats that as a failure. This check applies only to query endpoints, because component values may legitimately contain an error member.

Error bodies

FLECS doesn't JSON-escape error messages, so error bodies may be invalid JSON. The client extracts the message anyway.

Paging

limit must be ≥ 1, because FLECS treats limit=0 as unlimited. The FLECS default is 1000; this server defaults to 100.

type_info

The docs say "204 if no reflection". The implementation returns HTTP 200 with body 0. Both are handled.

Query plans

FLECS colors plans and writes the escape character as [. The color codes are stripped.

Stats

In 9e874bc, /stats/* dereferences the stats component without a null check, and crashes the application if FlecsStats isn't imported (reproduced against 4.1.6). The client first checks that flecs.stats exists and refuses otherwise.

Caching

FLECS caches identical GET responses for 0.2 s, so data may be up to one frame old.

Deliberately not exposed:

  • DELETE /entity: cascades to children and can trigger (OnDelete, Panic) aborts.

  • PUT /script and GET /call: execute FLECS script code and can write files on the host.

  • GET /commands/capture: replaces the world's command hook.

  • PUT /action: maintenance only.

  • GET /world: an unbounded dump; use the paged flecs_query instead.

  • GET /tables.

Safety

  • The FLECS REST API has no authentication. Anyone who can reach its port can read the world and, through the REST API itself, modify or delete it, whether or not this server is involved. Don't expose it on untrusted networks.

  • This server exposes only explicitly implemented operations, with typed arguments. It never forwards arbitrary paths, methods or bodies.

  • Mutation tools are off by default and are never registered unless enabled. Even then, destructive operations (entity deletion, script execution) aren't available.

  • In debug builds, FLECS asserts on some invalid operations, for example changing builtin components. Only enable mutations against development builds.

  • Credentials embedded in FLECS_REST_URL are redacted from logs and errors. Request values, such as component values, are never logged, and httpx request logging is silenced.

Architecture

MCP Client (Claude Code, Claude Desktop, ...)
    ↓  MCP over stdio or streamable HTTP
FastMCP Server              server.py, tools/, resources.py
    ↓  typed tool calls
FLECS REST Client           client.py (one reused httpx.AsyncClient)
    ↓  HTTP
FLECS REST API              EcsRest, port 27750
    ↓
FLECS ECS World

Module

Responsibility

config.py

Typed Config loaded and validated from the environment; logging setup.

errors.py

FlecsError hierarchy. It derives from FastMCP's FastMCPError, so messages reach the client verbatim and are logged without tracebacks.

client.py

FlecsRestClient: URL building, encoding, error translation, response validation, FLECS version quirks.

tools/read.py

READ tools (thin adapters that shape results for agents).

tools/mutations.py

MUTATION tools, only registered when enabled.

resources.py

flecs:// resources.

server.py

create_server() wiring, the lifespan that closes the HTTP client on shutdown, and main().

Development

uv sync                         # install runtime and dev dependencies
uv run pytest                   # tests (with coverage); no FLECS instance needed
uv run ruff check .             # lint
uv run ruff format --check .    # formatting (uv run ruff format . to fix)
uv run pyright                  # type checking (strict for src/)

Run the complete quality gate (format, lint, types, tests) with:

uv run python scripts/check.py

It runs every step and exits non-zero if any of them fails.

  • Tests replace the HTTP layer with httpx.MockTransport (see tests/conftest.py), drive the MCP tools through FastMCP's in-memory client, and start the real server over stdio once.

  • Ruff enables pycodestyle, pyflakes, isort, bugbear, pyupgrade, simplify, comprehensions, pytest-style, async and Ruff's own rules.

  • Pyright runs in standard mode, with strict mode for src/.

Troubleshooting

Symptom

Cause and fix

Unable to connect to the FLECS REST API at ...

The application isn't running, REST isn't enabled, or the host or port is wrong. Check FLECS_REST_URL (default port 27750) and open http://<host>:27750/ in a browser: it should answer "You've reached the REST API for Flecs".

Timed out ... while connecting

Wrong host, a firewall, or the host is unreachable. From Docker, use host.docker.internal instead of localhost.

did not respond within N s

The connection works but FLECS isn't serving requests: the main loop is paused, stopped at a breakpoint or stuck in a long frame. Resume it, or raise FLECS_REST_TIMEOUT for large queries.

Entity 'X' not found

Use dotted paths (Sun.Earth) and full parent paths. Find the entity with flecs_query, e.g. $this ~= "Earth".

FLECS rejected GET /query: ...

The FLECS query parser error, including the position. Check names (full paths), commas and parentheses; flecs_explain_query helps.

HTTP 404 ... endpoint not found

The FLECS build doesn't include the addon that serves this endpoint (e.g. stats).

Statistics are not available

Import the stats module (FlecsStats) in the application.

Component values are null

The component has no reflection data. Register it with the meta addon.

MCP client can't start the server

Check that uv is on the client's PATH, that --directory points to this repository, and that FLECS_REST_URL is set in the client's env. Configuration errors are printed to stderr with exit code 2. Run the same command in a terminal to see them.

TLS / certificate errors

For self-signed certificates on a trusted network, set FLECS_REST_VERIFY_TLS=false.

Too much or too little logging

Set MCP_LOG_LEVEL (DEBUG shows each REST call without its values).

License

MIT

Available Tools

11 tools
flecs_explain_queryFlecs Explain QueryA
Read-onlyIdempotent

[READ] Explain how FLECS parses and plans a query, without returning results.

Returns 'query_info' (resolved terms: operator, source, traversal flags), 'field_info' (id, type and member schema per field), 'query_plan' (the FLECS query plan as text) and optionally 'query_profile'. Use it to debug queries that match nothing or match too much. Invalid queries return the FLECS parser error.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesQuery in the FLECS query language, e.g. 'Position, Velocity'.
profileNoAlso measure evaluation time and result/entity counts ('query_profile'). Evaluates the query repeatedly for up to ~1 ms.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/non-destructive, so the safety profile is covered and the description needn't restate it. It adds genuine behavior beyond that: no results are returned, structured debug sections are produced, and invalid queries surface the FLECS parser error rather than failing opaquely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and the 'no results' constraint are front-loaded, followed by the returned sections and then the debugging use case. Four short sentences, each carrying distinct information with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return shapes, yet it still names the key sections and the optional profile. Purpose, usage trigger, result structure, and error behavior are all covered, leaving nothing an agent needs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'query' and 'profile' are already fully documented in the schema, including the ~1 ms repeated-evaluation behavior. The description only echoes the optional 'query_profile' output, adding no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Explain how FLECS parses and plans a query') and immediately scopes it ('without returning results'), which cleanly separates it from the sibling flecs_query that presumably executes queries. An agent can identify the tool's role without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it to debug queries that match nothing or match too much' gives a concrete condition for reaching for this tool. It stops short of naming the alternative (e.g. flecs_query) or stating when not to use it, so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_get_componentFlecs Get ComponentA
Read-onlyIdempotent

[READ] Get the value of a single component of an entity.

Returns {entity, component, value}; 'value' is the component serialized by FLECS reflection (usually an object of members, e.g. {"x": 10, "y": 20}). Errors when the entity does not have the component, when the id is a tag (no data) or when it cannot be resolved.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesEntity path in FLECS dotted notation, e.g. 'Sun.Earth' or 'flecs.core.World' (the 'parent' and 'name' fields of a query result joined with '.'), or a numeric entity id such as '#523'. Escape a literal '.' inside a name as '\.'.
componentYesComponent, tag or pair, written as FLECS prints it: a full path such as 'planets.Mass', or a pair such as '(flecs.core.ChildOf, Sun)'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
valueYes
entityYes
componentYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds real behavioral context: the return shape {entity, component, value}, the fact that value is reflection-serialized (usually a member object like {"x":10,"y":20}), and the three error conditions. That is more than the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the purpose, then the return contract, then failure modes. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return-value explanation is technically redundant, but the note about FLECS reflection serialization adds nuance the schema likely cannot express. Combined with documented error conditions, the definition is complete enough, with only the sibling-differentiation gap remaining.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the dotted-entity-path and component/pair syntaxes are already fully documented in the schema. The description adds no additional parameter meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get the value of a single component of an entity') and the '[READ]' prefix signals the operation class. It does not, however, distinguish itself from the sibling flecs_get_entity, which an agent might reasonably pick instead when it wants a component's value.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 'single component' scoping and the error cases (missing component, tag with no data, unresolvable id), which hint at when the call will fail. There is no explicit guidance on when to prefer this over flecs_get_entity or flecs_query.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_get_entityFlecs Get EntityA
Read-onlyIdempotent

[READ] Get one entity with its tags, relationship pairs and component values.

Returns the FLECS entity JSON: 'parent', 'name', 'tags' (list), 'pairs' (relationship -> target), 'components' (component -> value; null when the component has no reflection data) and, when requested, 'id', 'type_info', 'inherited', 'matches' and doc info. Names use full paths.

Errors when the entity does not exist. Use flecs_query to find entities.

ParametersJSON Schema
NameRequiredDescriptionDefault
docNoInclude flecs.doc information (doc names, descriptions).
entityYesEntity path in FLECS dotted notation, e.g. 'Sun.Earth' or 'flecs.core.World' (the 'parent' and 'name' fields of a query result joined with '.'), or a numeric entity id such as '#523'. Escape a literal '.' inside a name as '\.'.
valuesNoInclude component values.
matchesNoInclude the queries, systems and observers that match this entity ('matches'). Useful to find out why a system does or does not process an entity.
entity_idNoInclude the numeric entity id ('id').
inheritedNoInclude tags/components inherited from prefabs (IsA).
type_infoNoInclude the reflection schema of the returned components.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so the bar is lower. The description adds real value beyond them: it errors when the entity does not exist, components are null when there is no reflection data, names are full paths, and entity_id/type_info/inherited/matches/doc are opt-in outputs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the [READ] prefix and a one-line purpose, followed by a compact enumeration of return fields and a closing routing sentence. The return-field list is somewhat lengthy given an output schema exists, but every sentence carries information and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter read tool with full schema coverage, complete annotations and an output schema, the description covers purpose, alternative routing, error behavior and return shape. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including escaping of literal '.' and the '#523' numeric-id form is already documented in the schema. The description only restates that optional fields are included 'when requested', adding no syntax or format detail beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Get one entity') plus the [READ] tag, and enumerates exactly what is returned (tags, relationship pairs, component values). It also names a sibling ('flecs_query') so an agent can differentiate read-one from search-many without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit guidance: use flecs_query to find entities, and use the 'matches' flag to diagnose why a system does or does not process an entity. It stops short of an explicit when-not-to-use rule, but the routing to the alternative is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_get_pipeline_statsFlecs Get Pipeline StatsA
Read-onlyIdempotent

[READ] Get per-system timing statistics from the FLECS stats module.

'entries' contains one object per system: 'name', 'disabled', 'time_spent' and (for non-task systems) 'matched_entity_count' and 'matched_table_count', each metric as {avg, min, max}. With a pipeline, sync points ('multi_threaded', 'immediate', 'time_spent', 'commands_enqueued') are interleaved in execution order. Times are in seconds. Requires the stats module (FlecsStats).

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNoSampling window: '1s', '1m', '1h', '1d' or '1w'. Each window holds 60 samples.1s
historyNoReturn all 60 samples of the window (oldest first) instead of only the latest sample.
pipelineNoDotted path of a pipeline, e.g. 'flecs.pipeline.BuiltinPipeline', to get its systems in execution order including sync points. Omit to get all systems (unordered).

Output Schema

ParametersJSON Schema
NameRequiredDescription
periodYes
entriesYes
historyYes
pipelineYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: the FlecsStats module prerequisite, units (seconds), and the interleaving of sync points in execution order with a pipeline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then structured detail on the entries payload and pipeline behavior. Every sentence carries information; slightly dense but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema and full annotation coverage, the description only needs to frame purpose and behavior, which it does well, including the prerequisite module and return-field semantics. The main omission is comparative routing to sibling stats tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so 'period', 'history' and 'pipeline' are already documented in the schema. The description reinforces pipeline ordering and the seconds unit but adds little syntax or semantics beyond what the schema fields state; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get per-system timing statistics from the FLECS stats module') with clear scope, and is readily distinguishable from the sibling flecs_get_world_stats by being per-system rather than world-level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Mentions the required stats module and explains that supplying a pipeline yields ordered systems including sync points, which is useful context. However, it never explicitly states when to prefer this over flecs_get_world_stats or the other stats/list siblings, leaving the choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_get_type_infoFlecs Get Type InfoA
Read-onlyIdempotent

[READ] Get the reflection schema (members, types, units) of a component type.

Returns {component, has_reflection, schema}. 'schema' maps member names to [type, {unit, ...}] descriptors, e.g. {"x": ["float"], "y": ["float"]}. has_reflection is false (and schema null) when the type has no reflection data, in which case FLECS cannot serialize or set its value.

ParametersJSON Schema
NameRequiredDescriptionDefault
componentYesDotted path of a component entity, e.g. 'planets.Mass'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
schemaYes
componentYes
has_reflectionYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so safety is covered. The description adds genuine behavioral meaning beyond structured fields: has_reflection=false yields schema null and means FLECS cannot serialize or set the value, which is a meaningful caveat for the caller.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb and resource in the first clause, then describes the return shape and edge case in tight prose. The inline example of the schema mapping is slightly verbose but earns its place by clarifying the descriptor format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't fully describe returns, yet it usefully highlights the has_reflection/schema-null edge case. Annotations cover safety. The only gap is routing guidance relative to sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema fully documents 'component' including its dotted-path format example. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('reflection schema of a component type'), and immediately scopes it with '(members, types, units)'. This distinguishes it from siblings like flecs_get_component or flecs_list_components, which the agent can tell apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the description explains the has_reflection=false case, which hints at when the tool is useful, but never says when to prefer it over flecs_get_component or flecs_get_entity, nor lists prerequisites. No explicit when/when-not guidance or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_get_world_infoFlecs Get World InfoA
Read-onlyIdempotent

[READ] Describe the connected FLECS world; call it first to test the link.

Returns:

  • rest_url: the FLECS REST API this server talks to.

  • build_info: FLECS version, compiler, enabled addons and build flags (flecs.core.BuildInfo), or null if unavailable.

  • world_summary: live counters such as entity, table, component and query counts, frame count, fps, target fps, time scale and uptime (flecs.stats.WorldSummary), or null if the stats module is not imported.

  • notes: why a section is null, if any.

Fails if the FLECS REST API cannot be reached.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesYes
rest_urlYes
build_infoYes
world_summaryYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and non-destructive, so the safety profile is covered. The description adds genuinely new behavioral context beyond the structured fields: the tool fails outright if the REST API is unreachable, and any return section can be null with a 'notes' field explaining why (e.g. stats module not imported).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the [READ] tag and the primary action before any detail, which is exactly the right ordering. The bulleted return list is scannable, though it partially restates structure that an output schema already provides, so it earns slightly less than full marks for economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with annotation coverage and an output schema, this is complete: it explains the failure mode, the nullability of each section, and the diagnostic 'notes' field, while leaving return-value shape to the output schema. An agent has everything needed to call it correctly and interpret a degraded response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics for the description to carry; the 4 baseline applies. Nothing about the empty schema is misrepresented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Describe the connected FLECS world') and immediately disambiguates it from siblings by scoping it to the world rather than an entity, component, query, or stats subset. It also fixes its ordinal role ('call it first'), so an agent can place it relative to the other ten tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use rule: call it first, as a connectivity/link test. It does not, however, address overlap with the close sibling flecs_get_world_stats, which appears to surface a subset of the same world_summary counters, and there are no explicit when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_get_world_statsFlecs Get World StatsA
Read-onlyIdempotent

[READ] Get world performance statistics from the FLECS stats module.

'metrics' maps names such as 'performance.fps', 'performance.frame_time', 'entities.count', 'tables.count', 'queries.system_count', 'commands.add_count' and 'memory.alloc_count' to {avg, min, max, brief}. Times are in seconds. Requires the FLECS application to import the stats module (FlecsStats); returns an explanatory error otherwise.

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNoSampling window: '1s', '1m', '1h', '1d' or '1w'. Each window holds 60 samples.1s
historyNoReturn all 60 samples of the window (oldest first) instead of only the latest sample.

Output Schema

ParametersJSON Schema
NameRequiredDescription
periodYes
historyYes
metricsYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/non-destructive, so the safety profile is covered; the description goes further by disclosing the stats-module dependency, the graceful failure mode, and that all time values are in seconds. That is exactly the kind of environment/auth-style context the annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the [READ] marker and the core action, then the metric key examples and the prerequisite. No filler, no repetition of the schema's parameter text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return-value description is a bonus rather than a requirement, and the description nonetheless sketches the metric map. Combined with the module prerequisite and failure behavior, an agent has everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so period and history are already fully documented in the schema and the baseline is 3. The description adds value above that baseline by explaining what the 'metrics' map keys look like and the {avg, min, max, brief} shape of each value, plus the seconds unit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Get world performance statistics from the FLECS stats module') and enumerates the metric namespaces returned (performance.*, entities.*, tables.*, queries.*, commands.*, memory.*), which pins down the scope precisely. It does not explicitly distinguish itself from the sibling flecs_get_pipeline_stats, so it falls just short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states a prerequisite ('requires the FLECS application to import the stats module') and what happens when unmet ('returns an explanatory error otherwise'), which is genuinely useful gating information. However, it never says when to prefer this over flecs_get_pipeline_stats or flecs_get_world_info, so the use-case routing is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_list_componentsFlecs List ComponentsA
Read-onlyIdempotent

[READ] List the component, tag and pair ids in use, with storage statistics.

Each entry has 'name', 'entity_count', 'entity_size', 'tables' (table ids), 'traits' (e.g. Exclusive, CanToggle, (OnDelete,Remove)), 'type' (size, alignment, which lifecycle hooks are set; absent for tags), 'sparse' (for sparse components) and 'memory' when the stats module is imported. Includes FLECS builtin ids and wildcard records. Returns {total, offset, limit, components}; total counts matches after filtering. Use flecs_get_type_info for a component's member schema.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of items to return.
offsetNoNumber of items to skip (for paging).
name_containsNoOnly return entries whose name contains this text (any case).

Output Schema

ParametersJSON Schema
NameRequiredDescription
limitYes
totalYes
offsetYes
componentsYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so this only needs to add context beyond that. It does: it discloses that FLECS builtin ids and wildcard records are included, that 'type' is absent for tags, and that 'memory' appears only when the stats module is imported — genuinely useful behavioral caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the [READ] tag and one-line purpose, then structured field-by-field detail. The entry-field enumeration is dense but earns its place by telling the agent what each record contains; slightly more verbose than strictly necessary given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a paginated read lister with annotations and an output schema, the agent has everything: what is listed, what each entry holds, what is included (builtins, wildcards), and which sibling to use for member schemas. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit/offset/name_contains are fully documented in the schema itself; baseline 3 applies. The description's note that 'total counts matches after filtering' marginally clarifies filter semantics but adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List the component, tag and pair ids in use, with storage statistics.' Scope (in use) and the [READ] marker distinguish it clearly from siblings like flecs_get_component (single component) and flecs_get_type_info (member schema).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use flecs_get_type_info for a component's member schema', clearly delineating this listing tool from the schema-introspection sibling. No explicit when-not conditions or guidance on paging toward other listers, but the primary alternative is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_list_queriesFlecs List QueriesA
Read-onlyIdempotent

[READ] List named queries, systems and observers with evaluation statistics.

Each entry has 'name' (usable with flecs_run_named_query), 'kind' (Query, System or Observer), 'expr' (the query expression), 'results' and 'count' (current matches), 'eval_count', 'eval_time', 'eval_mode', 'cache_kind', 'batched', 'empty_tables', 'plan_size' and, when the stats module is imported, 'memory'. FLECS evaluates every query to produce this list, which can take a while in very large worlds. Returns {total, offset, limit, queries}.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOnly return entries of this kind.
limitNoMaximum number of items to return.
offsetNoNumber of items to skip (for paging).
include_plansNoInclude the query plan text ('plan', 'cache_plan').
name_containsNoOnly return entries whose name contains this text (any case).

Output Schema

ParametersJSON Schema
NameRequiredDescription
limitYes
totalYes
offsetYes
queriesYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely new behavior: every query is evaluated to build the list (a real performance cost), and the 'memory' field appears only when the stats module is imported. That is useful operational context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and the [READ] tag, then enumerates fields efficiently. The field-by-field enumeration is somewhat redundant given an output schema exists, but the cross-reference and cost caveat justify most of the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a paginated listing tool with a full output schema, it covers what the schema cannot: the evaluation cost, the conditional stats-module field, and the downstream link to flecs_run_named_query. An agent has everything needed to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents kind, limit, offset, include_plans and name_contains. The description adds no parameter-level syntax or filtering guidance, so the baseline 3 applies; it earns no extra credit here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List named queries, systems and observers') and names the exact scope of the results (evaluation statistics). It even ties the returned 'name' field back to the sibling flecs_run_named_query, so an agent can distinguish this discovery tool from the execution tool without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: use this to discover named queries/systems/observers whose names feed flecs_run_named_query, and warns that listing triggers full evaluation and 'can take a while in very large worlds.' It does not explicitly say when to prefer flecs_query, flecs_explain_query, or flecs_get_pipeline_stats instead, so there are no hard exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_queryFlecs QueryA
Read-onlyIdempotent

[READ] Find entities with a query written in the FLECS query language.

Syntax cheat sheet (terms are comma separated and all must match):

  • 'Position, Velocity' entities with both components

  • 'Position, !Velocity' ... without Velocity

  • 'Position, ?Mass' Mass is optional

  • 'Planet || Moon' either tag

  • '(ChildOf, Sun)' children of entity Sun

  • '!(flecs.core.ChildOf, *)' root entities

  • '(Likes, *)' / '(Likes, $x)' wildcard / variable pair target

  • 'Position, Mass(up)' Mass on an ancestor (ChildOf)

  • 'SpaceShip, $this ~= "Uss"' name contains "Uss"

  • 'IsA(_, *)' instances of any prefab Use full paths (e.g. 'transform.Position') when names are ambiguous.

Returns {page, results, type_info?}. Each result has 'parent', 'name' and, depending on options, 'id', 'fields' ({values, ids, sources, is_set} per query term), 'vars' or, with table=true, 'tags'/'pairs'/'components'. page = {offset, limit, returned, may_have_more}; when may_have_more is true, call again with offset += limit. Invalid queries return the FLECS parser error (with position) so the query can be fixed.

ParametersJSON Schema
NameRequiredDescriptionDefault
docNoInclude flecs.doc information (doc names, descriptions).
limitNoMaximum number of items to return.
queryYesQuery in the FLECS query language, e.g. 'Position, Velocity'.
tableNoReturn every tag, pair and component of each matched entity (same format as flecs_get_entity) instead of only the query fields.
fieldsNoInclude per-term field data ('fields') for each result.
offsetNoNumber of items to skip (for paging).
valuesNoInclude component values.
inheritedNoWith table=true: include components inherited from prefabs.
type_infoNoInclude the reflection schema of the returned components.
entity_idsNoInclude numeric entity ids ('id').

Output Schema

ParametersJSON Schema
NameRequiredDescription
pageYes
resultsYes
type_infoNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description adds real value beyond them: it documents the failure mode ('Invalid queries return the FLECS parser error with position') and the paging completeness signal ('may_have_more'), which the annotations do not cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, followed by the syntax reference and then the return shape, which is a logical reading order. The cheat-sheet block is long but every example encodes a distinct query-language feature, so it earns its space rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex query-language tool with 10 parameters, the description covers syntax, output shape, paging semantics, and error behavior. With an output schema already present it doesn't need to enumerate return values, yet its summary of the result fields is accurate and sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so a baseline of 3 would suffice. The description earns a higher score by providing a full syntax cheat sheet for the core 'query' parameter and by explaining the table=true behavior and its effect on the returned shape, adding semantics beyond the terse schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Find entities with a query written in the FLECS query language'), making the mechanism clear. It is implicitly distinguished from the sibling flecs_run_named_query by taking an inline query, but it never names that sibling to make the contrast explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete operational guidance for paging ('when may_have_more is true, call again with offset += limit'), which is genuinely useful. However, it never states when to prefer this tool over alternatives like flecs_run_named_query or flecs_get_entity, so usage is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flecs_run_named_queryFlecs Run Named QueryA
Read-onlyIdempotent

[READ] Return what an existing named query, system or observer matches now.

Useful to check which entities a system processes. Same result format as flecs_query: {page, results, type_info?}.

ParametersJSON Schema
NameRequiredDescriptionDefault
docNoInclude flecs.doc information (doc names, descriptions).
nameYesDotted path of an existing named query, system or observer, as listed by flecs_list_queries (e.g. 'game.systems.Move').
limitNoMaximum number of items to return.
tableNoReturn every tag, pair and component of each matched entity (same format as flecs_get_entity) instead of only the query fields.
fieldsNoInclude per-term field data ('fields') for each result.
offsetNoNumber of items to skip (for paging).
valuesNoInclude component values.
inheritedNoWith table=true: include components inherited from prefabs.
type_infoNoInclude the reflection schema of the returned components.
variablesNoOptional values for query variables as 'var:entity' pairs, e.g. 'parent:Sun' or 'x:e1,y:e2'.
entity_idsNoInclude numeric entity ids ('id').

Output Schema

ParametersJSON Schema
NameRequiredDescription
pageYes
resultsYes
type_infoNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is fully covered and the '[READ]' tag is largely redundant. The description does add the snapshot semantics ('matches now') and points to the result shape {page, results, type_info?}, which the annotations do not. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the read marker and core purpose, then use case, then return format. Nothing is padded, though the '[READ]' tag duplicates the annotation rather than earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a full output schema and rich annotations, the description only needs to convey purpose, usage, and return shape, and it does all three. It stops short of routing between siblings (flecs_query vs. this vs. flecs_list_queries), which is the one gap an agent might still need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 11 parameters, including good detail on 'name' (dotted path, example, cross-reference to flecs_list_queries) and the interplay between table/inherited/values. The description adds no parameter-level meaning beyond this, so the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return what an existing named query, system or observer matches') and names the key constraint that the query already exists by name. It partially distinguishes itself from flecs_query by noting the shared result format, though it never explicitly contrasts the two (ad-hoc vs. named), leaving sibling differentiation implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful to check which entities a system processes' gives one concrete scenario, which is better than nothing. But there is no explicit when-to-use/when-not guidance, no mention of when to prefer this over flecs_query or flecs_list_queries, and no prerequisites (e.g. that the name must come from flecs_list_queries). Usage is implied rather than directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 11 tool updatesv0.1.0
    • First observedflecs_explain_query
    • First observedflecs_get_component
    • First observedflecs_get_entity
    • First observedflecs_get_pipeline_stats
    • First observedflecs_get_type_info
    • First observedflecs_get_world_info
    • First observedflecs_get_world_stats
    • First observedflecs_list_components
    • First observedflecs_list_queries
    • First observedflecs_query
    • First observedflecs_run_named_query

TDQS

A4/5.0

Scored across 11 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: entity vs component retrieval, ad-hoc vs named query execution vs query explanation, world info vs stats vs pipeline stats, and list vs schema inspection. The descriptions explicitly clarify boundaries (e.g., use flecs_query to find entities, flecs_run_named_query for existing named queries), so misselection risk is low.

Naming Consistency4/5

All tools use the flecs_ prefix and snake_case, with a consistent verb_noun pattern in nearly every case (get_world_info, get_entity, list_components, run_named_query, etc.). The one minor deviation is flecs_query, which is verb-only rather than verb_noun, but overall the convention remains predictable.

Tool Count5/5

Eleven tools is well-scoped for a FLECS inspection server, covering world info, entities, components, queries, and statistics without redundancy. Each tool earns its place and the surface is neither thin nor bloated.

Completeness4/5

The read-only introspection surface is thorough: world info, entity/component retrieval, query execution, query explanation, type schemas, lists, and performance stats. However, mutation operations (e.g., creating entities, setting components, running systems) are absent, which could be a gap if the server is intended for more than read-only diagnostics.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables live runtime inspection of any Python application, allowing MCP clients to query state, evaluate expressions, inspect objects, and read source code while the app runs.
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to drive the Unity Editor directly over HTTP, including scene and asset editing, diagnostics, play mode control, code execution, input simulation, and builds, with no separate MCP server process.
    325
    MIT