Skip to main content
Glama

agentic3d

An agent for instruction-guided editing of 3D indoor scenes. Given a scene mesh and a natural-language instruction such as "remove the table and move the chair to where it was" or "add a sofa against the wall", the system localizes the objects the instruction refers to, applies the corresponding geometric edits, and checks the result against the instruction from rendered views.

The implementation runs on CPU or Apple Silicon (MPS) with no CUDA dependency: rendering is done with ray casting and a lightweight shading pass rather than a GPU rasterizer, object localization uses frozen open-vocabulary 2D models (Grounding DINO and SAM), and meshes for "add" operations are retrieved from a local Objaverse index.

Method

The pipeline is organized into the following stages, each implemented as a module in agentic3d/:

  1. Rendering (render). Multi-view ray casting produces, per view, a depth map, a per-pixel hit-face index, and a shaded RGB image. The hit-face index makes the mapping from image pixels back to mesh faces exact.

  2. 2D segmentation (segment2d). Grounding DINO proposes boxes for the target phrases and SAM converts them to masks, on each rendered view.

  3. Mask lifting (lift). Per-view masks are back-projected onto mesh faces and aggregated by voting across views; the vote field is thresholded, cleaned, and split into connected components to yield one submesh per object instance.

  4. Scene editing (edits). Instances populate an editable scene graph that supports add, remove, move, rotate, scale, and replace operations, with floor snapping and axis-aligned collision handling. Composing the graph produces the output mesh.

  5. Agent loop (agent). A language model plans a sequence of edit operations from the instruction and the scene description, the operations are executed, and a vision-language model verifies the rendered result. Failed verifications trigger revision and re-planning.

Objects for "add" operations are supplied by retrieval, which matches the query to Objaverse-LVIS categories by CLIP text similarity and re-ranks a small set of candidate meshes by CLIP image similarity.

Related MCP server: Hayba

Installation

conda env create -f environment.yml
conda activate agentic3d
pip install -e .
python -m ipykernel install --user --name agentic3d --display-name "Python (agentic3d)"

The agent and evaluation code read an OpenAI API key from a .env file at the repository root (OPENAI_API_KEY=...).

Usage

A single instruction is applied with pipeline.edit_scene:

from agentic3d import pipeline

edited_mesh, scene_graph, trajectory = pipeline.edit_scene(mesh, "add a sofa against the wall")
edited_mesh.export("edited.glb")

mesh is a file path or a trimesh.Trimesh. The returned trajectory records the planned operations, each operator's report, and the verifier's verdict.

EditSession applies several instructions to the same scene, reusing the segmentation and scene graph between calls:

from agentic3d.pipeline import EditSession

session = EditSession(mesh, targets=["chair", "table"])
session.apply("remove the table")
session.apply("add a sofa where the table was")
session.mesh().export("edited.glb")

MCP server

agentic3d/mcp_server.py exposes the scene-editing pipeline over the Model Context Protocol as a stdio server, so an MCP client (for example Claude Desktop or the MCP Inspector) can drive it directly. The tools are load_scene, segment_scene, get_scene_info, add_object, remove_object, move_object, rotate_object, scale_object, replace_object, undo, render_scene (returns a rendered image), save_scene, and edit_with_agent, which runs the full plan / apply / verify / revise loop for one instruction with an internal model and returns a summary and a render. Two prompt templates are also provided: edit_scene(scene_path, instruction) and describe_scene(scene_path).

After pip install -e . the server is on the path as agentic3d-mcp (equivalently python -m agentic3d.mcp_server). Point an MCP client at the environment's executable directly:

{
  "mcpServers": {
    "agentic3d": {
      "command": "/path/to/conda/envs/agentic3d/bin/agentic3d-mcp"
    }
  }
}

conda run also works but needs --no-capture-output, otherwise it buffers the server's stdio and the connection never completes:

{
  "mcpServers": {
    "agentic3d": {
      "command": "conda",
      "args": ["run", "--no-capture-output", "-n", "agentic3d", "agentic3d-mcp"]
    }
  }
}

Or inspect it interactively:

npx @modelcontextprotocol/inspector conda run --no-capture-output -n agentic3d agentic3d-mcp

scripts/make_demo_scene.py writes a room.glb (a floor, two walls, and a chair, table, and potted plant retrieved from Objaverse) to try the server against:

python scripts/make_demo_scene.py
agentic3d-mcp --test room.glb chair table plant

A typical session calls load_scene, then segment_scene(["chair", "table", "plant"]), then the edit tools, then render_scene and save_scene. The client's own model reads the images returned by render_scene to decide whether the edit matched the request.

Evaluation

eval.evaluate_edit scores an edit:

from agentic3d import eval

scores = eval.evaluate_edit(input_mesh, edited_mesh, instruction)

By default the score is reference-free: a vision-language model rates the result from 0 to 5 on instruction adherence, physical plausibility, and spatial layout. Passing target_mesh= adds Chamfer distance and axis-aligned bounding-box IoU against a hand-authored target. The evaluation notebook runs a small sweep of tasks and plots the resulting scores.

Repository structure

agentic3d/               library modules (see Method)
agentic3d/mcp_server.py  Model Context Protocol server
notebooks/               one notebook per module, for inspecting intermediate results
scripts/                 helper scripts (demo scene generation)
pyproject.toml           package metadata and the agentic3d-mcp entry point

The notebooks depend on the agentic3d package and on the embreex ray-mesh intersection backend that the project environment installs.

Available Tools

13 tools
add_objectA

Retrieve a 3D asset from Objaverse by text query and place it on the floor at world (x, y) with the given yaw in degrees. label names the new object; category is an optional single noun to steer retrieval.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
labelYes
queryYes
yaw_degNo
categoryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose meaningful behavior: external retrieval from Objaverse, floor-level placement with world coordinates, and yaw in degrees. However, it does not disclose scene-mutation side effects, prerequisites (e.g., needing a loaded scene), failure behavior on a poor retrieval, or permissions. Partial but not complete disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence front-loads the primary action and placement semantics; the second efficiently packs parameter clarifications, which is justified given the 0% schema description coverage. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. With six parameters and no annotations, the description covers the full purpose and every parameter. Remaining gaps are non-critical for making a correct call: no sibling-routing guidance and no disclosure of side effects or scene prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% — the schema provides only bare titles. The description compensates fully: query is the Objaverse text query, label names the new object, x/y are world floor coordinates, yaw_deg is the yaw in degrees, and category is an optional single noun to steer retrieval. All six parameters receive semantic meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Retrieve a 3D asset from Objaverse by text query and place it on the floor at world (x, y)') with a clear resource and placement semantics. It is readily distinguishable from siblings like move_object (repositions existing objects), replace_object (swaps existing objects), and load_scene (loads a scene), so an agent can tell what this tool uniquely does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied — the description makes clear this adds a newly retrieved object to the scene — but it never explicitly says when to use this versus move_object, replace_object, or scale_object, nor does it state any exclusions or prerequisites (e.g., a scene must be loaded first). No alternative tool is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_with_agentA

Run the full editing agent on one instruction with an internal model: it picks the objects to segment, plans and applies the edits, verifies the result from renders, and revises. Pass scene_path to load and segment a fresh scene, or omit it to act on the scene already loaded. Returns a summary of what the agent did and a render.

ParametersJSON Schema
NameRequiredDescriptionDefault
scene_pathNo
instructionYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the disclosure burden. It transparently describes the agent's autonomy—selecting objects, planning, applying, verifying, and revising—and states the return value. It does not mention persistence, undoability, or cost/time, but the core behavioral cycle is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: a front-loaded statement of what the tool does, a conditional parameter explanation, and the return summary. There is no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex agent tool with no output schema and no annotations, it provides the essential invocation context: required instruction, optional scene_path behavior, workflow, and return value. It leaves minor gaps around persistence and relationship to save_scene/undo, but nothing that blocks a competent agent from calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully explains the scene_path option and its default behavior. The required instruction parameter is only described as 'one instruction,' leaving exact format and example unstated, though context makes it reasonably inferable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action—'Run the full editing agent on one instruction'—and enumerates the agent's pipeline (segment, plan, edit, verify, revise). This makes it unmistakably distinct from the granular sibling tools like add_object or segment_scene.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditional usage for scene_path ('pass to load... or omit to act on the loaded scene'), which is useful. However, it does not explicitly state when an agent should choose this full-agent tool over the individual sibling operations, or when the granular tools are preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_scene_infoA

Return the current scene info: object ids, labels, world centers and sizes, base heights, collisions, floor_z, and room_bounds.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It states the operation is a read ('Return') and lists the returned data, which signals read-only behavior. However, it does not elaborate on potential side effects, permissions, or error conditions, though for a getter this is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the verb and resource and then lists the returned fields. No filler words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, an output schema exists, and the description enumerates the returned scene attributes, an agent can call this tool correctly without additional information. The description is sufficient for this simple read-only tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description cannot add parameter meaning. Baseline for 0 params is 4, and the description appropriately focuses on the output instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('current scene info') and enumerates the exact content fields (object ids, labels, world centers, etc.), making it distinct from sibling tools like render_scene or load_scene. No ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage – call this to retrieve current scene state – but does not explicitly state when to use it over alternatives or mention any exclusions. No reference to sibling tools or when this is preferred, leaving the agent to infer from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_sceneA

Load a scene mesh (.glb/.obj/.ply). Creates an uneditable scene shell; call segment_scene next to make objects editable. Returns face count, bounds, and floor height.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden — and it delivers by revealing a non-obvious behavioral trait: the loaded scene is an uneditable shell that requires segment_scene. It also states the return values (face count, bounds, floor height). It doesn't mention side effects like replacing the current scene, but for a one-parameter load tool this is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero filler: the core action, the behavioral consequence, and the return values each earn their place. The purpose is front-loaded in the first sentence, and the workflow pointer in the second adds high-value routing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with an output schema present, the description covers purpose, file formats, workflow position, behavioral state, and return values. The only meaningful omission is whether loading a new scene replaces the existing scene, which is minor given the tool's simplicity and available output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented 'path' parameter. It partially does by specifying accepted file formats (.glb/.obj/.ply), implying path must reference a supported mesh file, but it never explicitly defines what path means, its constraints, or acceptable source locations. Adequate but not exhaustive compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Load a scene mesh') and enumerates supported formats (.glb/.obj/.ply). It differentiates from siblings by explicitly naming segment_scene as the follow-up, so an agent can distinguish loading from the 12 sibling tools without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit sequential guidance ('call segment_scene next to make objects editable'), which establishes when this tool fits into the workflow and why the result is not immediately editable. It stops short of enumerating when-not conditions for the other siblings, but the loading-then-segmenting context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

move_objectA

Move an object by id to world position (x, y). It stays on the floor and inside the room; the returned report notes any unresolved collision.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
object_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It states key behaviors: 'It stays on the floor and inside the room' (constraints) and 'the returned report notes any unresolved collision' (output behavior). This goes beyond a simple 'move object' and informs the agent about expectations and side effects. It could mention whether the operation is reversible or if it triggers any other side effects, but for a simple move it's quite informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary action and target, and includes critical behavioral constraints in a compact form. No wasted words; every phrase adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are not required in the description. With 3 required parameters and no schema descriptions, the description should clarify parameter semantics further. It vaguely implies id and x/y but not explicit mapping. However, given the output schema exists and the behavior is clear, it's nearly complete. Could be improved by explicitly mentioning parameter names and units, but not severely lacking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the parameters are entirely undocumented in the schema. The description mentions 'id' and 'world position (x, y)', but does not explicitly label the parameters or provide constraints (e.g., bounds for x and y). Since the schema has no descriptions, the tool description must compensate, but it only partially does so by implying the meaning but not detailing each parameter. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Move'), the resource ('object'), and the target ('world position (x, y)'). It distinguishes itself from siblings like rotate_object and scale_object, but doesn't explicitly mention alternative move-related tools, though none exist in the sibling list. It's clear and concise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (when moving object is needed), but provides no explicit exclusions or alternatives. It doesn't state conditions like 'not for moving the camera' or 'use replace_object for changing object type'. However, the sibling context shows no similar move tool, so clarity in purpose partially compensates. Still, a clear when-to-use statement is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_objectA

Remove an object by id. The floor and walls under it are filled in.

ParametersJSON Schema
NameRequiredDescriptionDefault
object_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly states the side effect that the floor and walls under the removed object are filled in, which is valuable behavioral context beyond the simple 'remove' action. It does not mention reversibility or other side effects, but the disclosed side effect is significant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The core action is front-loaded, and the important side effect is stated immediately after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description is mostly complete. It explains the action and the key side effect. It could mention whether the operation is reversible or what the output contains, but the output schema likely covers return values, and the side effect disclosure is sufficient for a simple removal tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions the object is removed 'by id', which aligns with the single object_id parameter, but it does not add detail about the id format or how to obtain it. The description adds minimal meaning beyond the schema's property name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Remove') and resource ('an object by id'), and the additional detail about floor/walls being filled in helps distinguish it from other object-manipulation siblings like move_object or replace_object. It is clear but does not explicitly name sibling alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: call it when you want to delete an object by its id. It does not explicitly state when not to use it or mention alternatives like replace_object or undo, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

render_sceneA

Render the current scene from one orbiting camera and return the image. azimuth_deg orbits around the vertical axis, elevation_deg tilts up.

ParametersJSON Schema
NameRequiredDescriptionDefault
azimuth_degNo
elevation_degNo

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It explains the camera parameters and that it returns an image, but doesn't disclose whether it modifies the scene, any permissions needed, or limitations like requiring a loaded scene. This is a moderate disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no filler. The main action is front-loaded, and the parameter explanations are immediately relevant. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two parameters and no output schema, the description covers the action and parameters. It omits explicit statement that it's a read-only operation, but the verb 'render' implies it. Minor gaps like error handling are not critical for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description fully compensates by explaining what azimuth_deg and elevation_deg control: orbiting and tilting. This is precisely the semantic meaning an agent needs beyond the schema's default values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('render'), the resource ('current scene'), and the output ('image'). It clearly distinguishes from sibling tools like load_scene or get_scene_info, which deal with loading or info retrieval. The camera behavior is also explained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context that this tool produces an image of the scene, implying it's for visual output. However, it doesn't explicitly mention when not to use it or alternative tools, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replace_objectA

Replace an object by id with a freshly retrieved Objaverse asset, keeping its footprint position. category is an optional single noun for retrieval.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
categoryNo
object_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It adds useful behavior details: the asset is freshly retrieved from Objaverse and the footprint position is kept. It does not explain whether the replacement is destructive, reversible via undo, or how other transforms like scale and rotation are affected, so transparency is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The operation and spatial behavior are front-loaded, and the parameter note about category is placed in a separate sentence where it is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter mutation tool with no annotations and 0% schema coverage, the description gives a solid overview but leaves the required query parameter under-defined. The existence of an output schema reduces the need to describe return values, but the query gap makes the definition incomplete for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains category ('optional single noun for retrieval') and implies object_id through 'by id', but it never explicitly defines the required query parameter or how it relates to category. The agent is left to infer the query semantics from 'freshly retrieved Objaverse asset'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Replace'), the target resource ('an object by id'), and the source ('freshly retrieved Objaverse asset'), plus the key spatial constraint that footprint position is preserved. This clearly distinguishes replace_object from sibling tools like add_object, remove_object, and move_object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: swap an existing object with a current Objaverse asset while keeping its footprint. However, it gives no explicit when-to-use/when-not-to-use guidance or alternatives, so an agent must infer the decision from the tool name and scene context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rotate_objectA

Rotate an object by id about the vertical axis by yaw_deg degrees.

ParametersJSON Schema
NameRequiredDescriptionDefault
yaw_degYes
object_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description itself carries the behavioral burden, and it largely succeeds: it specifies rotation about the vertical axis and describes yaw_deg as a delta in degrees. It does not disclose the coordinate reference frame or error behavior for unknown object IDs, but the core side effect is clearly communicated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. Every phrase adds meaning: target, axis, and angle.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter rotation operation with an output schema present, the description gives the essential semantics. The main omissions—reference-frame details and ID-format specifics—are relatively minor for such a simple operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds 'degrees' and 'vertical axis' context for yaw_deg and identifies object_id as the selection key, but it does not specify valid ranges, ID format, or angle conventions beyond the schema's type information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a precise operation (rotate), the target (object by id), the axis (vertical), and the control (yaw_deg degrees). This clearly differentiates it from sibling operations like move_object and scale_object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It is easy to infer when to use this tool—when an object needs a rotational adjustment—but the description gives no explicit guidance on when not to use it or how it relates to the move/scale/replace sibling tools. The intended usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_sceneB

Write the current composed scene to a .glb file at path.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It only states the write action and omits material behavior for a persistence operation: whether an existing file gets overwritten, whether the scene must be non-empty, whether the operation is reversible, and whether it mutates the in-memory scene state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 13-word sentence with no filler; the verb and resource are front-loaded. A minor deduction because 'at path' reads as truncated phrasing that could be clearer ('to the specified path'), but the sentence remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (1 parameter, output schema present) and the description covers the core action and destination. But for a mutation operation with zero annotation coverage, the missing overwrite semantics and preconditions (e.g., a scene must be composed before saving) keep it from being fully self-sufficient. Adequate, with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for the sole 'path' parameter, and it partially does: 'Write ... to a .glb file at path' clarifies that path is the destination file location. However, it leaves ambiguity about whether parent directories are created, whether the .glb extension is enforced, and what path formats are accepted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Write') and identifies a precise resource ('current composed scene') and target format ('.glb file at path'). This lets an agent instantly distinguish it from siblings like load_scene, render_scene, and get_scene_info without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no exclusions, and no mention of the natural inverse workflow (load_scene) or the sibling render_scene for visual output. Usage is only implied by the verb and resource; no conditions or recommended scenarios are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scale_objectA

Scale an object by id by a multiplicative factor (>1 larger, <1 smaller). The base stays on the floor and the size is clamped to the room.

ParametersJSON Schema
NameRequiredDescriptionDefault
factorYes
object_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does disclose meaningful traits: the base stays on the floor and the size is clamped to the room. It does not mention reversibility, failure conditions, or persistence, but for a simple scaling operation the provided behavior is informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the action and parameter behavior, with no wasted words. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential behavior needed to call the tool correctly, and the presence of an output schema reduces the need to describe return values. Minor gaps like error handling or id format remain, but they are not critical for this low-complexity tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the factor parameter well, including multiplicative semantics and size direction, but does not elaborate on object_id format or edge cases like zero or negative factors. The object_id meaning is fairly obvious from its name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('scale') with a resource ('an object by id') and defines the core semantics of the factor (>1 larger, <1 smaller). It clearly differentiates this from sibling operations like move_object and rotate_object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the name and description, but no explicit when-to-use or when-not-to-use guidance is given. Alternatives such as replace_object or edit_with_agent are not mentioned, so the agent must infer selection from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

segment_sceneC

Segment the loaded mesh into editable objects for the given names (e.g. ['chair', 'table']). Rebuilds the scene graph. Returns the scene info: object ids, labels, world centers and sizes, floor_z, room_bounds.

ParametersJSON Schema
NameRequiredDescriptionDefault
objectsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses one behavioral trait: 'Rebuilds the scene graph,' implying a mutating operation. However, with no annotations provided, it fails to state whether existing scene objects are destroyed, whether the operation is undoable, what happens if the mesh is not loaded, or any permission requirements. These are significant gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no wasted words; the action is front-loaded and the example is placed early. The structure is efficient and scannable, though the meaning of 'given names' could have been clarified without bloating the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return field list may be documented externally. However, the description misses crucial context: how to determine valid names, the prerequisite of a loaded mesh, relationship to get_scene_info for read-only access, and side-effect risks. This is insufficient for safe, correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero schema description coverage, the description must clarify the 'objects' parameter. It provides an example (['chair', 'table']) but does not explain what these names refer to (e.g., labels in the source mesh, arbitrary names to assign), valid constraints, or error behavior for unmatched names. The semantics remain ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('segment') and resource ('loaded mesh') with a clear goal: producing editable objects for named targets. The example adds concrete context, and the side effect of rebuilding the scene graph distinguishes it from sibling tools like add_object, remove_object, or get_scene_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. It mentions the 'loaded mesh' prerequisite but does not instruct the agent to ensure load_scene was called first, nor does it warn that get_scene_info should be used instead when only scene information is needed without modifying the graph.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

undoB

Undo the last edit.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure, and it only restates the action. It does not disclose what happens if there is no edit to undo, whether the operation is destructive, whether multiple undos are supported, or how undo interacts with load_scene/save_scene.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler or repetition. It is appropriately compact for a zero-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and has an output schema plus no parameters, so a one-line definition is mostly adequate. However, the description leaves ambiguity around what 'last edit' covers and what happens with no prior edit, which are meaningful gaps for an agent deciding whether and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so there is nothing for the description to add beyond the schema, which is already fully complete. The baseline of 4 applies because no parameter semantics need to be compensated for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Undo') and a clear target ('the last edit'), so an agent can tell what the tool does. It does not explicitly define what counts as an 'edit' or contrast itself with sibling mutation tools, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Undo the last edit' implies the tool is used after an edit has been made and needs reverting. It gives no explicit when-to-use guidance, exclusions, or mention of alternatives such as re-applying a replaced object, so it only reaches implied-usage level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.1.0
    • First observedadd_object
    • First observededit_with_agent
    • First observedget_scene_info
    • First observedload_scene
    • First observedmove_object
    • First observedremove_object
    • First observedrender_scene
    • First observedreplace_object
    • First observedrotate_object
    • First observedsave_scene
    • First observedscale_object
    • First observedsegment_scene
    • First observedundo

TDQS

A3.9/5.0

Scored across 13 tools

Disambiguation5/5

Each tool maps to a distinct action in the 3D scene editing workflow: loading, segmenting, querying, adding/removing, transforming, undo, rendering, saving, and high-level editing. Even load_scene and segment_scene are clearly separated by their roles despite both returning scene info. No two tools appear to perform the same operation.

Naming Consistency5/5

Tool names consistently follow a snake_case verb_noun pattern (load_scene, add_object, render_scene). Minor deviations like get_scene_info and edit_with_agent are still verb-first and fit the broader convention. The naming is predictable and easy to scan.

Tool Count5/5

13 tools is well-scoped for a 3D scene composition server. Each tool covers a core lifecycle or editing operation without unnecessary redundancy. The count feels appropriate for the domain.

Completeness5/5

The server provides full coverage of the scene editing workflow: load, segment, inspect, add/remove, transform, replace, undo, render, and save. Object manipulation covers the essential operations agents need, and edit_with_agent offers a high-level fallback for complex changes. No obvious dead ends or critical missing operations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server enabling AI agents to author Unreal Engine 5 scenes directly, with tools for spawning actors, building PCG graphs, validating physics, generating terrain, and more through a single MCP connection.
    13
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to construct, edit, and export 3D models using geometric primitives and boolean operations, with multi-view rendering to facilitate spatial reasoning.
    7
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Enables natural language creation and refinement of Blender scenes through structured MCP tools, with persistent object identity, visual validation, and reversible edits.
    23
    MIT