Skip to main content
Glama

CAD Agent

AI-generated mechanical CAD product showcase

CAD Agent helps coding agents turn mechanical requirements into validated FreeCAD models. It combines requirement reasoning, knowledge reuse, standard-part provenance, deterministic CAD state, automatic validation, correction, final confirmation, and reusable Design Lessons.

Nothing here guesses a number. Sizing follows published standards, the result carries its own assumptions and limitations, and a finished model is confirmed by a process the agent cannot reach or forge. A learned surrogate may offer a fast screening estimate, but only as a calibrated interval with declared coverage that is refused outside its fitted domain and can never gate completion.

The package provides the mech-cad-design CLI and mech-cad-design-mcp server. A compatible coding agent performs the design reasoning, while an external FreeCAD GUI MCP performs interactive CAD work. The package does not embed a language model and does not replace engineering review.

Design process

User request
  → requirement clarification
  → short design proposal
  → one natural-language direction approval
  → knowledge retrieval
  → CAD modeling
  → automatic validation and correction
  → correction capture
  → final result
  → natural-language final confirmation
  → automatic Design Lesson evaluation
  → finish, or one decision before durable lesson publication

Related MCP server: fcgen-mcp

Core capabilities

  • Create new designs or edit read-only snapshots of existing FCStd/STEP models.

  • Retrieve matching Product Family Knowledge and Design Lessons when available.

  • Continue CAD work when knowledge has no match or its backend is unavailable.

  • Model interactively in FreeCAD and keep FCStd as the source of truth.

  • Size a spur gear drive from duty inputs (power, speeds, material, duty, life, safety factor) through ratio, module, forces, bending and contact stress, shaft diameter, and required bearing capacity, then model and validate the resulting pair against the same calculated values. Preliminary sizing evidence, not a strength certification; see its limitations.

  • Find purchasable standard parts through configured structured providers and, when they miss, extend the search to authoritative manufacturer, standards body, industry association, and attributable authorized-distributor sources.

  • Register selected CAD components with provider, manufacturer, part identity, source, license, validation evidence, and SHA-256 provenance.

  • Calculate from Shigley's Mechanical Engineering Design independently of any one component: stress states and principal stresses, static and fatigue failure criteria, deflection and columns, shafts, bolted joints, springs, rolling and journal bearings, gears, welds, clutches, brakes and belts. Pure standard library, with every shipped constant checked against an independent anchor and every curve fit marked as one.

  • Record a surrogate screening estimate produced outside this package, as a calibrated interval with its coverage, calibration provenance and declared training domain. An estimate whose query falls outside that domain is refused and the refusal is kept. Screening never gates completion and never counts as a passed validation check; see Surrogate screening.

  • Test a screening interval against the closed-form mechanics result for the same load case, so a miscalibrated surrogate is caught by the analytical answer rather than trusted over it.

  • Validate geometry, dimensions, placements, interfaces, assemblies, fasteners, BOM consistency, and visual evidence.

  • Bind completion to the exact FCStd SHA-256 and passed JSON, Markdown, and PNG evidence.

  • Re-verify the finished model in a separate process the agent does not control, under a SHA-256-pinned FreeCAD executable running digest-pinned scripts, and accept the result only when a host-issued nonce and the recorded model digest both come back unchanged. The agent writes its own validation report; this is the part of the evidence it cannot author.

  • Record every validation attempt in an append-only correction ledger, so a failure that was fixed is not lost when the next attempt is recorded.

  • Evaluate reusable lessons automatically after the user confirms the final model, including lessons derived from the mandatory checks this design failed and then corrected.

  • Feed published correction lessons back through ordinary knowledge retrieval, so a later design finds the defect before repeating it.

  • Store long-term Product Family profiles, Knowledge Assertions, and Design Lessons in a local SQLite database by default, with optional PostgreSQL for shared team use and an optional rebuild-only Neo4j projection.

MCP surfaces

The default design surface contains the complete design flow:

  • design_system_status

  • design_start

  • design_status

  • design_knowledge_retrieve

  • design_record_result

  • design_mistakes

  • design_gear_size

  • design_screening_record

  • design_screening_status

  • design_confirm

  • design_lesson_decide

  • standard_part_providers_get

  • standard_part_sources_status

  • standard_part_download_register

The separate knowledge-admin surface manages Product Family onboarding, knowledge search, Design Lesson supersession or revocation, and explicit Neo4j projection rebuilds.

Architecture

CAD Agent architecture

Design sessions live under designs/<design-id>/ as atomic JSON state, one authoritative model.FCStd, optional source snapshots, validation evidence, outputs, and an optional lesson review card. CAD creation and validation do not depend on PostgreSQL.

The knowledge store holds only durable Product Families, Knowledge Assertions, and Design Lessons, in local SQLite by default or PostgreSQL when configured. Neo4j is optional, rebuildable, and never authoritative. See Architecture and trust boundaries.

Learning from corrected mistakes

Every call to design_record_result appends one entry to the design's append-only correction ledger: the model hash, the validation outcome, and each failed check. Nothing in a later attempt rewrites an earlier one, so the record of what went wrong survives the fix.

When the user confirms the final model, the package groups the mandatory checks that failed on earlier attempts and passed on the confirmed model. Each such defect becomes one deterministic Design Lesson candidate carrying origin: validation_correction, its check signature, and how many attempts it cost. These candidates join any the agent proposes on the same immutable review card and follow the same single publication decision.

Derivation runs without a language model and never blocks: a design that made no mistakes derives nothing, an advisory-only failure derives nothing, and a malformed derivation is dropped rather than holding up a completed model. Once published, correction lessons are ordinary Design Lessons, so design_knowledge_retrieve returns them to later designs in the same scope.

Install and run

Python 3.12 or newer is required. There is no PyPI release yet, so install from a clone:

git clone https://github.com/bloodreaper005/cad-agent
cd cad-agent
python -m pip install .

mech-cad-design init \
  --workspace /path/to/mech-cad-design-workspace \
  --actor engineer \
  --organization example-org \
  --design-group example-group
mech-cad-design knowledge bootstrap \
  --workspace /path/to/mech-cad-design-workspace
export MECH_DESIGN_WORKSPACE=/path/to/mech-cad-design-workspace
mech-cad-design-mcp

The knowledge store is a local SQLite database inside the workspace, so no service has to be running. mech-cad-design status reports workspace, FreeCAD, and knowledge readiness as structured JSON and exits non-zero when setup is incomplete.

Windows PowerShell:

mech-cad-design init --workspace "D:\Mechanical Design Workspace" --actor engineer --organization example-org --design-group example-group
$env:MECH_DESIGN_WORKSPACE = "D:\Mechanical Design Workspace"
mech-cad-design-mcp

Command line

Every command prints one JSON document and exits 0 ready, 1 warning, 2 setup required, or 3 blocked. gear size uses the same scale: 0 sized, 1 sized with warnings, 3 rejected.

mech-cad-design init --workspace W --actor A --organization O --design-group G
mech-cad-design status --workspace W
mech-cad-design migrate --workspace W [--dry-run]

mech-cad-design design start --workspace W --design-id ID --title T \
  --requirements-json '{"capacity": 4}' --proposal P --approve "yes"
mech-cad-design design list --workspace W
mech-cad-design design open --workspace W --design-id ID
mech-cad-design design status --workspace W --design-id ID
mech-cad-design design mistakes --workspace W --design-id ID

mech-cad-design gear size --power-kw 7.5 --pinion-rpm 1450 --gear-rpm 480 \
  --pinion-material 20MnCr5_carburised_G2 --gear-material 20MnCr5_carburised_G2 \
  --duty moderate --life-hours 20000 --safety-factor 1.5 [--out sizing.json]

mech-cad-design screening record --workspace W --design-id ID \
  --screening-file screening.json \
  --query-json '{"material_class": "linear_elastic"}'
mech-cad-design screening status --workspace W --design-id ID

mech-cad-design family start --workspace W --onboarding-id OB \
  --family-id F --family-name N [--alias A]
mech-cad-design family analyze --workspace W --onboarding-id OB --analysis-file P
mech-cad-design family review --workspace W --onboarding-id OB --decision "approved"
mech-cad-design family publish --workspace W --onboarding-id OB
mech-cad-design family status --workspace W --onboarding-id OB

mech-cad-design knowledge bootstrap --workspace W
mech-cad-design knowledge import-postgres --workspace W --source-env E
mech-cad-design standard-parts providers [--category C]

design mistakes reports which mandatory validation checks this design failed and later corrected, and which defects are still outstanding.

design start is idempotent: repeating it with the same design intent resumes the existing job instead of creating a second one. design open reports where an existing job stands and which step comes next. Starting a design needs a configured FreeCADCmd; listing, opening, and reading design jobs do not.

Select knowledge administration only when needed:

MECH_DESIGN_MCP_TOOL_PROFILE=knowledge-admin mech-cad-design-mcp

The current acceptance target is official FreeCAD 1.1.3. Configure the exact FreeCADCmd path and SHA-256 in the workspace or environment. Durable knowledge works out of the box on the local SQLite store; set MECH_DESIGN_DATABASE_URL to use PostgreSQL instead, which has no pgvector requirement. Install the neo4j extra (python -m pip install '.[neo4j]') only when the optional relationship projection is wanted.

Project-owned Agent Skills

Operating boundaries

  • Generated models, reports, screenshots, databases, credentials, and customer-specific evidence stay outside the public repository.

  • Local MCP and database services remain bound to loopback interfaces.

  • A passed validation report proves only the checks that ran against one exact model revision. It is not FEA, manufacturing release, safety certification, or legal standards certification.

  • A surrogate screening estimate is a learned prediction with declared coverage, not an analysis. It is not FEA either, it proves nothing about the model, and no screening result can complete, block, or confirm a design.

  • Final engineering responsibility remains with the user or an authorized engineer.

Documentation

License

Project source is released under Apache-2.0. External dependencies, integrations, and assets retain their own licenses; see Third-Party Notices.

Available Tools

13 tools
design_confirmC

Confirm the model, evaluate lessons, or revise its pending review.

ParametersJSON Schema
NameRequiredDescriptionDefault
design_idYes
confirmation_textYes
review_revision_textNo
lesson_candidates_jsonNo[]

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It fails to explain side effects (e.g., whether confirmation is recorded, if revisions replace existing data), return format, or any state changes. The phrase 'evaluate lessons' hints at processing but lacks detail on outcomes or requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise in length, but it lacks clarity and structure. It front-loads multiple ambiguous actions instead of stating a primary purpose. While not verbose, the conciseness is not effective because it sacrifices comprehensibility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters (2 required) and an output schema, the description is grossly inadequate. It doesn't specify required inputs, optional behavior, or what the tool returns. An agent would need to open the schema and infer semantics from parameter names alone, which is risky. The description adds minimal context beyond the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all parameters. It indirectly references confirmation_text and review_revision_text ('confirm' and 'revise its pending review'), but it does not clarify design_id (which is required) or lesson_candidates_json (which is a JSON string with default '[]'). The description fails to define the role of each parameter or their expected formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Confirm the model, evaluate lessons, or revise its pending review' is vague and ambiguous. It lists three verbs but doesn't clearly define the resource ('the model' is unclear) or how these actions relate. It doesn't differentiate from sibling tools like design_lesson_decide or design_status, leaving an agent uncertain about the tool's exact function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any conditions, prerequisites, or exclusions. An agent cannot determine whether to call design_confirm or design_lesson_decide from the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_gear_sizeA

Size a spur gear drive: ratio, teeth, module, forces, stresses, shaft, bearings.

Preliminary sizing evidence, not a strength certification. See the result's limitations for what is not evaluated.

ParametersJSON Schema
NameRequiredDescriptionDefault
dutyYes
power_kwYes
life_hoursYes
options_jsonNo{}
gear_materialYes
safety_factorYes
gear_speed_rpmYes
pinion_materialYes
pinion_speed_rpmYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the behavioral disclosure burden. It honestly discloses that results are preliminary and not a strength certification, and directs the agent to the result's limitations for what is not evaluated. This is useful transparency beyond the raw schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core purpose is stated in the first sentence, and the limitation/qualification follows immediately. There is no filler, and every sentence contributes actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose and key limitation, and an output schema exists to explain return values. However, with 9 parameters, no annotations, and no input guidance beyond names, the description leaves the agent to infer several input requirements, making it only minimally complete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no parameter-specific meaning. Although some parameter names are self-explanatory (power_kw, pinion_speed_rpm), the description does not clarify expected units, value ranges, material formats, or the purpose of options_json. The description must compensate for low schema coverage but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Size a spur gear drive' and enumerates the covered aspects (ratio, teeth, module, forces, stresses, shaft, bearings). This clearly distinguishes the tool from the sibling workflow tools, none of which perform gear sizing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context by stating this is 'Preliminary sizing evidence, not a strength certification,' which tells the agent when the tool is appropriate and when it is not. It does not name explicit alternative tools, but the exclusion is meaningful and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_knowledge_retrieveC

Retrieve applicable knowledge without making retrieval a CAD gate.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
requiredNo
design_idYes
features_jsonNo{}
used_knowledge_ids_jsonNo[]

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only hints at a possible non-blocking intent with 'without making retrieval a CAD gate,' but it does not disclose side effects, retrieval scope, response behavior, or operational constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, with no redundant filler. However, it is more under-specified than usefully concise, and the CAD-gate phrase adds ambiguity rather than structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, 0% parameter coverage, and 5 parameters including two required fields, this one-line description is severely incomplete. It does not explain how an agent should construct a query, what the required inputs mean, or what response behavior to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with 5 parameters, and the description mentions none of them. The agent is left without any explanation of design_id, query, required, features_json, or used_knowledge_ids_json.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Retrieve applicable knowledge.' It is specific enough to convey the basic action, though it does not explicitly distinguish this tool from its siblings, and the phrase 'without making retrieval a CAD gate' is somewhat jargon-heavy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool instead of siblings such as design_status, design_record_result, or design_lesson_decide. There are no conditions, exclusions, or alternative tool references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_lesson_decideC

Approve all or selected review-card lessons, or decline publication.

ParametersJSON Schema
NameRequiredDescriptionDefault
design_idYes
decision_textYes
selected_lesson_numbers_jsonNo[]

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits but only reveals that the tool changes publication state. It does not mention authorization requirements, reversibility, effects on non-selected lessons, or how the default decision applies. The 'decline publication' hint is a thin signal with no detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with a clear verb and scope. It is concise and readable, but the brevity is achieved at the cost of leaving important operational details unexplained. Still, it earns a 4 for clarity of structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating decision tool with no annotations, the description lacks critical context: how to specify approval vs decline, how to choose which lessons to act on, and what design_id refers to. The presence of an output schema does not compensate for missing parameter semantics and workflow guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for undefined parameters. It does not explain what decision_text should contain, how selected_lesson_numbers_json selects lessons, or what the default '[]' means. The phrase 'all or selected' hints at the selection parameter but leaves format and values ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action — approving or declining publication of review-card lessons — and distinguishes between approving all or selected lessons. It is specific about the resource and operation, though it does not contrast with sibling tools like design_confirm or design_record_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as design_confirm or design_record_result. The description implies it is used after reviewing lessons but does not state prerequisites, workflow position, or when to choose this tool over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_listA

List every design job in the workspace with its current state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly communicates a read-only listing behavior, the workspace-wide scope, and the fact that each returned item includes current state. It does not discuss pagination or ordering, but that is reasonable for such a simple tool with an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that delivers all key information immediately: what is listed, at what scope, and what information is included. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with an output schema, the description is nearly complete. It clearly states the operation and the returned content. It falls slightly short only by not addressing whether all workspace design jobs include completed/failed/archived jobs or just currently active ones.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parametersches and 100% schema description coverage, so the baseline is 4. The description adds no parameter-specific meaning, but there are no parameters to explain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List'), a clear resource ('every design job in the workspace'), and a qualifier ('with its current state'). It does not explicitly differentiate from sibling tools like design_status, but 'every design job' signals a broad listing operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when an agent needs an overview of all design jobs, but it provides no explicit when-to-use guidance or distinction from sibling status-related tools. It leaves the choice between design_list and design_status to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_mistakesC

Report validation defects this design corrected or still carries.

ParametersJSON Schema
NameRequiredDescriptionDefault
design_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavior. 'Report' implies a read-only operation, but the description does not explicitly state that it is safe, whether it requires prior steps (e.g., a design must exist), or whether it might return stale data. The behavioral implications are only implicit, leaving room for misinterpretation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no fluff. It front-loads the key action ('Report validation defects') and the scope ('this design'). It earns high marks for efficiency, though the brevity borders on under-specification, which is why it is not a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the description is minimally adequate: it identifies the resource and the deliverable. However, it lacks context about the design lifecycle, what constitutes a 'validation defect', and how this tool relates to siblings like design_record_result or design_lesson_decide. The output schema covers return values, but the description does not fully equip an agent to call it correctly in a broader workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description should compensate by explaining design_id. It only alludes to 'this design' without clarifying how to obtain or format the identifier, what values are valid, or any relationship to other design tools. The parameter is left as an opaque string, providing minimal additional meaning beyond the schema's type information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'report' and resource 'validation defects' are specific, and the phrase 'this design corrected or still carries' indicates the scope is tied to a particular design. It is clearly distinct from sibling tools like design_status or design_list, which focus on status or listing rather than defect reporting. A minor ambiguity is what 'validation defects' exactly encompasses, but the core purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives. It does not mention prerequisites, typical scenarios, or any exclusions (e.g., 'use design_status for overall health'). An agent must rely on the tool name and vague wording to decide when to call it, which is insufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_record_resultB

Bind the exact FCStd hash to passed JSON, Markdown, and PNG evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
design_idYes
model_pathYes
evidence_paths_jsonYes
validation_report_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It implies a write operation ('Bind') but does not state consequences, reversibility, permission requirements, or side effects. It also fails to disclose what happens to existing records or whether the operation is idempotent. This leaves significant uncertainty for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that leads with the core action and resource. There is no redundant phrasing, and it is appropriately sized for the tool's function. Every word contributes to understanding the primary purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description omits critical operational context: what inputs are expected (parameter meanings), how they interrelate, and any workflow prerequisites. Given the tool has four required parameters and no annotations, this is insufficient. An agent cannot confidently construct a valid call without further elaboration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It only mentions 'passed JSON, Markdown, and PNG evidence,' which hints at the evidence_paths_json parameter but does not map it explicitly. It completely omits design_id, model_path, and validation_report_path, leaving their roles unexplained. The description adds minimal value beyond the schema's bare property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Bind' and clearly states the action: linking an exact FCStd hash to evidence files (JSON, Markdown, PNG). It specifies the resource and the operation, and it is distinct from sibling tools that deal with design lifecycle events rather than result recording. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any preconditions, sequencing with other design tools, or exclusions. An agent has to infer that this is used after validation produces evidence, but no explicit context is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_startC

Start or resume a design after natural-language direction approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
design_idYes
source_pathNo
approval_textYes
proposal_summaryYes
requirements_jsonYes
model_classificationYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. 'Start or resume' implies a state-changing operation, but it does not disclose side effects, required permissions, what happens to an existing design on resume, whether validation occurs, or how approval_text is used. This is a significant transparency gap for a mutation-like tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no filler words; the main action and condition are front-loaded. It loses a point only because it is so terse that it omits information the agent needs, but judged purely on conciseness and structure, it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with 7 parameters, 6 required, no annotations, and 0% schema coverage, yet the description provides almost no operational context. While an output schema exists, it does not offset the absence of parameter semantics, behavioral side effects, or workflow guidance. The description is far from sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 7 parameters, yet the description adds no parameter-level meaning. It only alludes to 'natural-language direction approval,' which loosely maps to approval_text, but design_id, title, model_classification, requirements_json, proposal_summary, and source_path are entirely unexplained. The description fails to compensate for the absent schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Start or resume a design') and the triggering condition ('after natural-language direction approval'). It is a specific verb+resource statement, though it does not explicitly contrast itself with sibling tools like design_confirm or design_status, so it falls just short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after natural-language direction approval' gives an implied usage context: the tool should be invoked once approval has been obtained. However, there is no explicit guidance about when not to use it, when to use design_confirm instead, or what prerequisites must be satisfied (e.g., whether design_id must already exist for resume).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_statusC

Read current model, validation, confirmation, and lesson state.

ParametersJSON Schema
NameRequiredDescriptionDefault
design_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. 'Read current... state' clearly signals a non-mutating snapshot operation, which is helpful. However, it does not disclose freshness, authentication needs, error behavior, or whether the design must be active, though for a simple status read this is a moderate gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, front-loaded with the operation verb 'Read'. Every word earns its place, and it is appropriately sized for a simple status tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists and there is only one required parameter, so the baseline completeness is decent. However, the description fails to explain design_id or clarify when to choose design_status over design_system_status, leaving a clear gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions design_id or how to obtain it. The required parameter is completely undocumented in both the schema and the description, so the agent gets no guidance on its meaning, format, or source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and names concrete state domains: model, validation, confirmation, and lesson state. This makes the tool's purpose clear and distinguishes it from mutation-focused siblings like design_start, design_confirm, and design_record_result. It does not explicitly distinguish itself from design_system_status, but the listed state categories provide enough specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use design_status versus related tools such as design_system_status or design_list. No prerequisites, exclusions, or alternative-selection rules are provided, leaving the agent to infer the right context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_system_statusB

Inspect local design, FreeCAD, knowledge, and provider readiness.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavior. It implies a read-only inspection via 'inspect', but it does not state whether external calls are made (e.g., for provider readiness), what local data is accessed, whether any side effects occur, or what permissions are required. This is minimal behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of nine words, front-loading the verb and the key domains with no filler. It is maximally concise while still conveying the purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description is largely sufficient, but it does not address the relationship to the similarly named 'design_status' or clarify whether the readiness check is purely local or involves external services. This leaves some ambiguity for an agent deciding between sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the description correctly avoids inventing any. Per the rubric, a schema with no parameters earns a baseline of 4, and the empty input schema already covers this fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'inspect' and names the resource domains ('local design, FreeCAD, knowledge, and provider readiness'), clearly conveying a status/health-check purpose. It does not explicitly distinguish itself from the sibling 'design_status', so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'design_status' or 'standard_part_sources_status'. The intended invocation context is only implied by the name and domain list, with no explicit conditions, exclusions, or routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

standard_part_download_registerC

Register a validated downloaded standard part with full provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
standardYes
file_pathYes
source_urlYes
part_numberYes
provider_idYes
nominal_sizeYes
metadata_jsonNo{}
validation_report_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose side effects, authorization needs, or data requirements. It only says 'register' (implying a write) and 'full provenance,' but doesn't say whether registration is idempotent, overwrites existing parts, or enforces validation requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence is economical and front-loaded, but for an 8-parameter registration tool it is under-specified rather than appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even with an output schema, the agent lacks guidance on required parameter semantics, registration behavior, and prerequisites. The description conveys only the high-level intent, leaving too much for the agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no parameter-level meaning. The phrase 'full provenance' only vaguely gestures at source URL/validation metadata without mapping to any of the eight parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('register') with a well-defined object ('validated downloaded standard part') and a scope qualifier ('with full provenance'). It is clearly distinct from sibling tools, which query providers or source status rather than creating a registration record.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when this tool should be used versus alternatives, and no exclusions. The only implied context is that the part must already be validated and downloaded, but the description does not explain prerequisites or workflow placement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

standard_part_providers_getC

List standard-part sources in configured trust order.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description indicates this is a listing/read operationabbvie and implies no mutation, which is behaviorally useful. However, with no annotations, it does not disclose whether any authentication, trust-order configuration, or side-effect behaviors apply; it stays at a basic level.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One clear sentence that immediately conveys the action, resource, and ordering constraint. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and has an output schema, so the missing return-value documentation is acceptable. However, the complete silence on the optional 'category' parameter and the lack of sibling differentiation leave gaps that could lead to an incorrect invocation or tool choice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention the 'category' parameter at all. The agent cannot determine whether category filters, sorts, or otherwise affects the list of standard-part sources.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and clearly identifies the resource ('standard-part sources') and a distinguishing characteristic ('in configured trust order'). It does not explicitly differentiate itself from its similar sibling standard_part_sources_status, so the uniqueness is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus standard_part_sources_status or standard_part_download_register. The phrase 'configured trust order' hints at context, but no conditions, exclusions, or alternatives are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

standard_part_sources_statusB

Inspect the configured standard-part catalog binding.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. 'Inspect' implies a read-only operation, but the description does not disclose potential error states, whether the binding can be unconfigured, or what kind of data the status check returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundant phrasing. Every word adds semantic value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, has no parameters, and an output schema is present, so the description does not need to explain return values. The only slight gap is not defining what 'standard-part catalog binding' means, but for a low-complexity status tool this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so the schema fully covers this dimension. The description does not need to explain parameter semantics, and the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('inspect') and names a concrete resource ('configured standard-part catalog binding'), making the tool's purpose clear. It reads as a status/read operation distinct from sibling tools like standard_part_providers_get or standard_part_download_register, though it does not explicitly point out the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus its siblings. There is no mention of alternatives, exclusions, or the larger workflow context in which the binding status matters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.11.0
    • First observeddesign_confirm
    • First observeddesign_gear_size
    • First observeddesign_knowledge_retrieve
    • First observeddesign_lesson_decide
    • First observeddesign_list
    • First observeddesign_mistakes
    • First observeddesign_record_result
    • First observeddesign_start
    • First observeddesign_status
    • First observeddesign_system_status
    • First observedstandard_part_download_register
    • First observedstandard_part_providers_get
    • First observedstandard_part_sources_status

TDQS

B3.2/5.0

Scored across 13 tools

Disambiguation5/5

Each tool targets a distinct action or resource: readiness checks, lifecycle management, knowledge retrieval, evidence recording, gear sizing, and standard-part management are clearly separated. Even closely related tools like design_confirm and design_lesson_decide have non-overlapping responsibilities (model confirmation vs. lesson publication).

Naming Consistency4/5

All tools follow a consistent domain prefix (design_ or standard_part_), but the suffix pattern mixes verb-first (design_start, design_confirm) with noun-first (design_status, design_mistakes) forms. The readability is high and the prefix convention provides strong predictability, so the minor deviations are not confusing.

Tool Count5/5

13 tools is well within the ideal 3-15 range and matches the server's broad but focused scope—covering design workflow, knowledge, validation, and standard-part integration. Each tool offers a distinct function and none feel redundant or extraneous.

Completeness4/5

The tool surface covers the core design lifecycle (start, status, list, confirm, record, mistakes) and standard-part management (providers, sources, register). Minor gaps exist, such as no explicit design update/delete or a standard-part search/download tool, but the existing tools allow agents to work around these via design_start/confirm and the provider/source inspection tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables LLMs to safely generate parametric CAD parts (STEP/STL) using verified templates and FreeCAD, with validation and assembly support.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to convert natural language engineering prompts into editable parametric CAD models with deterministic parsing, validation, and edit support.
    6
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables engineering agents to perform parametric CAD operations through a controlled MCP gateway built on FreeCAD, with transactions, diagnostics, and reproducible verification.
    2
    GNU Lesser General Public v2.1 only