Skip to main content
Glama

TypedDecisionMCP

Repositório: hectorlutero/TypedDecisionMCP.

MCP de uma tool (decide, formerly decidir) para o Cursor. Motor local default: head-mlp + MiniLM (embedding curto + probe). Opt-in: DECIDE_ENGINE=logits + Qwen. Sem TypeSafe. Sem JSON gerado no caminho de 10×.

Alvo v1 em PLAN.md: ≥ 10× menos tokens gerados e ≥ 10× menos tempo no hop de decisão, com accuracy ≥ 0,75 no ouro authored. Implementação do alvo seguinte (authored 0,85 e 15×) em PLAN-melhorias.md.

Uso

npm install
npm test
npm run fetch-model
npm run build

examples/cursor.mcp.json aponta para dist/index.js. A rule está em cursor/rule.mdc.

npm run bench:logits
npm run bench:report

bench:report separa qualidade de 10×. Qualidade: authored ≥ 0,75 e generated_tokens === 0. 10× tokens exige bench/baseline.json com source: measured e method: cursor-ui (Composer) ou cursor-subagent (DECIDE_ACCEPT_PROXY=1, text em todos os hops). Os hops são cmd-01 sub-01 diff-02 file-01 commit-01. Prompts em bench/hop-prompts.md. Mesma régua chars/4 nos dois lados quando há text. 10× de tempo só com relógio da UI ou spawn de três hops ok.

Ship actual nesta VM (head-mlp + all-MiniLM-L6-v2 Q8): authored 40/40, held-out 19/20, generated_tokens === 0, p50 ~16 ms, 0 destructive false auto (miss só h-diff-01). Baseline logits (Qwen3-0.6B Q8) no mesmo commit: authored 33/40, held-out 15/20, p50 ~381 ms. Tempo 10× tipicamente 3–4/5 (UI cursor-ui). Aceite dual: held-out ≥ max(0,75, logits) e p50 ≤ 200 ms ou ≤ logits — MiniLM passa. Detalhe: models/README.md, snapshot bench/out/quality-minilm-ship.md.

Related MCP server: ContextBridge

Presets

command · subagent · diff · file · commit · bundle

file e bundle exigem state.candidates (2–20 caminhos). Portuguese aliases still parse: comando, subagente, ficheiro, pacote.

Available Tools

1 tool
decidirC

Typed local decision over state. Presets: comando, subagente, diff, ficheiro, commit, pacote. Returns yesno/choice/score plus action auto|review|stop. Do not invent probabilities in text.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
presetNo
questionsNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose that the decision is 'local', specifies the return contract (yesno/choice/score plus action auto|review|stop), and adds a meaningful constraint: 'Do not invent probabilities in text.' It omits details like side effects or failure modes, but the local qualifier and output contract give useful transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is compact and front-loaded: purpose, presets, output contract, and a constraint in four short clauses. There is no filler or redundancy. The closing probability instruction earns its place because it is a rule for correct use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is moderately complex (union-typed state, nested questions object, preset enum) and has no output schema, no parameter descriptions, and no sibling context. The description leaves key invocation details implicit, such as the semantics of state, how questions are structured, and what the action values mean. It is minimally usable but not complete enough for confident call construction.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only repeats the preset values that already exist in the enum. It does not define the required 'state' parameter, explain what 'questions' means, or clarify how presets affect behavior. The preset names and action values add a little context, but not enough to compensate for the undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource ('state') and the output shape (yesno/choice/score plus action), and lists presets, so it is more than a tautology. However, it lacks a concrete verb: 'decision' essentially restates the tool name, and what a 'typed local decision' actually does is left fuzzy without sibling tools for contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool, when to prefer a preset versus questions, or what conditions should route to it. Even with no sibling tools, the description could describe intended scenarios but does not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observeddecidir

TDQS

C2.9/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion or misselection between tools. The single tool 'decidir' is unique, so disambiguation is trivially perfect.

Naming Consistency4/5

The tool name 'decidir' is a clear, single verb, and since there is only one tool, there is no naming pattern to violate. It is slightly inconsistent with typical verb_noun conventions, but internally consistent and readable.

Tool Count2/5

The server has only one tool that handles a wide range of presets (comando, subagente, diff, etc.), making it a monolith. This is too few tools for the apparent scope, as an agent would need to parse presets and cannot use distinct tools for different decision types.

Completeness2/5

The single tool covers only a generic 'decide' operation with presets, but lacks any other operations like listing available presets, getting metadata, or handling workflow-specific actions. There are significant gaps that would limit an agent's ability to perform complex tasks without workarounds.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers