Skip to main content
Glama

module_observe

Capture and record reproducible observable evidence from executed commands or system states for Vibe Engineering MCP, including stdout, stderr, exit codes, timing, and artifact hashes.

Instructions

Captures and records reproducible observable evidence from an executed command or system state

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
stderrYesStandard error captured
stdoutYesStandard output captured
taskIdNoAssociated task ID
commandYesThe command or harness that was observed
exitCodeYesExit code resulting from execution
moduleIdNoAssociated module ID
evidenceTypeNoObservation modality (CLI, HTTP_API, DATABASE, etc.)
artifactPathsNoFile paths of generated artifacts to hash and verify
executionTimeMsYesDuration of execution in milliseconds

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says the tool records evidence, which implies a write/mutation operation, but it does not disclose permissions, persistence behavior, idempotency, whether existing records are overwritten, error behavior, or rate limits. The behavioral profile is therefore substantially incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundant text. It is appropriately sized, though it does not provide structural cues for complex usage or sequencing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given nine parameters and no output schema, the description is minimally adequate: it conveys the core action but omits return values, storage semantics, and post-capture behavior. With no annotations and no output schema, more context would be expected for a recording/mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all nine parameters, including evidenceType and artifactPaths. The description adds no parameter-level meaning beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Captures and records reproducible observable evidence from an executed command or system state.' It distinguishes the tool from execution-focused siblings like module_run and terminal_exec by focusing on evidence recording rather than execution. It does not explicitly name alternatives, so a 4 is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used after a command has executed or when recording system state, but it gives no explicit when-to-use guidance, prerequisites, or alternatives. It does not clarify when to choose module_observe over contract_verify, integration_verify, or checkpoint_create.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.