Skip to main content
Glama

osnova_tests

Read-onlyIdempotent

Identify tests affected by a symbol or the dependencies a test file calls, so you know what to run before editing.

Instructions

Tests: given symbols, the indexed test files for each one in two separate tiers with separate counts: files with a resolved call or reference edge to the symbol (exact file:line and resolution basis), then files that only import the symbol's file and contain no indexed call or reference to it. An empty resolved tier is stated on its own line; import-only files are leads, not tests of the symbol. Given one test file, the non-test symbols it calls and the files it imports. Exactly one of symbols or file. Use before editing to find the tests to run. No indexed test is not proof of no test, and a listed test is not coverage.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fileNoRepo-relative test file whose symbols under test to list
limitNoMaximum test files per symbol (default 20) or symbols per file (default 50)
symbolsNoSymbol names or qualified names (file#Class.method) to find tests for
includeImportOnlyNoWith symbols: also list test files that only import the symbol's file (default true); false lists only files with a resolved edge

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.12.0

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, etc., so safety and idempotency are covered. The description adds valuable context beyond annotations: it explains the two-tier result structure, that an empty resolved tier is stated on its own line, that import-only files are leads not tests, and provides an important caveat about the limitations ('No indexed test is not proof of no test, and a listed test is not coverage'). It doesn't discuss rate limits or auth, but none are relevant here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that packs a lot of information, but it's not front-loaded with a clear summary sentence. The purpose is buried after listing tier details. It could be better structured with a leading sentence stating the core function, then details. However, it is not overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (dual-mode, two-tier results) and the lack of an output schema, the description does a good job covering return semantics and caveats. It could be improved by explaining what happens when both parameters are omitted (since required parameters is 0), but it largely addresses what an agent needs to know.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters thoroughly. The description adds a bit of meaning by clarifying the mutual exclusivity ('Exactly one of symbols or file') and the tier interpretation, but it doesn't add new syntax or format details for parameters. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: finding indexed test files for given symbols across two tiers, and conversely listing symbols under test for a given file. It specifies the exact dual-mode behavior precisely. It does not distinguish this tool from its siblings (e.g., doesn't say why to use this over osnova_unreferenced or osnova_ground).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: 'Use before editing to find the tests to run.' It also specifies the mutual exclusivity constraint: 'Exactly one of symbols or file.' However, it doesn't compare to sibling tools or state when NOT to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.