Seshat
OfficialAllows Seshat to access and analyze repositories from GitHub, including public repos cloned automatically and private repos via connecting a GitHub account, to build structural code intelligence and history insights.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Seshatwhat breaks if I change the process_payment function?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Seshat — structural code intelligence for AI agents
Your agent reconstructs your codebase from likelihood. Seshat compiles it.
Seshat turns a repository into a typed symbol graph (every function, class, route, and table, with its real dependency edges, data flow, and constraints) and serves it to your agent over MCP. So before your agent edits a function, it can ask what will actually break instead of guessing.
Backed by a compiled intermediate representation, not text search or embeddings: if Seshat says a function has three callers, it has exactly three.
Why
AI coding agents are fast but structurally blind. When an agent changes one function, it often cannot see everything that depends on it, so it silently breaks callers it never looked at. A large share of AI-introduced regressions come from exactly this. Seshat gives the agent the map.
Related MCP server: io.github.pmgarg/cgraphy
Try it first (no install, no login)
Paste any public repo at https://seshat.papyruslabs.ai/try and watch it trace what a change would break.
Install
npx -y @papyruslabsai/seshat-mcp setup <your-key>Get a free key at https://seshat.papyruslabs.ai (first extraction is free). The setup command writes your MCP config and stores the key.
Or configure manually (Claude Code, Cursor, or any MCP client):
{
"mcpServers": {
"seshat": {
"command": "npx",
"args": ["-y", "@papyruslabsai/seshat-mcp"],
"env": { "SESHAT_API_KEY": "your-key" }
}
}
}Tools
Point your agent at a repo with sync_project, then it investigates the way a senior
engineer does: orient, trace, verify.
Orient
list_projects— what is syncedlist_modules— how the codebase is organized, by layer or modulequery_entities— find functions, classes, and routes by name, layer, or modulefind_entry_points— routes, exports, and the public API surface
Investigate a symbol
get_entity— signature, callers, callees, data flow, side effects, and tables touchedget_dependencies— the real call chain, callers and calleesget_blast_radius— everything that breaks if you change it, transitivelyget_data_flow— what a function reads, returns, and mutatesget_optimal_context— the minimal, ranked set of files to read before editingfind_by_constraint— every function that touches a given table (or carries a given trait)find_dead_code— unreachable symbols, safe to delete
Read the history (from the repo's commit record, backfilled on first sync)
get_lineage— how one function has actually changed: each commit typed by what moved (body, calls, data, signature), CI pass/fail and reverts, what changes alongside it, and what last forced a change here. Ask it before touching anything load-bearing.get_hotspots— where development happens and where it fails: the most-changed code, thrash spots where changes keep getting reverted or landing on red CI, and heavily used code nobody has touched (stability pressure)get_co_change_clusters— the hidden modules: code that changes together across files even when no import connects it, so a change to one member usually means the rest
Every answer comes from the compiled graph and discloses the coverage behind it. History is commit-resolution correlation and says so; it never claims causation it can't show.
Cross-cutting audit tools (test coverage, topology, semantic clones) are being hardened and will be added to this list as they land.
Privacy
Analysis runs in the Papyrus Labs cloud, by design: the compiled graph is the product's moat, and keeping extraction server-side is how that stays protected. Public repos are cloned from GitHub; private repos require you to connect your GitHub account. Source is processed to build the graph and cached to serve queries. See the privacy policy at seshat.papyruslabs.ai.
Pricing
First extraction is free. $0.03 per query after a free tier; a typical investigation is 5 to 15 queries. Details at https://seshat.papyruslabs.ai.
MIT licensed server. Named for the goddess who kept the records, built so your agent can read them.
Available Tools
27 toolsfind_by_constraintFind by ConstraintARead-onlyIdempotent
Find every function with a specific syntactic constraint tag — AUTH (requires authentication), DB_ACCESS (touches database), THROWS (explicit throw statement), PURE (no side effects), NETWORK_IO (makes HTTP calls), VALIDATED (has input validation). Also supports table-level queries: pass table="walks" to find every function that reads or writes the walks table (answers "what touches this table?" for schema migrations). Constraints are extracted from source syntax, not inferred. For semantic/behavioral properties (e.g., "can fail transitively"), use query_traits instead.
| Name | Required | Description | Default |
|---|---|---|---|
| table | No | Optional: filter to functions that touch a specific database table (e.g., "walks", "users"). Returns structured db_operations showing read/write/mutate per function. | |
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| constraint | Yes | Constraint tag to search for: AUTH, VALIDATED, PURE, THROWS, DB_ACCESS, NETWORK_IO, IMP, etc. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety profile is covered. The description adds a useful provenance fact – 'constraints are extracted from source syntax, not inferred' – which matters for trust in results, but does not disclose return format or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and tag list, then adds table-level mode, then the semantic alternative. Three sentences, each earning its place, though the tag enumeration is lengthy and somewhat schema-duplicative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters, no output schema, and rich annotations, the description covers both query modes, provenance, and the sibling routing clearly. It could note interaction between constraint and table parameters when used together, but the essential guidance is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are already documented. The description adds the constraint tag vocabulary and the table-mode semantics, which enriches meaning beyond raw schema strings, but the schema itself carries the parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find) and resource (functions) with the exact mechanism (syntactic constraint tags), enumerating the tag values so an agent knows what's queryable. It also distinguishes itself from query_traits by naming the semantic alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-to-use for two distinct modes (constraint-based and table-based) with a concrete example (table="walks"), and routes semantic queries to query_traits. Lacks explicit when-not guidance beyond the semantic case, but coverage is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_dead_codeFind Dead CodeARead-onlyIdempotent
Find functions that nothing calls. Walks the call graph from all entry points (routes, exports, tests) and flags symbols that are unreachable. Use this during cleanup or before a release to find safe deletion candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| include_tests | No | Include test entities in dead code results (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds meaningful context beyond the safety profile: it explains internals ('walks the call graph from all entry points') and the purpose ('safe deletion candidates'). It doesn't describe output format, but the description is still substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero waste; action, mechanism, and usage are front-loaded in that order. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-param analysis tool with no output schema and full annotation coverage, the description covers purpose, mechanism, and when to call. It does not explain return value shape or pagination, but output_schema is absent so that is less critical. Slightly incomplete on result format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are fully documented in the schema. The description adds no additional parameter details. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find functions that nothing calls') and elaborates the mechanism ('walks the call graph from all entry points'). Clearly distinguishable from siblings like find_entry_points or get_dependencies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear intended contexts ('during cleanup or before a release') for when to use it. However, no explicit when-not-to-use guidance or named alternatives (e.g. vs get_dependencies) is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_entry_pointsFind Entry PointsARead-onlyIdempotent
List the ways into the system: route/controller handlers, test entries, framework plugin registrations, and exported symbols (the public API surface), classified by kind and ranked by reach. Call it first when orienting in an unfamiliar codebase — it answers "where does execution start, and what is the public surface?" Complements list_modules (structure) and find_dead_code (its exact inverse: these are the reachability roots).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and openWorld. The description adds important context: the classification by kind, ranking by reach, and the conceptual role as reachability roots (the inverse of dead code). It doesn't cover return format or pagination, but the read-only safety profile is well established via annotations, so this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two information-dense sentences that are front-loaded with what the tool lists, followed by when to use it and how it relates to siblings. Every clause earns its place; no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-parameter tool with an output-free description, the definition is complete: it tells the agent what it returns (classified entry points), when to call it (orientation), and how it fits with other tools. No important behavioral or usage details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter (project), and schema coverage is 100%, so the schema already documents it adequately. The description doesn't add additional parameter semantics, but with a single well-documented parameter, the baseline is already high. No need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (List) and a clearly-defined resource (the ways into the system: route/controller handlers, test entries, framework plugin registrations, exported symbols), and explicitly characterizes it as the public API surface. It distinguishes itself from siblings by naming list_modules (structure) and find_dead_code (its inverse).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call it first when orienting in an unfamiliar codebase' and explains what question it answers. It names sibling tools and clarifies its complementary/inverse relationship to them, giving clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_error_gapsFind Error GapsARead-onlyIdempotent
Find crash risks: functions that throw or have network/DB side effects whose callers don't catch errors. Returns the specific caller→callee pairs where exceptions can propagate unhandled. Use this before shipping to find missing error handling.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds what the analysis detects and returns, but does not disclose any behavioral traits beyond that, such as performance characteristics or result limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with no wasted words. The core function and return value are front-loaded, and the usage recommendation is placed last as a natural call to action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single optional parameter, full schema coverage, and no output schema, the description does enough: it defines the risk being detected, the return shape, and when to use it. It could be slightly more complete by noting multi-project context, but that is adequately handled by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'project' parameter is fully documented in the schema, including multi-project mode and the list_projects reference. The description adds no additional parameter meaning, which is the baseline when the schema carries the parameter explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Find') and resource ('crash risks' / 'error gaps') and explains the precise detection condition: functions that throw or have network/DB side effects whose callers don't catch errors. It also specifies the return content (caller→callee pairs), making the tool's output purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this before shipping to find missing error handling,' giving clear situational context. However, it does not name alternative analysis tools or state when not to use it, so it lacks explicit alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_exposure_leaksFind Exposure LeaksARead-onlyIdempotent
Find places where public/API code directly accesses private internals, bypassing the intended abstraction boundary. Use this during API design reviews or before extracting a module.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safe, non-mutating nature is already covered. The description adds the scoping concept (abstraction boundary leaks) but doesn't disclose behavioral traits like how many results to expect, pagination, or what counts as a 'leak' (e.g., direct field access vs. method calls). With annotations carrying the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with a precise definition of the leak type, followed by the concrete usage context. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with one fully documented parameter and rich annotations, the description covers the what and the when well. It lacks an explicit tie to sibling tools (e.g., find_layer_violations) that an agent might confuse it with, but otherwise it's complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is only one parameter, so the baseline should be high. The description doesn't repeat the parameter, but the schema already explains 'project' with a pointer to list_projects. No additional semantics are needed or missing, so a 4 is warranted for adequately letting the schema do the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Find') and a precisely defined resource ('places where public/API code directly accesses private internals, bypassing the intended abstraction boundary'). This is a concrete architectural concern that distinguishes it from siblings like find_layer_violations or find_ownership_violations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use it: 'during API design reviews or before extracting a module.' This is a specific, actionable trigger. However, it doesn't explicitly name an alternative tool or state when-not to use it, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_layer_violationsFind Layer ViolationsARead-onlyIdempotent
Find places where the architecture is broken — a database repository calling a route handler, a utility importing a component. Returns every backward or skip-layer dependency that violates clean architecture.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is read-only, idempotent, and open-world, so the safety profile is covered. The description adds the 'returns every' completeness claim, but says nothing about scope, result size, ordering, or cost of a whole-architecture scan. With annotations carrying the safety burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first defines the violation type with concrete examples, the second states what is returned. Zero filler and the core concept is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-parameter analysis tool with no output schema, the description gives the agent enough to know what it does and what comes back. It stops short of noting result format or the multi-project requirement nuance, but the annotations and schema cover most of the rest.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and only one optional parameter exists, so the schema already documents 'project' fully (including the list_projects hint). The description adds no parameter detail, but with coverage this high and a single parameter, the baseline is satisfied and no gap remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Find places where the architecture is broken') and defines the target precisely — backward or skip-layer dependencies violating clean architecture. Sibling tools like find_ownership_violations and find_exposure_leaks are different violation classes, and the description makes clear this one is specifically about layer dependency direction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The examples ('a database repository calling a route handler, a utility importing a component') establish exactly when this tool applies. However, there is no explicit guidance on when NOT to use it or which sibling to reach for instead (e.g., find_ownership_violations), leaving the boundary to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_ownership_violationsFind Ownership ViolationsARead-onlyIdempotent
Find memory and lifecycle issues — entities with complex ownership, unsafe blocks, escaping references, or illegal mutability on borrowed data. Returns 0 for most JS/Python codebases — a non-zero result in those languages indicates a serious boundary violation worth investigating. Most detailed results for Rust and C++.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, open-world behavior. The description adds valuable behavioral context beyond annotations: expected result distribution by language and the interpretation that a non-zero result in JS/Python indicates a serious boundary violation worth investigating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then result behavior by language, then language-specific detail. No sentence restates the name or title, and each adds useful information without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey return expectations; it does so via the 0/non-zero result behavior and language detail. Annotations cover the safety profile, and the schema covers the only parameter, leaving only explicit output structure as a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single 'project' parameter is fully documented in the schema. The description adds no additional meaning about parameter usage or format beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find memory and lifecycle issues') and enumerates examples such as complex ownership, unsafe blocks, escaping references, and illegal mutability. However, it does not distinguish this from sibling analysis tools like find_exposure_leaks, find_layer_violations, or find_dead_code.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful language-context guidance: returns 0 for most JS/Python codebases and is most detailed for Rust/C++, which implies when results are meaningful. But it does not state when to choose this over sibling tools or any prerequisites beyond what the schema already says.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_semantic_clonesFind Semantic ClonesARead-onlyIdempotent
Find duplicated logic across the codebase. Normalizes variable names and compares code structure to catch identical algorithms in different files — even across different languages. Use this before a DRY refactor.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| min_complexity | No | Minimum logic expressions to count as a match (default: 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it normalizes variable names and compares structure across different languages, which is non-obvious detection behavior beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first defines the capability, the second gives detection details and a usage directive. Front-loaded with the core purpose, zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with full param coverage and no output schema, the description covers purpose, detection method, and when to use. Return format is unspecified, but that's a minor gap given the low-risk read operation and sibling patterns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents both parameters fully, including defaults and cross-tool reference ('Use list_projects'). The description does not add parameter syntax, but the baseline for full coverage is 3; the description's mention of cross-language comparison indirectly clarifies why min_complexity matters. Slightly above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find duplicated logic'), and uniquely distinguishes itself from siblings like get_co_change_clusters or find_dead_code by explaining the semantic normalization approach. An agent can immediately tell what this does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use context: 'Use this before a DRY refactor.' This gives a concrete scenario for invocation. It doesn't name alternatives or state when-not to use, but the pre-refactor timing is actionable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_statusAccount StatusARead-onlyIdempotent
See your current plan, available tools, and credit balance. Call this if a tool returns a tier error or you want to know what tools are available.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare read-only, open-world, idempotent, and non-destructive. The description adds that it reveals plan, tools, and credits, but does not elaborate on whether the information is cached, real-time, or subject to rate limits. With annotations covering safety, the added context is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with what it does, then the when-to-use condition. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters, annotations covering safety, and no output schema, the description adequately tells the agent what it returns and when to call it. It could mention that this is a lightweight diagnostic tool, but it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters, so baseline is 4. The description correctly implies no input is needed, matching the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns the current plan, available tools, and credit balance; a specific verb+resource. This is distinct from siblings like get_auth_matrix or list_modules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: 'Call this if a tool returns a tier error or you want to know what tools are available.' This provides a clear trigger condition, though no explicit alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_auth_matrixGet Auth MatrixARead-onlyIdempotent
Audit authentication coverage. Shows which API routes and controllers require auth and which don't, plus inconsistencies like database access without auth checks. For large codebases, use the module parameter to drill into a specific module/directory. Most useful for backend codebases with middleware-based auth. Frontend frameworks (React, Vue) handle auth via component wrappers, which this tool won't detect as AUTH constraints.
| Name | Required | Description | Default |
|---|---|---|---|
| layer | No | Filter to a specific layer: "route" or "controller" (optional). | |
| module | No | Filter to routes/controllers in a specific module or directory path (optional). | |
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and destructiveHint=false, so the agent knows this is a safe read. The description adds important behavioral context beyond annotations: it explains the tool's detection limits (frontend auth wrappers won't be detected) and the module-drilldown behavior for large codebases. It does not cover return format (no output schema), but this is acceptable given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: first states the purpose, second describes output scope, third gives a scaling tip, fourth warns about applicability limits. Front-loaded with the core purpose and no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately conveys what the tool does, its scope, its limitations, and how to use its parameters. It also provides behavioral context (detection gaps, module drilldown) that compensates for the lack of output schema. Nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters clearly, including enum-like values for 'layer' and multi-project mode for 'project'. The description does not dwell on parameter syntax but does explain the practical use of the 'module' parameter ('drill into a specific module/directory'), adding meaningful context beyond the schema's generic description. Baseline 3 is exceeded slightly, though it could more explicitly mention when to use 'layer' or 'project'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Audit authentication coverage.' It then states exactly what the tool shows (routes/controllers requiring auth, inconsistencies like DB access without auth checks), which clearly distinguishes it from siblings like get_account_status or find_exposure_leaks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: most useful for backend codebases with middleware-based auth, and names an alternative (module parameter) for large codebases. It also states a clear exclusion: frontend frameworks (React, Vue) handle auth via component wrappers, which this tool won't detect. This is textbook when/when-not/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_blast_radiusGet Blast RadiusARead-onlyIdempotent
Before modifying a function, call this to see everything that could break. Returns all transitively affected symbols — both upstream callers and downstream callees — with distance from the change point. Like git log --follow but for runtime impact. Designed for repeated use: as you discover new symbols with other tools, call this again on them to expand your understanding of the affected surface.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| entity_ids | Yes | Array of entity IDs or names to compute blast radius for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds behavioral context beyond annotations: it explains the transitive scope, the distance metric returned, and the intended iterative expansion pattern. It does not mention performance or rate limits, but the annotation coverage lowers the burden and the added usage pattern is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the trigger condition ('Before modifying a function'), then the return scope, then an analogy, then iterative usage. Four sentences, each carrying weight, with no fluff. Slightly longer than minimal but justified by the complex transitive semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers when to call, what it returns conceptually, and how to use iteratively. The lack of an output schema means the description should ideally explain the return format more, but it does state 'distance from the change point' and categories of affected symbols. Complete enough for correct invocation, with minor room to specify output shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are fully documented in the schema. The description explains that the tool computes blast radius for the given entity IDs but does not add format details, batching guidance, or constraints beyond what the schema provides. Baseline 3 is correct when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get blast radius) and clearly defines the resource as all transitively affected symbols, distinguishing both upstream callers and downstream callees. The analogy 'like git log --follow but for runtime impact' makes the scope concrete and differentiates it from siblings like get_dependencies or get_lineage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Before modifying a function, call this' and 'Designed for repeated use: as you discover new symbols with other tools, call this again.' This gives both the prime trigger condition and iterative usage guidance, which is rare and valuable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_co_change_clustersGet Co-Change ClustersARead-onlyIdempotent
The codebase's hidden modules: groups of symbols that historically change together (≥3 shared commits, sweep commits excluded), computed from real commit history. Clusters that span multiple directories reveal coupling the file tree doesn't show. Call it when planning a change or splitting work across agents — touching one member of a cluster usually means touching the rest, even when no static dependency connects them. Correlation evidence, honestly framed; complements get_blast_radius (static reach) with empirical reach.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, and non-destructive, so safety is covered. The description adds valuable behavioral context the annotations cannot: the ≥3-commit threshold, sweep commit exclusion, and the caveat that clusters are correlation evidence, not static dependency. It doesn't mention performance characteristics or result ordering, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, each earning its place: definition with thresholds, architectural insight, when-to-use call, and relationship to a sibling. Front-loaded with what the tool actually returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does well by explaining what clusters represent and their empirical basis, though it doesn't describe the result shape (e.g., cluster membership fields). The parameter and safety burdens are fully covered by schema and annotations, so the remaining gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter's description already covers multi-project mode and points to list_projects. The description adds no parameter detail, so the schema fully carries the burden; given zero required parameters and complete coverage, this is appropriately high but not exceptional.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (groups of symbols that historically change together) with the exact algorithmic basis (≥3 shared commits, sweep commits excluded). It goes on to name the sibling it complements (get_blast_radius) and contrasts static vs empirical reach, so the agent can tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to call it: 'when planning a change or splitting work across agents,' plus the consequence of touching one cluster member. It also positions it against the alternative by naming get_blast_radius and contrasting their types of reach.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_coupling_metricsGet Coupling MetricsARead-onlyIdempotent
Measure how tangled your code is. Returns coupling (cross-boundary dependencies), cohesion (within-group dependencies), and instability scores. High coupling + low cohesion = refactoring candidates. Start with group_by: "layer" for the architectural health view ("are my controllers more coupled than my services?"), then drill into group_by: "module" for specific hotspots.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| group_by | No | Group entities by module or layer (default: module) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds analytical context (the coupling/cohesion heuristic, the drill-down sequence) but says nothing about computation cost, result size, or freshness. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with what is measured, then the interpretation heuristic, then the recommended call sequence. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, read-only analytical tool with full schema coverage and annotations, the description covers purpose, output semantics, and usage sequencing well. It could be more explicit that the tool returns aggregates rather than per-entity lists, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the enum values for group_by and the project parameter are already documented. The description reinforces the practical use of each enum value ("layer" for architecture, "module" for hotspots), which is useful for choosing between them but adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource ("Measure... coupling metrics") and enumerates exactly what is returned: coupling, cohesion, and instability scores. This is distinct from siblings like get_dependencies or get_topology, which return structural graphs rather than quantified metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It prescribes a concrete workflow: start with group_by: "layer" for architectural health, then drill into group_by: "module" for hotspots. This tells the agent not just when to use the tool, but how to sequence calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_data_flowGet Data FlowARead-onlyIdempotent
See what data a function reads, returns, and mutates (DB writes, state changes). Use this when debugging data bugs or when you need to verify whether a function has side effects before refactoring it.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| entity_id | Yes | Entity ID or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a read-only, idempotent, non-destructive call, so the safety profile is covered. The description usefully clarifies that despite being a read, the report surfaces mutating behavior such as DB writes and state changes. It says nothing about scope of analysis, cost, or result size, so it adds only moderate value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first defines what the tool reveals, the second gives when to reach for it. The purpose is front-loaded and no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only two well-documented parameters and no output schema, the description needs to carry purpose and usage, which it does. It is nearly complete; the missing piece is how this relates to closely related siblings such as trace_data_path or get_dependencies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself documents project (including the list_projects hint) and entity_id. The description adds no parameter-level meaning of its own, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it reveals what data a function reads, returns, and mutates (DB writes, state changes). That is concrete and well beyond a restatement of the name. It does not differentiate itself from the sibling 'trace_data_path', which an agent could easily confuse with this tool, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear triggering conditions: debugging data bugs, or verifying side effects before refactoring. That is genuinely actionable context. It names no alternatives or exclusions (e.g., when to use trace_data_path or get_dependencies instead), so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dependenciesGet DependenciesARead-onlyIdempotent
Trace who calls a function and what it calls, up to N levels deep. Use this instead of grep-for-function-name when you need the actual call chain — returns the dependency graph, not text matches. Covers callers (upstream), callees (downstream), or both.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | How many levels deep to traverse (default: 2) | |
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| direction | No | Which direction to traverse (default: both) | |
| entity_id | Yes | Entity ID or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and open-world, so the safety profile is covered. The description adds value by clarifying the return is a dependency graph rather than text matches and that traversal can go either direction, though it says nothing about result size, truncation, or depth limits in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: capability, distinction from the alternative, and the direction options. The core purpose is front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the necessary work of indicating what comes back (a dependency graph vs text matches) and the traversal directions. It is nearly complete for a read-only traversal tool, with only result-size/cost behavior missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so depth, direction, project, and entity_id are already documented in the schema. The description's gloss on depth ('up to N levels') and direction ('callers, callees, or both') restates the enum and property docs rather than adding format or defaulting details. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (trace) and resource (who calls / what it calls a function), plus the depth scope. It explicitly contrasts with text matching, so an agent can distinguish it from a grep-style search without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (grep-for-function-name) and the condition that selects this tool over it ('when you need the actual call chain'). However, it does not differentiate from closely related siblings like get_blast_radius, get_lineage, or get_data_flow, which an agent could easily confuse with dependency traversal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_entityGet Entity DetailsARead-onlyIdempotent
Get everything about one function or class — its signature, callers, callees, data flow, constraints, source location, and database operations (which tables it reads/writes). Use this when you need to deeply understand a single symbol before modifying it. Returns more than reading the source file because it includes the dependency context.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity ID or name | |
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds valuable output context—signature, callers, callees, data flow, constraints, source location, and database read/write operations—but does not disclose pagination, size limits, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core operation and followed by a clear use case and a differentiator against reading source. No wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, 2-parameter tool with no output schema, the description plus schema and annotations give an agent enough to call it correctly: what it returns, when to use it, and that it is safe. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear descriptions for both id and project. The description adds no parameter-level meaning beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: getting everything about one function or class, and enumerates the returned facets. Very clear what the tool does, but it does not explicitly name or compare against sibling tools like query_entities or get_dependencies, so sibling differentiation is inferred from scope rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case: when you need to deeply understand a single symbol before modifying it. Clear context for use, but no when-not conditions or alternative tool suggestions (e.g., when to prefer get_dependencies or get_data_flow).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_hotspotsGet HotspotsARead-onlyIdempotent
Project-level change-history orientation: the most-changed entities (and what kind of change dominates each), low-survival thrash spots where changes don't stick, heavily-depended-on entities that haven't changed all window (interface freeze), and directories with no recent changes. Call it after list_modules when orienting in a codebase — it answers "where does development actually happen, and where does it fail?" from commit history rather than current structure. Complements get_lineage (one entity's story) with the project-wide map.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety and mutability profile is covered. The description adds context about relying on commit history rather than current structure, which is useful behavioral framing. But it doesn't disclose data window semantics, whether project is required in single-project mode, or result size/performance characteristics beyond what's in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and scope, then lists the four output categories compactly and finishes with a usage note and differentiation from get_lineage. Slightly dense but no filler; each sentence adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional parameter, full schema coverage, and no output schema, the description adequately tells the agent what the tool returns and when to call it. It lacks explicit return format details, which is a minor gap for an analytical tool where the agent may need to know what the output looks like for downstream processing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds no parameter syntax details. However, this is a zero-required-parameter tool where the schema itself carries semantics. Baseline 4 applies since no params = full schema coverage and description doesn't need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific analytical verb+resource ('Project-level change-history orientation') and enumerates the four outputs: most-changed entities, low-survival thrash spots, interface freeze, no-recent-change directories. It explicitly distinguishes from the sibling get_lineage ('one entity's story' vs 'project-wide map'), so an agent can route correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear usage trigger: 'Call it after list_modules when orienting in a codebase.' It also clarifies scope versus get_lineage. However, it doesn't state when NOT to use it or mention any alternatives like get_co_change_clusters or get_coupling_metrics, which could also be relevant for change-history questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lineageGet LineageARead-onlyIdempotent
The change history of one symbol, typed by what kind of change each commit made (body, calls, data, signature, constraints…), with CI verdicts, reverts, rename tracking, co-change partners, and rejected PRs that touched it. Call before modifying anything load-bearing: it answers "how does this entity usually change, and what happened last time someone tried?" — which no text diff can. Complements get_blast_radius (current impact) with history (past behavior).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| entity_id | Yes | Entity ID or name to fetch change history for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description goes further by disclosing what the result actually contains — typed change categories, reverts, rename tracking, rejected PRs — which materially shapes expectations. It omits pagination/volume behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the resource and its payload before explaining why to call it. The long enumeration in sentence one is dense but earns its place by defining scope; only minor trimming is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so well by naming the categories of history returned. Combined with annotations covering safety, an agent has enough to call correctly; only result-shape details like ordering or limits are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so the schema already documents entity_id and the multi-project 'project' flag with its list_projects hint. The description adds no parameter syntax or format detail beyond that, which makes 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('The change history of one symbol') and enumerates the distinguishing content: change-type typing, CI verdicts, reverts, rename tracking, co-change partners, rejected PRs. It explicitly contrasts itself with get_blast_radius, so an agent can tell the two siblings apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Call before modifying anything load-bearing'), an explicit framing question it answers, and names the alternative (get_blast_radius) with the axis of difference (current impact vs. past behavior). Nothing about when-to-use is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_optimal_contextGet Optimal ContextARead-onlyIdempotent
Before working on a function, call this to get the most relevant related code ranked by importance and fitted to a token budget. Returns a prioritized reading list of symbols you should understand — better than guessing which files to open. Designed for iterative use: call it on your target, read the top results, then call it again on any surprising dependencies to build a complete picture.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| strategy | No | Traversal strategy: bfs (faster, local neighborhood) or blast_radius (full affected set) | |
| max_tokens | No | Token budget for the context window (default: 8000) | |
| target_entity | Yes | Entity ID or name to build context around |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, covering safety and predictability. The description adds that results are 'ranked by importance' and 'fitted to a token budget,' but does not disclose tie-breaking, edge cases, or rate limits. With annotations doing the heavy lifting on behavior, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and followed by value proposition and iterative workflow. No filler, though the phrase 'better than guessing' is slightly promotional rather than informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter read-only tool with full schema coverage and no output schema, the description covers purpose, ranking, budgeting, and iterative usage. It gives the agent everything needed to call it correctly without redundancy.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters like target_entity, strategy, and max_tokens are already documented with descriptions and an enum. The description adds no parameter-specific syntax or default details beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('get') and resource ('optimal context'), and immediately clarifies what that means: 'the most relevant related code ranked by importance and fitted to a token budget.' This distinguishes it from siblings like get_entity or get_dependencies, which return raw results without ranking or budgeting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to call it ('Before working on a function') and how to iterate ('read the top results, then call it again on any surprising dependencies'). It fills the gap left by annotations by giving a concrete workflow, which is valuable for a context-building tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_test_coverageGet Test CoverageARead-onlyIdempotent
See which production functions are actually exercised by tests via the call graph — semantic coverage, not line coverage. Optionally ranks uncovered functions by blast radius so you know which missing tests are riskiest.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| weight_by_blast_radius | No | Rank uncovered entities by blast radius to prioritize testing (default: false, slower) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context — it's semantic coverage via call graph, and the blast-radius ranking is flagged as slower in the schema — but it doesn't describe output format or the openWorldHint implications. Baseline 3 is appropriate given annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core differentiator ('semantic coverage, not line coverage') followed by the optional ranking behavior. Every clause earns its place; no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a 2-param read-only analysis tool. The description explains the novel concept (call-graph-based semantic coverage), clarifies what the optional flag buys you, and the schema plus annotations cover the remaining requirements. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema, including the default and slowdown note for weight_by_blast_radius. The description reinforces the intent of that flag ('so you know which missing tests are riskiest') but adds no syntax or format detail beyond the schema. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('See which production functions are actually exercised') and resource (semantic coverage via call graph), and explicitly contrasts with a common alternative concept ('not line coverage'). No sibling tool covers test coverage, so the agent can identify its role immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (assess coverage, prioritize missing tests) but offers no explicit when-to-use vs. when-not-to-use guidance, no prerequisites, and names no alternative sibling tool. A 3 is fair for implied but unstated routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_topologyGet TopologyARead-onlyIdempotent
Get the full API surface map in one call — all routes, middleware, auth patterns, and database tables. Use this when you need to understand the overall architecture without reading every file. Returns the information you'd normally piece together from dozens of file reads.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world. The description adds useful behavioral context beyond that: this is an aggregate call that replaces 'dozens of file reads', signaling breadth and cost-savings. It doesn't discuss result size or pagination, but for a safe read the bar is lower.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what it returns, followed by when to use it and the value proposition. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns, and it does so by enumerating the four artifact categories. It is complete enough to call correctly, though it says nothing about volume or shape of the aggregate response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'project' parameter is fully documented in the schema, including the pointer to list_projects. The description adds no further meaning about the parameter, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (full API surface map), then enumerates the contents: routes, middleware, auth patterns, database tables. This distinguishes it from siblings like get_auth_matrix or get_data_flow, which cover narrower slices of the same surface.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition: use it when you need overall architecture understanding without reading every file. It does not name an alternative tool or state when not to use it (e.g. narrow questions better served by get_auth_matrix), so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modulesList ModulesARead-onlyIdempotent
Get a bird's-eye view of how the codebase is organized. Groups all symbols by architectural layer (route/service/component), module, file, or language with counts. Use this to orient yourself in an unfamiliar codebase before diving into specifics.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| group_by | No | How to group entities (default: layer) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds that the output is an aggregated grouping with counts rather than raw entities, which is useful context, but nothing about scale limits, cost, or result size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences; the overview purpose leads, the grouping dimensions follow, and the usage cue closes. No padding or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by stating the return shape (groups with counts). It is sufficient for a two-param read-only aggregation tool, though the exact structure of the grouped result remains implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description goes beyond it by clarifying that 'layer' means architectural layers such as route/service/component, giving the enum value concrete meaning the schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list/group symbols) and enumerates the grouping axes (layer, module, file, language) plus the returned counts. Clearly distinguishable from granular siblings like query_entities or get_entity, which fetch individual symbols rather than an organizational overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the use case: 'orient yourself in an unfamiliar codebase before diving into specifics,' and the schema points to list_projects for multi-project mode. It gives clear context but doesn't state when an alternative (e.g., query_entities) would be preferable once oriented.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsList ProjectsARead-onlyIdempotent
Start here. Returns all synced codebases with their size, language, and project name. You need the project name for every other tool. If this returns empty, use sync_project to import the current repo.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive, and open-world behavior, so the burden is lower. The description adds the critical workflow context that this is the entry point and that an empty result should trigger sync_project, which is genuinely useful behavioral guidance beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with 'Start here.' Every sentence earns its place by conveying scope, output fields, dependency for other tools, and failure handling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no params, no output schema, annotations cover safety), the description is complete enough to call and interpret correctly. It could mention pagination or the exact output shape, but for a bootstrap listing tool this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 per the rubric. The description correctly implies no arguments are needed and focuses on output semantics instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (projects/codebases), and explicitly lists the returns (size, language, project name). It also positions itself relative to siblings by noting its output is required for every other tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Start here' and 'You need the project name for every other tool', establishing when to call it and why. It also names the remediation path (sync_project) when the result is empty, which is a clear when-and-what-else instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_entitiesQuery EntitiesARead-onlyIdempotent
Like grep but for code structure. Find functions, classes, and routes by name, architectural layer (route/service/component), or module. Returns matching symbols with their type, file, and layer — use this instead of grep when you need to find code by what it does, not by text content.
| Name | Required | Description | Default |
|---|---|---|---|
| layer | No | Filter by architectural layer: route, controller, service, repository, utility, hook, component, schema | |
| limit | No | Max results to return (default: 50) | |
| query | No | Search term — matches against symbol name, ID, source file, and module | |
| module | No | Filter by module (partial match) | |
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| language | No | Filter by source language: javascript, typescript, python, go, rust, etc. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds that it finds symbols by architectural layer and returns type, file, and layer — useful context. However, it doesn't mention pagination, result limits (beyond the default in schema), or whether results are ranked/filtered. With annotations covering the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste, front-loading the analogy and core functionality, then the usage guideline. Every clause earns its place, and the structure moves from what it is to when to use it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with no output schema and 100% schema coverage, the description is complete enough: it explains what it finds, how to filter, what it returns, and when to prefer it over grep. It misses guidance on result ordering or pagination, which could matter for an agent planning calls, but against a rich schema and annotations, the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are fully documented in the schema. The description mentions the primary query dimensions (name, layer, module) but doesn't add syntax, format, or interaction details beyond what the schema provides. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('find functions, classes, and routes') and explains the mechanism ('by name, architectural layer, or module'). The 'grep but for code structure' analogy is vivid and immediately conveys the tool's purpose. It distinguishes itself from other query tools but doesn't explicitly contrast with siblings like get_entity or find_by_constraint, which also retrieve code entities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: 'use this instead of grep when you need to find code by what it does, not by text content.' This is an explicit when-to-use statement. However, it doesn't specify when to use this versus other sibling tools like get_entity (which might retrieve a single entity) or find_by_constraint (which might be for more complex queries). The guidance is strong but incomplete for a tool among 26 siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_traitsQuery TraitsARead-onlyIdempotent
Find functions by inferred behavioral trait — "fallible" (can fail, including transitively via callees that throw), "asyncContext" (carries async state), "generator" (yields values). Traits are semantic properties inferred from the call graph, not just syntax. Use this when you need to find all code with a specific capability. For syntactic tags (explicit throw statements, DB access), use find_by_constraint instead.
| Name | Required | Description | Default |
|---|---|---|---|
| trait | Yes | The trait or capability to search for (e.g., "fallible", "asyncContext") | |
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the key behavioral context that traits are call-graph-inferred semantics rather than syntax, but does not discuss performance, result format, or limitations of inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each with a distinct role: definition, trait examples with semantics, when-to-use, and alternative tool. The parenthetical trait definitions are dense but front-loaded relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, semantics, trait vocabulary, and sibling routing for a read-only query tool with full schema coverage and no output schema. Missing only the shape of returned results and multi-project mode behavior implied by the 'project' param.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description supplies example trait values ('fallible', 'asyncContext', 'generator') that enrich the enum-less 'trait' parameter, adding modest value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Find functions by inferred behavioral trait') and enumerates example traits. It distinguishes from sibling find_by_constraint by contrasting semantic traits vs. syntactic tags, though the contrast arrives late.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('when you need to find all code with a specific capability') and names an alternative for a different case ('For syntactic tags ... use find_by_constraint instead'). Lacks exclusions for other sibling tools like find_error_gaps or find_entry_points.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_projectSync ProjectAIdempotent
Import a public GitHub repo into Seshat for structural analysis. Call this when list_projects returns empty or when the user wants to analyze a new repo. Detects the git remote automatically if no URL is provided. Extraction typically takes 5-30 seconds. After syncing, call list_projects to confirm the project is available. Use force: true to re-extract even if a cached snapshot exists (useful after code changes or when Seshat extraction has been updated).
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Force re-extraction even if a cached snapshot exists. Default: false. | |
| repo_url | No | Public GitHub repo URL (e.g., https://github.com/org/repo). If omitted, tries to detect from the current git remote. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare it is a non-read-only, idempotent, non-destructive, open-world write, so safety is covered. The description adds genuinely new behavioral context: 5-30s extraction latency, automatic git remote detection, and cache/force re-extraction semantics. It stops short of stating failure modes or auth/network requirements, which openWorldHint implies but does not spell out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five tightly packed sentences, each carrying distinct information: purpose, trigger, default behavior, latency, follow-up, and force semantics. The core action is front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only two optional parameters, the description supplies everything an agent needs: trigger conditions, default remote detection, expected duration, verification step, and override behavior. Nothing material is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining the rationale for force ('useful after code changes or when Seshat extraction has been updated') rather than merely restating the default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Import a public GitHub repo into Seshat for structural analysis'), naming both the action and the target system. The scope is narrow and immediately distinguishable from read-only siblings like list_projects and query_entities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions ('when list_projects returns empty or when the user wants to analyze a new repo'), a post-condition step ('call list_projects to confirm'), and the condition for re-running with force. Alternatives and follow-ups are fully covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_data_pathTrace Data PathARead-onlyIdempotent
Follow data from one function through the call graph to its sinks. From a start entity it walks callees, records the data each hop consumes/produces/mutates, and reports the chain from the start to every sink the data reaches — database writes, network egress, filesystem writes — plus the tables touched. Call it before changing a function that handles real data: it answers "where does this data end up?", the cross-call composition get_data_flow (one entity) cannot give. Flags untrusted inputs at the source.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project name (required in multi-project mode). Use list_projects to see available projects. | |
| entity_id | Yes | Entity ID or name to trace data flow from | |
| max_depth | No | How many call hops to follow downstream (default: 5, max: 8) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds real value beyond that: it explains the traversal mechanism (walks callees hop by hop), what is recorded at each hop, which sink classes are reported, and that untrusted inputs are flagged at the source.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and mechanism in the first sentence, then usage and differentiation. Dense but every clause carries information; the only mild excess is the sink enumeration, which is still useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the return shape: the chain from start to every sink, sink categories, and tables touched. Prerequisites are implied by the usage sentence. A brief note on how the chain is ordered or truncated at max_depth would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema (including the max_depth default and cap). The description reinforces the start-entity semantics and mentions the tables-touched output but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Follow data from one function through the call graph to its sinks') and precisely delineates scope: walks callees, records consumed/produced/mutated data, reports chains to sinks plus tables touched. It explicitly contrasts with the sibling get_data_flow, so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when: 'Call it before changing a function that handles real data.' It also names the alternative (get_data_flow) and the exact capability gap that selects trace_data_path over it ('the cross-call composition get_data_flow (one entity) cannot give').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
27 tool updates
v0.20.2- First observed
find_by_constraint - First observed
find_dead_code - First observed
find_entry_points - First observed
find_error_gaps - First observed
find_exposure_leaks - First observed
find_layer_violations - First observed
find_ownership_violations - First observed
find_semantic_clones - First observed
get_account_status - First observed
get_auth_matrix - First observed
get_blast_radius - First observed
get_co_change_clusters - First observed
get_coupling_metrics - First observed
get_data_flow - First observed
get_dependencies - First observed
get_entity - First observed
get_hotspots - First observed
get_lineage - First observed
get_optimal_context - First observed
get_test_coverage - First observed
get_topology - First observed
list_modules - First observed
list_projects - First observed
query_entities - First observed
query_traits - First observed
sync_project - First observed
trace_data_path
TDQS
Scored across 27 tools
Tools have clearly distinct purposes, and descriptions actively cross-reference complementary tools (e.g., get_data_flow vs trace_data_path, query_traits vs find_by_constraint, get_dependencies vs get_blast_radius). A few pairs still require careful reading to distinguish—such as dependency tracing, blast radius, and optimal context—but the boundaries are explicitly stated.
Every tool follows a consistent snake_case verb_noun pattern: get_, list_, find_, query_, sync_, trace_. No mixed conventions or vague names appear.
27 tools is heavy for a single MCP server, though the deep code-intelligence domain justifies many specialized analyses. Several tools could likely be consolidated or grouped (e.g., multiple history and impact tools), but each addresses a distinct question.
Coverage is broad across structure, dependencies, data flow, history, architecture, security, testing, and dead code. Minor gaps exist for project lifecycle management (delete/update) and raw text/file search, but core analysis workflows are well covered.
Related MCP Connectors
Hosted code graph over MCP: exact callers, dependencies, and cross-repo blast radius for AI agents.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Codebase intelligence for AI agents — dead code, blast radius, ownership.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides code intelligence for AI coding agents by indexing repositories into a hybrid knowledge graph, enabling agents to query dependencies, impact, and context through 28 MCP tools.3Apache 2.0
- AlicenseNot gradedqualityAmaintenanceEnables AI coding agents to query a codebase as a knowledge graph, providing token-budgeted context, search, and impact analysis via MCP tools.23 PyPIMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to query a local, continuously updated symbol graph of a codebase, providing ranked search, caller/callee exploration, dependency paths, and git-diff impact analysis through MCP tools.8 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to query a local, multi-repo code graph for symbol exploration, blast radius analysis, co-change mining, and durable code-anchored memory via MCP tools.143 npm14Business Source 1.1