Skip to main content
Glama

Server Quality Checklist

58%
Profile completionA complete profile improves this server's visibility in search results.
  • Latest release: v0.1.0

  • Disambiguation4/5

    Each tool targets a distinct stage of leakage detection: profiling, dataset auditing, code auditing, and feature availability review. However, profile_dataset and audit_dataset could be confused by an agent since both operate on datasets, though their descriptions help separate general profiling from leakage-specific auditing.

    Naming Consistency5/5

    All tool names consistently follow a verb_noun pattern using lowercase snake_case: profile_dataset, audit_dataset, audit_training_code, and review_feature_availability. This makes the tool set predictable and easy to navigate.

    Tool Count5/5

    Four tools is well-scoped for a specialized leakage-detection server. Each tool has a clear, non-redundant role, and the count feels appropriate rather than thin or bloated.

    Completeness4/5

    The tool surface covers the core leakage-detection workflow: profile data, audit data for leakage, audit training code, and verify feature availability. Minor gaps exist around remediation or actionable reporting after an audit, but the main detection loop is complete.

  • Average 3.1/5 across 4 of 4 tools scored. Lowest: 2.4/5.

    See the Tool Scores section below for per-tool breakdowns.

    • No community issues in the last 6 months
    • 1 commit in the last 12 weeks
    • No stable releases found
    • No critical vulnerability alerts
    • No high-severity vulnerability alerts
    • No code scanning findings
    • CI is passing
  • This repository is licensed under MIT License.

  • This repository includes a README.md file.

  • No tool usage detected in the last 30 days. Usage tracking helps demonstrate server value.

    Tip: use the "Try in Browser" feature on the server page to seed initial usage.

  • Add a glama.json file to provide metadata about your server.

  • If you are the author, simply .

    If the server belongs to an organization, first add glama.json to the root of your repository:

    {
      "$schema": "https://glama.ai/mcp/schemas/server.json",
      "maintainers": [
        "your-github-username"
      ]
    }

    Then . Browse examples.

  • Add related servers to improve discoverability.

How to sync the server with GitHub?

Servers are automatically synced at least once per day, but you can also sync manually at any time to instantly update the server profile.

To manually sync the server, click the "Sync Server" button in the MCP server admin interface.

How is the quality score calculated?

The overall quality score combines two components: Tool Definition Quality (70%) and Server Coherence (30%).

Tool Definition Quality measures how well each tool describes itself to AI agents. Every tool is scored 1–5 across six dimensions: Purpose Clarity (25%), Usage Guidelines (20%), Behavioral Transparency (20%), Parameter Semantics (15%), Conciseness & Structure (10%), and Contextual Completeness (10%). The server-level definition quality score is calculated as 60% mean TDQS + 40% minimum TDQS, so a single poorly described tool pulls the score down.

Server Coherence evaluates how well the tools work together as a set, scoring four dimensions equally: Disambiguation (can agents tell tools apart?), Naming Consistency, Tool Count Appropriateness, and Completeness (are there gaps in the tool surface?).

Tiers are derived from the overall score: A (≥3.5), B (≥3.0), C (≥2.0), D (≥1.0), F (<1.0). B and above is considered passing.

Tool Scores

  • Behavior2/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    With no annotations, the description must carry behavioral disclosure, but it only says the tool asks Gemini about availability. It does not state whether the operation has side effects, requires external connectivity, returns a report, or behaves differently depending on inputs, leaving the agent to guess.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness4/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is one sentence with no redundant clauses and the core action appears first. 'Optionally' is a minor filler, but overall the structure is appropriately compact.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness1/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a tool with three required parameters, no output schema, and no annotations, this description is far too thin: it leaves parameter meanings, expected input formats, return behavior, and relationship to siblings unspecified. An agent cannot invoke it correctly from this text alone.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters1/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 0% and the description does not explain path, target, or prediction_time. With three required parameters and no semantics beyond their names, an agent cannot confidently construct valid arguments.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose4/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description names a specific action (asking Gemini) and subject (features that may be unavailable at prediction time), so an agent can roughly understand the tool's purpose. However, it does not define what 'features' refers to or differentiate it from the sibling dataset/audit tools, so it stops short of a 5.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines2/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    It offers no guidance on when to invoke this tool, what conditions warrant it, or why an agent might choose it over profile_dataset, audit_dataset, or audit_training_code. The word 'Optionally' hints it is not mandatory but does not provide a decision rule.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior2/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    With no annotations provided, the description carries the full burden of behavioral disclosure. It does not state whether the tool is read-only, what side effects or outputs are produced, or how it behaves with partial parameters. The term 'Detect' implies analysis, but this is not explicit and important behavioral details are missing.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness3/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    The description is extremely concise with no fluff, but it is a flat list of risk categories rather than a structured explanation. It is brief and front-loaded, yet the brevity comes at the cost of needed detail.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness2/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    For a 5-parameter tool with no output schema and no annotations, the description is incomplete. It does not explain what the audit produces, how to provide required inputs, or what constraints exist. The agent lacks critical context needed to invoke it correctly.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters2/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 0% and there are 5 parameters, so the description must compensate. It mentions 'target', 'entity', and 'temporal', which vaguely map to the target, entity, and time_column parameters, but it does not explain their formats, relationships, or the required path parameter.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description uses a specific verb ('Detect') and a clear resource (dataset risks), enumerating concrete risk categories: target, entity, and temporal leakage, PII, identifiers, and metric risks. This distinguishes it from siblings like profile_dataset and audit_training_code, which focus on profiling and code auditing respectively.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines2/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description gives no guidance on when to use this tool versus profile_dataset, audit_training_code, or review_feature_availability. It does not state conditions, exclusions, or alternatives, leaving the agent to infer the appropriate selection context.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior3/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    With no annotations, the description carries the full burden. It discloses that the check is static ('statically detect') and that it targets preprocessing/resampling placement, but it does not describe side effects, failure modes, or what happens when no issues are found. The output schema covers some return expectations.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    A single focused sentence with no filler. The verb and object are front-loaded and every word adds meaning.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness3/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    The description is adequate for a one-parameter tool with an output schema, but it omits path semantics and usage routing to dataset-focused siblings. It is not as complete as it could be for a first-time agent.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters2/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    Schema description coverage is 0% and the description never explains the 'path' parameter. It does not clarify whether path points to a file, directory, or specific language, so the agent must infer from the tool name.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose5/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description names a specific action ('statically detect') and a specific target ('preprocessing or resampling before the train/test split'). This clearly differentiates it from sibling dataset-focused tools like audit_dataset or profile_dataset.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines3/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description implies the tool is for auditing training code for leakage before splitting, but it does not explicitly state when to prefer it over siblings or mention exclusions. No alternative tools are named.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

  • Behavior3/5

    Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

    With no annotations provided, the description carries the burden of behavioral disclosure. It usefully reveals that the full dataset is not sent to the model, implying local processing, but it does not state whether the operation is read-only, what side effects occur, or what the profiling process entails beyond that privacy guarantee.

    Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

    Conciseness5/5

    Is the description appropriately sized, front-loaded, and free of redundancy?

    A single, well-structured sentence that front-loads the core purpose and adds the important privacy qualifier. Every word contributes value, with no filler or redundancy.

    Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

    Completeness2/5

    Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

    With no annotations, no output schema, and no schema-level parameter descriptions, the one-line description leaves significant gaps: it does not describe the returned profile structure, clarify sample_rows semantics, or help the agent choose between this and audit_dataset. The simple input schema keeps it from being a 1, but an agent would still be guessing about important call behavior.

    Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

    Parameters3/5

    Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

    The description adds meaning to the path parameter by specifying it must be a local file in CSV, Parquet, or JSONL format. However, it does not explain sample_rows or how it interacts with profiling, and schema description coverage is 0%, so the description only partially compensates.

    Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

    Purpose4/5

    Does the description clearly state what the tool does and how it differs from similar tools?

    The description clearly identifies the action ('Profile') and the resource ('a local CSV, Parquet, or JSONL'), and adds a meaningful qualifier ('without sending the full dataset to the model'). It does not explicitly differentiate from the sibling audit_dataset, which also operates on datasets, so it falls just short of a full 5.

    Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

    Usage Guidelines4/5

    Does the description explain when to use this tool, when not to, or what alternatives exist?

    The description gives clear context: use this tool to profile local tabular/structured files while avoiding sending the full dataset to the model. However, it does not state exclusions or explicitly point to alternatives such as audit_dataset when a full audit is needed.

    Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

GitHub Badge

Glama performs regular codebase and documentation scans to:

  • Confirm that the MCP server is working as expected.
  • Confirm that there are no obvious security issues.
  • Evaluate tool definition quality.

Our badge communicates server capabilities, safety, and installation instructions.

Card Badge

LeakageLens MCP MCP server – quality and maintenance score on Glama

Copy to your README.md:

Score Badge

LeakageLens MCP MCP server – quality and maintenance score on Glama

Copy to your README.md: