Apache Health MCP
Apache Health MCP
Este repositorio contiene un pequeño servidor MCP para consultar los informes de salud de Apache Incubator desde tools/health/reports.
Analiza el formato de informe Markdown utilizado por las herramientas de salud de Apache y expone herramientas MCP para:
listar los informes de podlings disponibles
buscar nombres de podlings
obtener un resumen analizado para un podling
devolver el informe Markdown original
devolver métricas para una ventana específica
comparar un podling en dos o tres ventanas
listar métricas y ventanas compatibles
clasificar podlings por una métrica dentro de una ventana como
3m,6mo12m
Entrada esperada
Apunte el servidor a un directorio local que contenga archivos Markdown como:
reports/
Amoro.md
Iggy.md
...El analizador está diseñado en torno a la estructura actual de informes de Apache, especialmente la sección ## Window Details.
Related MCP server: IPMC MCP
Instalación
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install .Para desarrollo local:
make install-devEjecución
health-mcp --reports-dir /path/to/incubator/tools/health/reportsEl servidor utiliza stdio, por lo que está diseñado para ser iniciado por un cliente MCP.
Para desarrollo local sin instalar primero, aún puede iniciar el servidor stdio directamente:
python3 server.pyEl paquete también mantiene apache-health-mcp como un alias de comando compatible con versiones anteriores.
Claude Desktop
Edite ~/Library/Application Support/Claude/claude_desktop_config.json y añada:
{
"mcpServers": {
"apache-health": {
"command": "health-mcp",
"args": [
"--reports-dir",
"/path/to/incubator/tools/health/reports"
]
}
}
}Luego reinicie Claude Desktop. Si instaló en un entorno virtual que no está en su PATH, utilice la ruta absoluta al comando health-mcp de ese entorno.
Herramientas MCP
health_overview
Devuelve el directorio de informes, el recuento de informes, la lista de podlings y la última fecha de generación.
list_podlings
Devuelve los nombres de los podlings disponibles en el directorio de informes.
search_podlings
Busca nombres de podlings mediante una subcadena que no distingue entre mayúsculas y minúsculas con un límite de resultados opcional.
get_report_summary
Devuelve métricas de ventana analizadas para un solo podling.
get_report_markdown
Devuelve el Markdown original para un informe de un solo podling.
get_window_metrics
Devuelve métricas para un podling y una ventana como 3m, 6m, 12m o to-date, incluyendo palabras de tendencia normalizadas como up, down y flat bajo trends.
compare_windows
Devuelve métricas lado a lado para un podling en dos o tres ventanas, incluyendo palabras de tendencia normalizadas bajo trends de cada ventana.
query_metric_rankings
Clasifica los podlings por una métrica analizada como commits, prs_merged, dev_messages, bus50 o median_merge_days.
list_metrics
Devuelve los nombres de las métricas compatibles y las ventanas disponibles para consultar.
Ejemplos de uso
Estos ejemplos muestran los tipos de preguntas que un usuario puede hacer a un cliente MCP conectado a este servidor.
Revisión de una instantánea de informe
"¿Qué informes de salud de Apache Incubator están disponibles en este checkout?"
"¿Cuántos informes de salud de podlings tenemos y cuándo se generaron?"
"¿Qué podlings tienen informes de salud que puedo consultar?"
"¿Sobre qué métricas de salud y ventanas de informe puedo preguntar?"
Investigación de un podling
"Muéstrame el resumen de salud de Amoro."
"¿Qué dice el último informe de salud sobre Iggy?"
"Busca podlings con nombres que contengan 'stream' y resume la mejor coincidencia."
"Para este podling, muestra las métricas de salud recientes de 3 meses."
"Muéstrame el informe Markdown original de Amoro para que pueda verificar la fuente."
Comparación de tendencias entre ventanas
"Compara la actividad de 3 meses, 6 meses y 12 meses de Amoro."
"¿La actividad de desarrollo de Iggy está mejorando o disminuyendo?"
"Compara la actividad reciente de la lista de correo con la tendencia a largo plazo de este podling."
"¿Ha cambiado la actividad de fusión de PR de este podling entre las ventanas de 3 y 12 meses?"
"¿El factor bus de este podling está mejorando o empeorando a través de las ventanas de informe?"
Búsqueda de podlings por señal de actividad
"¿Qué podlings tuvieron la mayor cantidad de mensajes en la lista de desarrollo en los últimos 3 meses?"
"Muéstrame podlings sin commits en los últimos 3 meses."
"¿Qué podlings tienen el tiempo medio de fusión de PR más largo?"
"Clasifica los podlings por PRs fusionados durante la ventana de 6 meses."
"Encuentra podlings con baja diversidad de revisores en la ventana de informe reciente."
Preparación de una cola de revisión humana
"Dame una lista corta de podlings que pueden necesitar atención de mentores según la actividad reciente."
"¿Qué podlings parecen tranquilos en cuanto a commits, PRs y mensajes de la lista de desarrollo?"
"Encuentra podlings con baja actividad reciente y compáralos con su tendencia de 12 meses."
"¿Qué podlings debería revisar manualmente por preocupaciones sobre el factor bus o la diversidad de revisores?"
Desarrollo
Las tareas comunes están disponibles a través de make:
make format
make lint
make typecheck
make test
make coverage
make checkNotas
Este servidor consulta archivos de informe ya generados. No ejecuta el script de recopilación ascendente de Apache.
El espacio de trabajo aquí no incluía un directorio
reports/local, por lo que el servidor está diseñado para aceptar cualquier clon local o instantánea copiada del directorio de informes de Apache.
Available Tools
9 toolscompare_windowsC
Compare one podling across two or three windows.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions 'compare' but doesn't specify whether this is a read-only analysis, if it requires specific permissions, what the output format is, or any rate limits. For a tool with zero annotation coverage, this is a significant gap in transparency about how the tool behaves operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's appropriately sized for a tool with no parameters and is front-loaded with the core action. Every part of the sentence contributes to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no annotations, no output schema, and 0 parameters, the description is incomplete for effective use. It doesn't explain what 'compare' entails (e.g., metrics compared, output format), behavioral traits, or usage context relative to siblings. For a comparison tool in a metric-focused server, more detail is needed to guide the agent adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, meaning no parameters are documented in the schema. The description doesn't add parameter details, but since there are no parameters, this is acceptable. The baseline for 0 parameters is 4, as the description doesn't need to compensate for missing param info.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action ('compare') and target resource ('one podling across two or three windows'), which is clear but somewhat vague. It doesn't specify what aspects are compared or how the comparison is performed. However, it distinguishes from siblings like 'list_podlings' or 'get_window_metrics' by focusing on comparison rather than listing or retrieving metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives is provided. The description implies usage for comparing podlings across windows, but doesn't mention prerequisites, when-not-to-use scenarios, or how it differs from siblings like 'query_metric_rankings' or 'search_podlings' that might involve podling analysis. This leaves the agent without clear contextual boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_report_markdownB
Return the raw markdown for one podling report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool returns raw markdown but doesn't explain how the report is selected, if authentication is needed, potential errors, or response format details. This leaves significant gaps for a tool that likely involves data retrieval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without any wasted words. It's front-loaded with the core action and resource, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of retrieving a specific report (implied by 'one podling report'), no annotations, and no output schema, the description is incomplete. It doesn't explain how to specify which report, what the markdown contains, or error handling, leaving the agent with insufficient context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add param info, but that's acceptable here. A baseline of 4 is appropriate since the schema fully handles the lack of parameters without requiring compensation from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return') and resource ('raw markdown for one podling report'), making the tool's purpose understandable. However, it doesn't differentiate from sibling tools like 'get_report_summary' or explain what distinguishes 'raw markdown' from other report formats, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_report_summary' or 'search_podlings'. It lacks context about prerequisites, such as how to identify the specific podling report, or any exclusions, leaving usage unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_report_summaryB
Get parsed metrics for one podling report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but offers minimal behavioral insight. It implies a read operation ('Get') but doesn't disclose authentication needs, rate limits, error conditions, or what 'parsed metrics' entails (format, structure, or completeness). For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded with the core action and resource, making it immediately understandable without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a tool that presumably returns parsed metrics, the description is incomplete. It doesn't explain what 'parsed metrics' includes, how the podling report is identified, or the return format. For a tool in a metric-heavy context with multiple siblings, more detail is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the lack of inputs. The description adds value by specifying the resource ('one podling report'), implying it operates on a single, implicitly identified report. This contextual meaning goes beyond the empty schema, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('parsed metrics for one podling report'), making the purpose understandable. It distinguishes from siblings like 'get_report_markdown' (which likely returns raw markdown) by specifying 'parsed metrics', but doesn't explicitly differentiate from 'get_window_metrics' or other metric-related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_report_markdown', 'get_window_metrics', or 'list_metrics'. It doesn't mention prerequisites, context for 'podling report', or when this tool is preferred over other metric-retrieval options.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_window_metricsB
Return metrics for a single podling/window combination.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but only states what the tool returns without behavioral details. It doesn't disclose whether this is a read-only operation, potential errors, rate limits, or authentication needs, leaving significant gaps for a tool that likely queries metrics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose with no wasted words. It is appropriately sized and front-loaded, making it easy to understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters and no output schema, the description is minimally adequate but lacks completeness. It doesn't explain what metrics are returned, their format, or error handling, which are important for a metrics query tool with no structured output documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by specifying that metrics are for a 'single podling/window combination', which clarifies the scope beyond what the empty schema provides, justifying a score above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Return' and the resource 'metrics for a single podling/window combination', making the purpose specific and understandable. It doesn't explicitly distinguish from siblings like 'list_metrics' or 'query_metric_rankings', but the focus on a single combination provides some implicit differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description implies it's for a specific podling/window pair, but it doesn't mention prerequisites, when not to use it, or refer to sibling tools like 'list_metrics' for broader queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
health_overviewB
Return a high-level summary of the available Apache health reports.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool returns a summary but doesn't specify format, data freshness, rate limits, or authentication needs. This leaves critical operational details unclear for a tool that likely involves data retrieval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It's front-loaded with the core action and resource, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters and no output schema, the description adequately covers the basic purpose. However, for a health reporting tool in a server with multiple related siblings, it lacks context on output format or how it complements other tools, leaving gaps in overall understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing on the tool's purpose instead, which aligns well with the schema's simplicity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Return') and resource ('high-level summary of available Apache health reports'), making the purpose specific and understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_report_summary' or 'get_report_markdown', which might offer similar functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_report_summary' or 'list_metrics'. It lacks context about scenarios where a high-level overview is preferred over detailed reports, leaving the agent without usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_metricsB
Return the supported metrics and windows for querying.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool returns data but doesn't disclose behavioral traits like whether it's a read-only operation, if it requires authentication, rate limits, or what the return format looks like (e.g., list, object). This leaves significant gaps for an agent to understand how to handle the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any unnecessary words. It's front-loaded and wastes no space, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and no output schema, the description is minimally adequate but lacks depth. It doesn't explain the return values (e.g., structure of metrics/windows) or any behavioral context, which could be important for querying tools. However, the simplicity of the tool (0 params) means the description isn't severely lacking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add parameter details, which is appropriate here, and it implies the tool takes no inputs, aligning with the schema. A baseline of 4 is given since no parameters exist and the schema fully covers them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return') and the target ('supported metrics and windows for querying'), making the purpose understandable. However, it doesn't explicitly differentiate this tool from its siblings like 'get_window_metrics' or 'query_metric_rankings', which appear related to metrics/windows, so it doesn't reach the highest score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings such as 'get_window_metrics' and 'query_metric_rankings' that might overlap in functionality, there's no indication of context, prerequisites, or exclusions for using 'list_metrics'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_podlingsB
List podlings that have a parsed markdown report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the filter condition ('have a parsed markdown report') but doesn't disclose behavioral traits such as pagination, rate limits, permissions needed, or what happens if no podlings meet the criteria. For a tool with zero annotation coverage, this leaves significant gaps in understanding its operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any redundant information. It is appropriately sized and front-loaded, making it easy to understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, no annotations, and no output schema, the description is minimal but adequate for a simple listing tool. It specifies a filter condition, which adds some context, but lacks details on behavior, output format, or integration with siblings, leaving room for improvement in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by specifying the filter condition ('have a parsed markdown report'), which provides context beyond the schema, though it doesn't detail how this filtering is applied internally.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and resource ('podlings'), specifying that they must 'have a parsed markdown report'. This distinguishes it from generic listing tools by adding a filter condition, though it doesn't explicitly differentiate from sibling tools like 'search_podlings'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'search_podlings' or other siblings. The description implies usage for podlings with parsed markdown reports but doesn't specify exclusions, prerequisites, or comparative contexts with other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_metric_rankingsC
Rank podlings by one parsed metric for a specific window.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool performs ranking, which implies a read-only operation, but doesn't disclose behavioral traits like whether it requires authentication, has rate limits, returns paginated results, or what the output format looks like. This is inadequate for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose ('Rank podlings') and adds necessary qualifiers ('by one parsed metric for a specific window') without any wasted words. It's appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0 parameters, the description is incomplete. It lacks details on behavioral aspects like authentication needs, rate limits, or output format, and doesn't clarify how it differs from siblings. For a ranking tool with no structured data, this leaves significant gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds context by specifying 'by one parsed metric for a specific window', which implies inputs might be inferred from context or defaults, but since there are no parameters, a baseline of 4 is appropriate as it doesn't need to compensate for gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Rank podlings by one parsed metric for a specific window', which provides a clear verb ('Rank'), resource ('podlings'), and scope ('by one parsed metric for a specific window'). However, it doesn't explicitly differentiate from siblings like 'get_window_metrics' or 'list_podlings', leaving ambiguity about when to use this versus those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description mentions ranking by a metric for a window, but it doesn't specify prerequisites, exclusions, or compare it to siblings such as 'get_window_metrics' or 'list_podlings', leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_podlingsB
Search podling names by case-insensitive substring.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the search is 'case-insensitive', which is useful, but fails to describe other critical behaviors like response format, error handling, or performance characteristics. This leaves significant gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without any redundant or unnecessary information. It is front-loaded and appropriately sized for its purpose, earning a perfect score for conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is adequate but incomplete. It covers the basic search action but lacks details on output format or behavioral context, making it minimally viable but with clear gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by specifying the search mechanism ('case-insensitive substring'), which is not captured in the schema, justifying a score above the baseline of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search') and resource ('podling names') with a specific constraint ('by case-insensitive substring'), making the purpose evident. However, it does not explicitly differentiate from sibling tools like 'list_podlings', which could serve a similar listing function, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as 'list_podlings' for unfiltered listing or other search-related tools. It lacks context on prerequisites, exclusions, or typical use cases, offering minimal usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
compare_windows - First observed
get_report_markdown - First observed
get_report_summary - First observed
get_window_metrics - First observed
health_overview - First observed
list_metrics - First observed
list_podlings - First observed
query_metric_rankings - First observed
search_podlings
TDQS
Scored across 9 tools
Each tool has a clearly distinct purpose with no overlap: compare_windows compares podlings across windows, get_report_markdown retrieves raw markdown, get_report_summary provides parsed metrics, get_window_metrics gives metrics for a single podling/window, health_overview offers a high-level summary, list_metrics enumerates supported metrics/windows, list_podlings lists podlings with reports, query_metric_rankings ranks podlings by metric, and search_podlings searches podling names. The descriptions unambiguously differentiate each tool's function.
All tools follow a consistent verb_noun or verb_adjective_noun pattern using snake_case: compare_windows, get_report_markdown, get_report_summary, get_window_metrics, health_overview, list_metrics, list_podlings, query_metric_rankings, and search_podlings. The naming is predictable and readable throughout, with no deviations or mixed conventions.
With 9 tools, the count is well-scoped for the Apache health reporting domain. Each tool earns its place by covering distinct aspects such as listing, retrieving, comparing, searching, and ranking podling health data. This is neither too thin nor too heavy, providing comprehensive functionality without bloat.
The tool surface offers complete coverage for querying and analyzing Apache podling health reports. It includes listing and searching podlings, retrieving raw and parsed report data, comparing across windows, getting metrics and rankings, and providing overviews. There are no obvious gaps; agents can perform full workflows from discovery to detailed analysis without dead ends.
Maintenance
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP server for medicaid-intelligence
MCP server for the Seline Analytics API
MCP server for Withings health data — sleep, activity, heart, and body metrics.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server for accessing and analyzing Apache Software Foundation Incubator podling data from podlings.xml. It provides tools to query podling metadata, generate statistics, and analyze incubation trends over time.22MIT
- AlicenseAqualityCmaintenanceA dependency-free MCP server for Apache Incubator PMC oversight that helps identify podlings needing attention, assess graduation readiness, and generate podling briefings by combining lifecycle data and community health signals.212MIT
- AlicenseAqualityCmaintenanceAn MCP server for analyzing startup financial health and generating metrics reports locally.2MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for comprehensive PyPI package intelligence, providing tools for dependency analysis, security scanning, health scoring, license compliance, and trend tracking.MIT