Apache Health MCP
Apache Health MCP
이 저장소는 tools/health/reports에서 Apache Incubator 상태 보고서를 쿼리하기 위한 작은 MCP 서버를 포함하고 있습니다.
이 서버는 Apache의 상태 도구에서 사용하는 마크다운 보고서 형식을 파싱하며, 다음과 같은 MCP 도구를 제공합니다:
사용 가능한 포들링 보고서 나열
포들링 이름 검색
특정 포들링에 대한 파싱된 요약 정보 가져오기
원본 마크다운 보고서 반환
특정 기간에 대한 지표 반환
두세 개의 기간에 걸쳐 포들링 비교
지원되는 지표 및 기간 나열
3m,6m,12m과 같은 기간 내 지표별 포들링 순위 매기기
예상 입력
서버를 다음과 같은 마크다운 파일이 포함된 로컬 디렉토리를 가리키도록 설정하십시오:
reports/
Amoro.md
Iggy.md
...이 파서는 현재 Apache 보고서 구조, 특히 ## Window Details 섹션을 중심으로 설계되었습니다.
Related MCP server: IPMC MCP
설치
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install .로컬 개발 환경:
make install-dev실행
health-mcp --reports-dir /path/to/incubator/tools/health/reports이 서버는 stdio를 사용하므로 MCP 클라이언트에 의해 실행되도록 설계되었습니다.
먼저 설치하지 않고 로컬에서 개발하는 경우, 여전히 stdio 서버를 직접 실행할 수 있습니다:
python3 server.py이 패키지는 하위 호환성을 위해 apache-health-mcp라는 명령 별칭도 유지합니다.
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json을 편집하고 다음을 추가하십시오:
{
"mcpServers": {
"apache-health": {
"command": "health-mcp",
"args": [
"--reports-dir",
"/path/to/incubator/tools/health/reports"
]
}
}
}그런 다음 Claude Desktop을 다시 시작하십시오. PATH에 없는 가상 환경에 설치한 경우, 해당 환경의 health-mcp 명령에 대한 절대 경로를 사용하십시오.
MCP 도구
health_overview
보고서 디렉토리, 보고서 수, 포들링 목록 및 최신 생성 날짜를 반환합니다.
list_podlings
보고서 디렉토리에 있는 포들링 이름을 반환합니다.
search_podlings
대소문자를 구분하지 않는 부분 문자열로 포들링 이름을 검색하며, 선택적으로 결과 제한을 설정할 수 있습니다.
get_report_summary
단일 포들링에 대해 파싱된 기간별 지표를 반환합니다.
get_report_markdown
단일 포들링 보고서의 원본 마크다운을 반환합니다.
get_window_metrics
3m, 6m, 12m 또는 to-date와 같은 특정 포들링 및 기간에 대한 지표를 반환하며, trends 아래에 up, down, flat과 같은 정규화된 추세 단어를 포함합니다.
compare_windows
각 기간의 trends 아래에 정규화된 추세 단어를 포함하여, 두세 개의 기간에 걸쳐 포들링에 대한 지표를 나란히 비교하여 반환합니다.
query_metric_rankings
commits, prs_merged, dev_messages, bus50, median_merge_days와 같이 파싱된 지표별로 포들링 순위를 매깁니다.
list_metrics
쿼리에 지원되는 지표 이름과 사용 가능한 기간을 반환합니다.
사용 예시
이 예시들은 사용자가 이 서버에 연결된 MCP 클라이언트에 질문할 수 있는 유형을 보여줍니다.
보고서 스냅샷 검토
"이 체크아웃에서 사용할 수 있는 Apache Incubator 상태 보고서는 무엇인가요?"
"포들링 상태 보고서가 몇 개나 있으며, 언제 생성되었나요?"
"쿼리할 수 있는 상태 보고서가 있는 포들링은 무엇인가요?"
"어떤 상태 지표와 보고서 기간에 대해 질문할 수 있나요?"
단일 포들링 조사
"Amoro의 상태 요약을 보여줘."
"Iggy에 대한 최신 상태 보고서 내용은 무엇인가요?"
"이름에 'stream'이 포함된 포들링을 찾아서 가장 적합한 것을 요약해줘."
"이 포들링에 대해 최근 3개월간의 상태 지표를 보여줘."
"출처를 확인할 수 있도록 Amoro의 원본 마크다운 보고서를 보여줘."
기간별 추세 비교
"Amoro의 3개월, 6개월, 12개월 활동을 비교해줘."
"Iggy의 개발 활동이 개선되고 있나요, 아니면 둔화되고 있나요?"
"이 포들링의 최근 메일링 리스트 활동과 장기적인 추세를 비교해줘."
"이 포들링의 PR 병합 활동이 3개월 기간과 12개월 기간 사이에서 변했나요?"
"이 포들링의 버스 팩터(bus factor)가 보고서 기간 동안 개선되고 있나요, 아니면 악화되고 있나요?"
활동 신호로 포들링 찾기
"지난 3개월 동안 개발자 메일링 리스트 메시지가 가장 많았던 포들링은 무엇인가요?"
"지난 3개월 동안 커밋이 없는 포들링을 보여줘."
"중앙값 PR 병합 시간이 가장 긴 포들링은 무엇인가요?"
"6개월 기간 동안 병합된 PR 기준으로 포들링 순위를 매겨줘."
"최근 보고서 기간에 리뷰어 다양성이 낮은 포들링을 찾아줘."
사람의 검토 대기열 준비
"최근 활동을 바탕으로 멘토의 주의가 필요할 수 있는 포들링 목록을 짧게 알려줘."
"커밋, PR, 개발자 메일링 리스트 메시지 전반에서 활동이 저조해 보이는 포들링은 무엇인가요?"
"최근 활동이 저조한 포들링을 찾아 12개월 추세와 비교해줘."
"버스 팩터나 리뷰어 다양성 문제로 수동 검토해야 할 포들링은 무엇인가요?"
개발
일반적인 작업은 make를 통해 수행할 수 있습니다:
make format
make lint
make typecheck
make test
make coverage
make check참고
이 서버는 이미 생성된 보고서 파일을 쿼리합니다. Apache의 업스트림 수집 스크립트를 실행하지 않습니다.
이 작업 공간에는 로컬
reports/디렉토리가 포함되어 있지 않으므로, 서버는 Apache 보고서 디렉토리의 로컬 클론이나 복사된 스냅샷을 수용하도록 빌드되었습니다.
Available Tools
9 toolscompare_windowsC
Compare one podling across two or three windows.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions 'compare' but doesn't specify whether this is a read-only analysis, if it requires specific permissions, what the output format is, or any rate limits. For a tool with zero annotation coverage, this is a significant gap in transparency about how the tool behaves operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's appropriately sized for a tool with no parameters and is front-loaded with the core action. Every part of the sentence contributes to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no annotations, no output schema, and 0 parameters, the description is incomplete for effective use. It doesn't explain what 'compare' entails (e.g., metrics compared, output format), behavioral traits, or usage context relative to siblings. For a comparison tool in a metric-focused server, more detail is needed to guide the agent adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, meaning no parameters are documented in the schema. The description doesn't add parameter details, but since there are no parameters, this is acceptable. The baseline for 0 parameters is 4, as the description doesn't need to compensate for missing param info.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action ('compare') and target resource ('one podling across two or three windows'), which is clear but somewhat vague. It doesn't specify what aspects are compared or how the comparison is performed. However, it distinguishes from siblings like 'list_podlings' or 'get_window_metrics' by focusing on comparison rather than listing or retrieving metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives is provided. The description implies usage for comparing podlings across windows, but doesn't mention prerequisites, when-not-to-use scenarios, or how it differs from siblings like 'query_metric_rankings' or 'search_podlings' that might involve podling analysis. This leaves the agent without clear contextual boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_report_markdownB
Return the raw markdown for one podling report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool returns raw markdown but doesn't explain how the report is selected, if authentication is needed, potential errors, or response format details. This leaves significant gaps for a tool that likely involves data retrieval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without any wasted words. It's front-loaded with the core action and resource, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of retrieving a specific report (implied by 'one podling report'), no annotations, and no output schema, the description is incomplete. It doesn't explain how to specify which report, what the markdown contains, or error handling, leaving the agent with insufficient context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add param info, but that's acceptable here. A baseline of 4 is appropriate since the schema fully handles the lack of parameters without requiring compensation from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return') and resource ('raw markdown for one podling report'), making the tool's purpose understandable. However, it doesn't differentiate from sibling tools like 'get_report_summary' or explain what distinguishes 'raw markdown' from other report formats, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_report_summary' or 'search_podlings'. It lacks context about prerequisites, such as how to identify the specific podling report, or any exclusions, leaving usage unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_report_summaryB
Get parsed metrics for one podling report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but offers minimal behavioral insight. It implies a read operation ('Get') but doesn't disclose authentication needs, rate limits, error conditions, or what 'parsed metrics' entails (format, structure, or completeness). For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded with the core action and resource, making it immediately understandable without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a tool that presumably returns parsed metrics, the description is incomplete. It doesn't explain what 'parsed metrics' includes, how the podling report is identified, or the return format. For a tool in a metric-heavy context with multiple siblings, more detail is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the lack of inputs. The description adds value by specifying the resource ('one podling report'), implying it operates on a single, implicitly identified report. This contextual meaning goes beyond the empty schema, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('parsed metrics for one podling report'), making the purpose understandable. It distinguishes from siblings like 'get_report_markdown' (which likely returns raw markdown) by specifying 'parsed metrics', but doesn't explicitly differentiate from 'get_window_metrics' or other metric-related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_report_markdown', 'get_window_metrics', or 'list_metrics'. It doesn't mention prerequisites, context for 'podling report', or when this tool is preferred over other metric-retrieval options.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_window_metricsB
Return metrics for a single podling/window combination.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but only states what the tool returns without behavioral details. It doesn't disclose whether this is a read-only operation, potential errors, rate limits, or authentication needs, leaving significant gaps for a tool that likely queries metrics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose with no wasted words. It is appropriately sized and front-loaded, making it easy to understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters and no output schema, the description is minimally adequate but lacks completeness. It doesn't explain what metrics are returned, their format, or error handling, which are important for a metrics query tool with no structured output documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by specifying that metrics are for a 'single podling/window combination', which clarifies the scope beyond what the empty schema provides, justifying a score above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Return' and the resource 'metrics for a single podling/window combination', making the purpose specific and understandable. It doesn't explicitly distinguish from siblings like 'list_metrics' or 'query_metric_rankings', but the focus on a single combination provides some implicit differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description implies it's for a specific podling/window pair, but it doesn't mention prerequisites, when not to use it, or refer to sibling tools like 'list_metrics' for broader queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
health_overviewB
Return a high-level summary of the available Apache health reports.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool returns a summary but doesn't specify format, data freshness, rate limits, or authentication needs. This leaves critical operational details unclear for a tool that likely involves data retrieval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It's front-loaded with the core action and resource, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters and no output schema, the description adequately covers the basic purpose. However, for a health reporting tool in a server with multiple related siblings, it lacks context on output format or how it complements other tools, leaving gaps in overall understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing on the tool's purpose instead, which aligns well with the schema's simplicity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Return') and resource ('high-level summary of available Apache health reports'), making the purpose specific and understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_report_summary' or 'get_report_markdown', which might offer similar functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_report_summary' or 'list_metrics'. It lacks context about scenarios where a high-level overview is preferred over detailed reports, leaving the agent without usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_metricsB
Return the supported metrics and windows for querying.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool returns data but doesn't disclose behavioral traits like whether it's a read-only operation, if it requires authentication, rate limits, or what the return format looks like (e.g., list, object). This leaves significant gaps for an agent to understand how to handle the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any unnecessary words. It's front-loaded and wastes no space, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and no output schema, the description is minimally adequate but lacks depth. It doesn't explain the return values (e.g., structure of metrics/windows) or any behavioral context, which could be important for querying tools. However, the simplicity of the tool (0 params) means the description isn't severely lacking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add parameter details, which is appropriate here, and it implies the tool takes no inputs, aligning with the schema. A baseline of 4 is given since no parameters exist and the schema fully covers them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return') and the target ('supported metrics and windows for querying'), making the purpose understandable. However, it doesn't explicitly differentiate this tool from its siblings like 'get_window_metrics' or 'query_metric_rankings', which appear related to metrics/windows, so it doesn't reach the highest score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With siblings such as 'get_window_metrics' and 'query_metric_rankings' that might overlap in functionality, there's no indication of context, prerequisites, or exclusions for using 'list_metrics'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_podlingsB
List podlings that have a parsed markdown report.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the filter condition ('have a parsed markdown report') but doesn't disclose behavioral traits such as pagination, rate limits, permissions needed, or what happens if no podlings meet the criteria. For a tool with zero annotation coverage, this leaves significant gaps in understanding its operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any redundant information. It is appropriately sized and front-loaded, making it easy to understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, no annotations, and no output schema, the description is minimal but adequate for a simple listing tool. It specifies a filter condition, which adds some context, but lacks details on behavior, output format, or integration with siblings, leaving room for improvement in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by specifying the filter condition ('have a parsed markdown report'), which provides context beyond the schema, though it doesn't detail how this filtering is applied internally.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and resource ('podlings'), specifying that they must 'have a parsed markdown report'. This distinguishes it from generic listing tools by adding a filter condition, though it doesn't explicitly differentiate from sibling tools like 'search_podlings'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'search_podlings' or other siblings. The description implies usage for podlings with parsed markdown reports but doesn't specify exclusions, prerequisites, or comparative contexts with other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_metric_rankingsC
Rank podlings by one parsed metric for a specific window.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool performs ranking, which implies a read-only operation, but doesn't disclose behavioral traits like whether it requires authentication, has rate limits, returns paginated results, or what the output format looks like. This is inadequate for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose ('Rank podlings') and adds necessary qualifiers ('by one parsed metric for a specific window') without any wasted words. It's appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0 parameters, the description is incomplete. It lacks details on behavioral aspects like authentication needs, rate limits, or output format, and doesn't clarify how it differs from siblings. For a ranking tool with no structured data, this leaves significant gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds context by specifying 'by one parsed metric for a specific window', which implies inputs might be inferred from context or defaults, but since there are no parameters, a baseline of 4 is appropriate as it doesn't need to compensate for gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Rank podlings by one parsed metric for a specific window', which provides a clear verb ('Rank'), resource ('podlings'), and scope ('by one parsed metric for a specific window'). However, it doesn't explicitly differentiate from siblings like 'get_window_metrics' or 'list_podlings', leaving ambiguity about when to use this versus those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description mentions ranking by a metric for a window, but it doesn't specify prerequisites, exclusions, or compare it to siblings such as 'get_window_metrics' or 'list_podlings', leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_podlingsB
Search podling names by case-insensitive substring.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the search is 'case-insensitive', which is useful, but fails to describe other critical behaviors like response format, error handling, or performance characteristics. This leaves significant gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without any redundant or unnecessary information. It is front-loaded and appropriately sized for its purpose, earning a perfect score for conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is adequate but incomplete. It covers the basic search action but lacks details on output format or behavioral context, making it minimally viable but with clear gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description adds value by specifying the search mechanism ('case-insensitive substring'), which is not captured in the schema, justifying a score above the baseline of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search') and resource ('podling names') with a specific constraint ('by case-insensitive substring'), making the purpose evident. However, it does not explicitly differentiate from sibling tools like 'list_podlings', which could serve a similar listing function, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as 'list_podlings' for unfiltered listing or other search-related tools. It lacks context on prerequisites, exclusions, or typical use cases, offering minimal usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
compare_windows - First observed
get_report_markdown - First observed
get_report_summary - First observed
get_window_metrics - First observed
health_overview - First observed
list_metrics - First observed
list_podlings - First observed
query_metric_rankings - First observed
search_podlings
TDQS
Scored across 9 tools
Each tool has a clearly distinct purpose with no overlap: compare_windows compares podlings across windows, get_report_markdown retrieves raw markdown, get_report_summary provides parsed metrics, get_window_metrics gives metrics for a single podling/window, health_overview offers a high-level summary, list_metrics enumerates supported metrics/windows, list_podlings lists podlings with reports, query_metric_rankings ranks podlings by metric, and search_podlings searches podling names. The descriptions unambiguously differentiate each tool's function.
All tools follow a consistent verb_noun or verb_adjective_noun pattern using snake_case: compare_windows, get_report_markdown, get_report_summary, get_window_metrics, health_overview, list_metrics, list_podlings, query_metric_rankings, and search_podlings. The naming is predictable and readable throughout, with no deviations or mixed conventions.
With 9 tools, the count is well-scoped for the Apache health reporting domain. Each tool earns its place by covering distinct aspects such as listing, retrieving, comparing, searching, and ranking podling health data. This is neither too thin nor too heavy, providing comprehensive functionality without bloat.
The tool surface offers complete coverage for querying and analyzing Apache podling health reports. It includes listing and searching podlings, retrieving raw and parsed report data, comparing across windows, getting metrics and rankings, and providing overviews. There are no obvious gaps; agents can perform full workflows from discovery to detailed analysis without dead ends.
Maintenance
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP server for medicaid-intelligence
MCP server for the Seline Analytics API
MCP server for Withings health data — sleep, activity, heart, and body metrics.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server for accessing and analyzing Apache Software Foundation Incubator podling data from podlings.xml. It provides tools to query podling metadata, generate statistics, and analyze incubation trends over time.22MIT
- AlicenseAqualityCmaintenanceA dependency-free MCP server for Apache Incubator PMC oversight that helps identify podlings needing attention, assess graduation readiness, and generate podling briefings by combining lifecycle data and community health signals.212MIT
- AlicenseAqualityCmaintenanceAn MCP server for analyzing startup financial health and generating metrics reports locally.2MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for comprehensive PyPI package intelligence, providing tools for dependency analysis, security scanning, health scoring, license compliance, and trend tracking.MIT