Tox Testing MCP Server
독성 테스트 MCP 서버
pytest를 사용하여 프로젝트 내에서 파이썬 테스트를 실행하기 위해 tox 명령을 실행하는 MCP 서버입니다. 이 서버는 모델 컨텍스트 프로토콜(MCP)을 통해 파이썬 테스트를 실행하고 관리하는 편리한 방법을 제공합니다.
특징
도구
run_tox_tests- 다양한 모드와 옵션으로 tox 테스트 실행다양한 실행 모드를 지원합니다:
all: 모든 테스트 또는 특정 그룹의 테스트를 실행합니다.file: 특정 파일에서 테스트 실행case: 특정 테스트 케이스를 실행합니다directory: 지정된 디렉토리의 모든 테스트를 실행합니다.
지원되는 테스트 그룹:
clients: 클라이언트 관련 테스트api: API 엔드포인트 테스트auth: 인증 테스트uploads: 업로드 기능 테스트routes: 경로 핸들러 테스트
Related MCP server: MCP Pytest Server
개발
종속성 설치:
지엑스피1
서버를 빌드하세요:
npm run build자동 재빌드를 사용한 개발의 경우:
npm run watch설치
VSCode와 함께 사용하려면 MCP 설정 파일( ~/.config/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json 에 서버 구성을 추가하세요.
{
"mcpServers": {
"tox-testing": {
"command": "node",
"args": ["/path/to/tox-testing/build/index.js"],
"env": {
"TOX_APP_DIR": "/path/to/your/python/project",
"TOX_TIMEOUT": "600"
}
}
}
}구성 옵션
env.TOX_TIMEOUT: (선택 사항) 테스트 실행이 완료될 때까지 대기하는 최대 시간(초)입니다. 테스트 실행이 이 제한 시간보다 오래 걸리면 종료됩니다. 기본값은 600초(10분)입니다.env.TOX_APP_DIR: (필수) tox.ini 파일이 있는 디렉터리입니다. tox 명령이 실행되는 위치입니다. 경로는 tox.ini 파일이 있는 Python 프로젝트의 루트를 가리켜야 합니다.
타임아웃은 특히 다음과 같은 경우에 중요합니다.
테스트 프로세스 중단 방지
장기 실행 통합 테스트 관리
CI/CD 파이프라인이 멈추지 않도록 보장
용법
서버는 다양한 모드에서 사용할 수 있는 단일 도구 run_tox_tests 제공합니다.
도구 인수
// Run all tests
{
"mode": "all"
}
// Run tests from a specific group
{
"mode": "all",
"group": "api"
}
// Run tests from a specific file
{
"mode": "file",
"testFile": "tests/test_api.py"
}
// Run a specific test case
{
"mode": "case",
"testFile": "tests/test_api.py",
"testCase": "test_endpoint_response"
}
// Run tests from a specific directory
{
"mode": "directory",
"directory": "tests/api/"
}Cline과 함께 사용
이 MCP를 Cline과 함께 사용하면 Cline의 사용자 지정 지침을 구성하여 테스트 실행을 효율적으로 처리할 수 있습니다. 권장되는 워크플로는 다음과 같습니다.
If asked to run tests on the project, use the tox-testing MCP. Follow these steps:
1. Run all tests across the project unless you are given instructions to run a specific test file or test case.
2. Review and rerun each failed test case individually as you troubleshoot and fix the issue from its output.
3. Repeat step 2 until the testcase passes.
4. Once all failed test cases from step 1 are passing rerun all tests again and repeat all steps until all tests pass.이 워크플로는 다음을 보장합니다.
모든 테스트를 먼저 실행하여 포괄적인 테스트 범위 제공
실패한 테스트 케이스를 분리하여 집중 디버깅
개별 사례를 다시 테스트하여 수정 사항 확인
모든 테스트를 다시 실행하여 최종 검증
클라인과의 상호작용 예시:
You: Run the tests for this project
Cline: I'll use the tox-testing MCP to run all tests:
{
"mode": "all"
}
You: Fix the failing test in test_api.py
Cline: I'll first run the specific test file:
{
"mode": "file",
"testFile": "tests/test_api.py"
}
Then address each failing test case individually:
{
"mode": "case",
"testFile": "tests/test_api.py",
"testCase": "test_endpoint_response"
}기여하다
행동 강령과 풀 리퀘스트 제출 프로세스에 대한 자세한 내용은 CONTRIBUTING.md를 참조하세요.
Available Tools
1 toolrun_tox_testsC
Run tox tests with different modes and options
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | Test execution mode | |
| directory | No | Directory containing tests to run (required for directory mode) | |
| group | No | Test group to run in all mode (defaults to clients) | |
| testFile | No | Specific test file to run (required for file and case modes) | |
| testCase | No | Specific test case to run (required for case mode) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool runs tests but fails to describe execution behavior, side effects, permissions needed, or output format. This leaves critical operational traits undocumented for a test-running tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is appropriately sized and front-loaded, making it easy to understand quickly with zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 parameters, no output schema, and no annotations), the description is incomplete. It does not explain return values, error handling, or behavioral nuances, leaving gaps that could hinder effective tool invocation in a testing context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'different modes and options,' which aligns with the parameters in the schema. However, with 100% schema description coverage, the schema already documents all parameters thoroughly. The description adds minimal semantic context beyond what the schema provides, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Run') and resource ('tox tests'), specifying the tool's purpose as executing tests with different modes and options. It distinguishes the tool's functionality but lacks explicit differentiation from siblings since none are provided, making it clear but not fully optimized for sibling comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, prerequisites, or contextual cues. It mentions 'different modes and options' but does not specify scenarios or exclusions, leaving usage entirely implicit and lacking actionable advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
run_tox_tests
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools, making disambiguation perfect. The single tool has a clear and distinct purpose that cannot be confused with any other tool in the set.
Since there is only one tool, naming consistency is inherently perfect as there are no other tools to compare it against. The tool name 'run_tox_tests' follows a clear verb_noun pattern, which is consistent with itself.
A single tool is generally too few for most server purposes, as it limits functionality and flexibility, making the server feel thin and underdeveloped. While it might suffice for a very narrow scope, it often indicates a lack of comprehensive coverage for the domain.
With only one tool, the server is severely incomplete for the domain of tox testing, as it lacks essential operations such as configuring environments, listing tests, checking results, or managing dependencies. This minimal surface will likely cause agent failures when more complex tasks are required.
Maintenance
Related MCP Connectors
An MCP server that provides access to Testiny projects, test cases and test runs
A MCP server built for developers enabling Git based project management with project and personal…
Official MCP server for Qase — manage test cases, runs, suites, defects via AI tools.
Related MCP Servers
- AlicenseAqualityDmaintenanceAn MCP server that enables AI assistants to discover and execute Nox sessions for project automation tasks like testing, linting, and building. It provides tools to list available sessions and run them using specific Python versions, tags, or keyword expressions.2MIT
- FlicenseNot gradedqualityDmaintenanceAn MCP-compliant server that enables the execution of pytest test suites and the storage of results into a QA platform database. It allows AI models to trigger test runs, track execution progress, and retrieve historical test data through specialized tool interfaces.1-
- AlicenseNot gradedqualityBmaintenanceThis MCP server enables automated maintenance and code analysis for Python/pytest repositories in isolated Docker environments. It supports read-only investigations, fix-and-verify tasks, and provides full audit trails with SQLite event history and artifact exports.MIT
- AlicenseDqualityDmaintenanceMCP server for deterministic local test execution and normalized test result reporting, supporting pytest and Jest with coverage summaries.6MIT