contribos
Integrates with the GitHub API (via GITHUB_TOKEN) to search and score open issues, inspect repo policies, vouch lists and outside-PR merge rates, look up similar past PRs and reviewer comments, track competing PRs and failing CI on a PR, and build a public verified record of merged PRs, review responses and AI disclosure for a user.
Optional LLM provider (CONTRIBOS_LLM=openai with OPENAI_API_KEY) used to generate grounded plain-language summaries of the evidence ContribOS has already collected, with any claim citing files outside that evidence discarded.
ContribOS
Get trusted. Get merged. Come back. ContribOS helps new open-source contributors earn trust. It finds where you're welcome, shows how this repo wants the change done, checks that you truly understand your change, coaches you through review, and turns merged PRs into a record the next maintainer can verify.
In 2026, writing code is not the hard part of contributing. AI made PRs cheap, maintainers answered with PR caps, vouch lists and AI policies, and newcomer merge rates fell. ContribOS makes your PR worth a maintainer's time.
ContribOS never opens PRs, posts comments or claims issues for you, and it never writes your explanation. It points, and you decide.
The journey
Step | Command | What you get |
Find |
| Open issues sorted into Good bet / Possible / Skip, each with evidence: claimed? open PR? maintainer active? clear? Plus repo welcome: archived, vouch list, AI policy, outside-PR merge rate, how long outside PRs wait for a first human reply, how many stall, a stale bot, and your own open PRs there |
Rules |
| The repo's AI policy, disclosure format ( |
Propose |
| A short proposal to post before coding: likely files, a similar-size past PR, where the test goes, and one real question. Warns about taken issues, no maintainer yet, good-first-issue AI rules and issue-first rules. If the repo has a vouch list, it drafts the introduction first |
Tone |
| Flags machine-written tells (delve, buzzwords, assistant openers, filler), length, a missing question and leftover placeholders in anything you're about to post |
Agent rules |
| Teaches your coding agent this repo's rules: a SKILL.md for Claude Code, Codex, Copilot and Cursor, a hook that stops the agent opening PRs or posting comments, and a commit-msg hook for sign-off and the AI-disclosure trailer. All local-only, excluded from git |
Setup |
| Exact steps taken from the repo's CI and contributing guide: runtime version, install, services, env vars, test and lint commands |
Diagnose |
| What a failure means: missing dependency, wrong version, service down, env var, native build tools, flaky test, or a real test failure (and how to check it isn't yours) |
Understand |
| Files to read with the reasons each was picked, similar past PRs and what reviewers said, related tests, look-alike files to leave alone, house rules, and likely review questions. Add |
Prove |
| Pre-submit review: scope versus similar past changes, tests, changelog, sign-off, debug leftovers, AI policy and disclosure trailers, competing PRs for the same issue. |
Respond |
| Each reviewer comment classified (blocking, change, question, nit…), the code it points at, the house rule it echoes, a reply draft, what's still unanswered and for how long, and failing CI |
Grow |
| A public, linkable record: merged PRs, change requests, whether you answered every review comment, whether you explained your change and disclosed AI use |
Agents |
| Nine MCP tools for Claude Code, Cursor, Codex and others (policy, brief, setup, diagnose, check, review, find, claim, tone), taking text as input where an agent has text: |
Related MCP server: OpenCollab MCP
Install
pip install contribos # Python 3.10+, git. Zero external dependencies.
export GITHUB_TOKEN=... # needed for find, claim/brief from an issue URL, review, recordEverything that can run offline does: policy, brief --title, claim --title, setup, check, bench and review --data need only git. Add --offline to skip the API, and --update to fetch new commits. Clones and indexes are cached in ~/.cache/contribos (CONTRIBOS_CACHE).
Optional AI summaries: CONTRIBOS_LLM=anthropic (with ANTHROPIC_API_KEY) or CONTRIBOS_LLM=openai (with OPENAI_API_KEY), plus optional CONTRIBOS_MODEL. The model only explains evidence ContribOS already gathered. Any sentence citing a file outside that evidence is dropped.
How it works
repo.pyhandles cached clones. It reads any revision withgit showandgit grep, so no checkout is needed.precedent.pybuilds a SQLite index of main-line changes (squash and merge PRs), the files they touched, and the issues they fixed.policy.pyis the policy radar, built from contribution docs, AI policies, templates, vouch lists, CI config and history.brief.pyranks files by rare-term matches in code, path matches, and files touched by similar past PRs. It also finds related tests via co-change history.check.pyis the pre-submit check, the proof-of-understanding check and the PR draft.setup_doctor.pyproduces setup steps from CI and docs, and diagnoses failures.find.pychecks issue takeability, repo welcome and responsiveness.claim.pydrafts the proposal and vouch introduction.proof.pyruns a test with and without your change (files restored infinally).tone.pychecks text you're about to post.agent_rules.pywrites the local agent skill and hooks.review.pyclassifies review comments, drafts replies and tracks follow-through.record.pybuilds the contribution record (markdown plus a self-contained HTML page).llm.pyis the optional provider layer (Anthropic or OpenAI) with citation grounding.mcp_server.pyis a dependency-free MCP stdio server.github.pyis the optional API client. Every caller handles its absence.bench.pyruns the history benchmark.
Benchmark (2026-09-28, offline, PR titles standing in for issues)
Repo | Brief hit@5 | Grep-only hit@5 |
pallets/flask (30 PRs) | 0.77 | 0.73 |
pytest-dev/pytest (30 PRs) | 0.87 | 0.87 |
Matching against past PRs barely improves file finding, which supports the plan's bet that finding files is a commodity. The value is in how past PRs did it and what reviewers asked. The next benchmark needs real issue text and review comments, which requires GITHUB_TOKEN.
Status
Tested on real repos (Flask, pytest, Ghostty):
policy,brief,claim --title,setup,check,bench,mcp.Tested against the live GitHub API on 2026-10-01:
find(scored issue takeability and maintainer response medians onmaximilianfeix/proxy-scraperandbrekkylab/backlot),claim <url>(drafted pre-coding proposals with past similar PRs and file targets), andrecord(verified public merged PR portfolio for@adityatiwari101104).Tested on temporary git repos:
agent-rules(git status stays clean, the guard blocksgh pr create, the commit-msg hook enforces sign-off and the trailer),check --verify-test,tone, policy radar v2.Tests:
python -m unittest discover -s tests(40 tests, CI matrix on Linux, macOS, and Windows).
Releasing
To publish a new release to PyPI:
python -m build
twine upload dist/*
git tag v<version>
git push origin v<version>Available Tools
9 toolscontribos_briefC
Evidence-backed contribution brief for an issue: files to read with reasons, similar past PRs, related tests, look-alike files to avoid, house rules and likely review questions.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | owner/name, when giving title instead of issue_url | |
| title | No | ||
| issue_url | No | GitHub issue URL (needs API access) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It describes output content but omits safety profile (read-only vs. mutating), side effects, auth/API requirements, rate limits, and error behavior. This is a significant gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the purpose and then lists the evidence components. It is appropriately sized with no filler, though the packed list makes it somewhat heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description enumerates the return contents, which is useful given there is no output schema. However, it omits key input semantics (repo+title vs. issue_url choice) and API access prerequisites, leaving the agent without usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description says nothing about the three parameters (repo, title, issue_url). It does not explain when to use repo+title versus issue_url, nor does it mention that issue_url needs API access. The schema carries most of the semantic load and the description adds none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific deliverable: an 'Evidence-backed contribution brief for an issue' and enumerates what it contains (files, similar PRs, tests, look-alike files, house rules, review questions). It is clear what the tool produces, though it does not explicitly differentiate itself from siblings like contribos_review or contribos_find.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance or comparison to alternatives. The description only states what the tool outputs; it never indicates when an agent should choose this over contribos_review, contribos_find, or other sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_checkB
Pre-submit review of the local branch against the repo's rules and history: scope versus similar past changes, tests, changelog, sign-off, debug leftovers, AI policy, and whether the contributor's own explanation covers the change, AI-disclosure trailers, competing PRs, and (with verify_test) whether the test fails without the fix. Do not write the explanation for the user.
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | branch the work started from (default: auto-detect) | |
| path | No | local checkout | . |
| issue | No | issue number, to look for competing PRs | |
| ai_note | No | how AI was used, e.g. 'Claude Code: drafted the test' | |
| explanation | No | the contributor's own explanation, as text | |
| verify_test | No | test command to run with and without the fix | |
| explain_file | No | or a file holding it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose real behavior: the enumerated check domains, the conditional verify_test behavior ('whether the test fails without the fix'), and a constraint on the agent's own output. It never states whether the tool is read-only, whether it mutates the checkout, or how failures surface, which leaves notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, but the body is one long comma-chained sentence listing a dozen checks, which is hard to scan. Every clause is substantive, yet the structure is a run-on rather than a prioritized list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no output schema and no annotations, the description covers the check surface well but omits return shape, side effects, and safety profile. It is adequate to call the tool, but an agent cannot predict what it gets back or whether the working tree is touched.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds meaning by tying parameters to intent: issue maps to 'competing PRs', verify_test to the fail-without-fix check, ai_note/explanation to AI-disclosure and contributor-explanation evaluation. Only 'path' and 'base' are left purely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: a pre-submit review of the local branch against the repo's rules and history. It enumerates the checks concretely (scope, tests, changelog, sign-off, AI policy, competing PRs), so an agent knows what the tool does. It does not, however, differentiate itself from the sibling contribos_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Pre-submit review' implies the moment to use it, and 'Do not write the explanation for the user' is a useful boundary. But no alternative is named and neither contribos_review nor contribos_brief is contrasted, so the agent must infer which sibling applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_claimB
Draft the short proposal a contributor posts on an issue before coding (and a vouch introduction where the repo needs one). The human rewrites and posts it; never post it.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | owner/name, when giving title instead of issue_url | |
| title | No | ||
| issue_url | No | GitHub issue URL (needs API access) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that this only drafts and never posts and that a human rewrites and posts it, plus conditional behavior for a vouch introduction. However, it omits any note on return format, permissions, or the API access the schema hints at for issue_url.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the core purpose front-loaded, followed by the critical never-post constraint. The parenthetical about the vouch introduction is slightly crammed but still earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-annotation, no-output-schema tool, the description explains the draft-only nature adequately, but it leaves the ambiguous 'title' parameter and the repo/issue_url interaction unexplained and gives no sense of what the draft output looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%: repo and issue_url are documented, but 'title' has no description in either place, and the description adds nothing about how repo/title/issue_url interact. The schema does most of the work, so the baseline 3 holds rather than being compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Draft) and a concrete artifact (the short proposal a contributor posts on an issue before coding, plus a conditional vouch introduction). The purpose is clear, but it never names or contrasts with siblings like contribos_brief or contribos_policy, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'before coding' timing and 'The human rewrites and posts it; never post it' constraint imply when the tool is used, but there is no explicit when-not guidance or routing to alternatives among the eight sibling tools. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_diagnoseB
Explain a failed install or test run: missing dependency, wrong version, service not running, missing env var, flaky test, or a real test failure.
| Name | Required | Description | Default |
|---|---|---|---|
| log | No | the failing command's output | |
| output_file | No | or a file holding that output |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It never says whether the tool is read-only, whether it executes or merely reads the supplied log, what permissions it needs, or how it handles the two input sources — for a diagnosis tool with zero annotation coverage this is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the core verb and resource lead, and the enumerated failure modes are compact and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should say more about what comes back and whether the two optional inputs are alternatives or can be combined. The failure taxonomy is helpful, but the return behavior is left entirely implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (log, output_file) are already documented in the schema. The description adds no syntax, format, or precedence guidance beyond the schema's own wording, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Explain) and resource (a failed install or test run), then enumerates the failure categories it can diagnose, so the agent knows exactly what it produces. It does not, however, distinguish itself from the closest sibling, contribos_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the trigger condition 'a failed install or test run', but there is no explicit when-to-use versus alternatives such as contribos_check, and no stated prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_findC
Find open issues where a newcomer is actually welcome, with takeability evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| repos | No | ||
| topic | No | ||
| language | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions 'takeability evidence' but does not disclose authentication needs, rate limits, pagination, return format, or how evidence is gathered or scored, leaving behavior largely opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler, so it is appropriately concise. It is not fully efficient because a phrase like 'takeability evidence' is left unexplained, but structure is otherwise sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter description coverage, the description is far too thin for a three-parameter search tool. It gives a domain hint, but omits how to use the parameters, what the search covers, and what the agent gets back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has three parameters with 0% description coverage, and the description does not mention repos, topic, or language at all. An agent receives no semantic guidance on how to populate any parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Find') and resource ('open issues') qualified by newcomer welcome and takeability evidence, so an agent can understand the tool's intent. It does not explicitly distinguish itself from sibling tools, which keeps it at a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'where a newcomer is actually welcome' implies a search context, but there is no explicit guidance on when to choose this tool over siblings like contribos_review or contribos_check. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_policyA
What a GitHub repository expects from outside contributors: AI policy, vouch/trust gates, claim-first norms, tests, changelog, sign-off, activity. Every finding cites file:line. Call this before starting work on any repo.
| Name | Required | Description | Default |
|---|---|---|---|
| json | No | return machine-readable rules instead of markdown | |
| repo | Yes | owner/name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses useful output behavior (every finding cites file:line) and implies read-only retrieval, but says nothing about permissions, auth requirements, rate limits, or cost of the call. Adequate but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resource and its scope, followed by the key behavioral promise (file:line citations) and the invocation cue. The enumerated policy list is dense but each item carries information, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of conveying what comes back — the policy categories and the file:line citation guarantee. Combined with a fully documented input schema, an agent has enough to call this correctly, though return format nuances (markdown vs the json flag) are only covered by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (json, repo) are already fully documented in the schema; baseline is 3. The description adds no parameter-level meaning beyond what the schema provides — its detail is about the returned content, not the inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (a GitHub repo's contributor expectations) and enumerates the concrete policy areas it returns (AI policy, vouch/trust gates, claim-first norms, tests, changelog, sign-off, activity). It is clearly distinguishable from a generic tool, but it never names or contrasts itself with siblings like contribos_brief or contribos_check, which likely overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance: 'Call this before starting work on any repo,' which tells the agent exactly when this tool belongs in a workflow. It does not, however, clarify when to prefer a sibling such as contribos_brief or contribos_check instead, so the routing is only partially resolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_reviewA
Review coach for an open PR: classifies each reviewer comment, shows the code it points at, drafts replies, and lists comments still waiting on the author.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | or saved review data: author, review_comments, issue_comments, reviews | |
| path | No | local checkout, to show the code each comment points at | |
| pr_url | No | PR URL (needs API access) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavior: comments are classified and replies are only *drafted*, not posted. However, it omits auth expectations, whether the operation is read-only, what happens with partial input, and any failure/rate-limit behavior for the API-backed path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence identifies the tool as a 'Review coach for an open PR' and then lists its outputs after a colon. Every clause earns its place and there is no filler or repetition of structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-annotation, no-output-schema tool with three optional inputs and nested objects, the description covers the high-level outputs but leaves the input-mode contract, error handling, and return shape unexplained. It is adequate but leaves the agent to infer how to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains data, path, and pr_url, including that the API path 'needs API access'. The description adds only the indirect hint that these are alternative review sources and does not state mutual exclusivity or which combination is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (review) and resource (an open PR's reviewer comments) and enumerates four concrete capabilities: classify comments, show the pointed-at code, draft replies, and list comments awaiting the author. It is clear what the tool does, but it never differentiates itself from siblings like contribos_check or contribos_tone, which plausibly overlap in the review workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for an open PR' implies the usage context, and 'waiting on the author' hints at the triage scenario, but there is no explicit when-to-use vs when-not, no routing to a sibling tool, and no statement about which of the three input modes (pr_url, path, saved data) to prefer or whether they combine.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_setupC
Exact steps to set up and test a repository, taken from its CI and contributing docs.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses provenance ('taken from its CI and contributing docs'), but never states that this is a read-only extraction, what happens when a repo has no CI or contributing docs, or whether the steps are cached/static versus generated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the source-of-truth clause is appended efficiently. It is appropriately sized for a one-parameter lookup tool, though it could carry one more clause of routing or format guidance without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description covers the essentials of what is returned ('exact steps'). It omits repo identifier format, behavior on repos lacking CI/docs, and any relation to sibling tools, leaving clear gaps for an agent to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter 'repo' has no description in either place. The description's phrase 'a repository' hints at the value's meaning but gives no format (owner/repo slug, URL, or name), so it does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete deliverable — exact setup and test steps for a repository — and discloses the source (CI and contributing docs), which is more than a restatement of the name. It does not, however, differentiate itself from siblings like contribos_brief or contribos_diagnose, which could plausibly also touch repo onboarding context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus the eight sibling tools, nor any stated prerequisites (e.g., repo must be public, must have CI config). The agent must infer usage entirely from the name and the one-line purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribos_toneB
Check text the human is about to post for machine-written tells, length, and a clear question. Points at problems; the human rewrites.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the non-mutating, advisory behavior ('Points at problems; the human rewrites') — the tool flags but does not fix. However, it says nothing about permissions, return format, or what kinds of problems it surfaces.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no waste. The core action leads, and the behavioral caveat follows compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param check tool with no output schema and no annotations, the description hints at the advisory return ('points at problems') but leaves the 'kind' parameter and the output shape underspecified. Adequate but with clear gaps given the zero schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. 'text' is self-evident, but the 'kind' enum (comment/intro/reply/explanation) is never mentioned or explained. The check dimensions named relate to no parameter, leaving the enum entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Check text the human is about to post' and enumerates the check dimensions (machine-written tells, length, a clear question). It's clear what the tool does, though it doesn't explicitly distinguish itself from siblings like contribos_check or contribos_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Text the human is about to post' implies a pre-post timing context, but there is no explicit when-to-use, when-not, or named alternative among the many sibling tools (check, review, policy, diagnose). Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.2.1- First observed
contribos_brief - First observed
contribos_check - First observed
contribos_claim - First observed
contribos_diagnose - First observed
contribos_find - First observed
contribos_policy - First observed
contribos_review - First observed
contribos_setup - First observed
contribos_tone
TDQS
Scored across 9 tools
Each tool targets a distinct stage of the contribution workflow (find, policy, brief, setup, diagnose, claim, tone, check, review). Overlaps like check-vs-review and claim-vs-tone are resolved by the descriptions, which clearly bound one to pre-submit local work and the other to open PRs, and one to drafting versus checking text.
Every tool follows the identical contribos_verb pattern with a single lowercase word (policy, brief, setup, diagnose, check, review, claim, tone, find). No deviations in casing or verb style.
Nine tools is well within the ideal 3-15 range and each maps to a discrete workflow step with no filler. The count feels deliberately scoped to the contribution lifecycle rather than padded.
The surface covers discovery, repo-policy intake, briefing, setup, diagnostics, claiming, tone, pre-submit review, and PR review commentary. Deliberately omits automated posting (human stays in the loop), leaving only minor gaps such as a way to locate repos or matching to review history.
Maintenance
Related MCP Connectors
Open-source licence risk checks for AI coding agents and dependency trees.
Git-native policy layer for AI agents: check_action verdicts against rules approved via PR.
Screens public GitHub repos and PRs to generate risk maps, findings, and merge-readiness signals.
Governance copilot for AI-assisted coding. 72 packs, 532 rules, proof bundles.
Related MCP Servers
- AlicenseBqualityAmaintenanceOpen source contribution manager — tracks PRs across repos, discovers contributable issues, diagnoses CI failures, and drafts maintainer responses. 21 MCP tools, 5 resources, 3 prompts. Ships as CLI, MCP server, and Claude Code plugin.2017MIT
- AlicenseAqualityAmaintenanceEnables developers to find personalized open-source contributions by analyzing GitHub profiles and matching them with relevant 'good first issues' and beginner-friendly repositories. Provides comprehensive contribution tooling including repository health scoring, setup difficulty assessment, impact estimation, and automated PR planning.2223MIT
- AlicenseAqualityCmaintenanceEnables AI assistants to review GitHub and GitLab pull requests and merge requests, supporting multiple transports and custom review rules.728 PyPI5MIT
- AlicenseNot gradedqualityAmaintenancePre-commit norm gate for AI-generated PRs — enforces coding and contribution norms, blocking violations before they land.MIT