octocounts-mcp
Provides access to SLOC reports for public GitHub repositories, enabling agents and developer assistants to retrieve file, line, code, comment, and blank counts per language for any branch, tag, or commit.
OctoCounts – GitHub SLOC Counter
GitHub shows language bars. OctoCounts shows the actual line counts.
GitHub's sidebar shows language percentages, but misses actual file and line counts. OctoCounts adds this missing SLOC (Source Lines of Code) view to public repos without cloning.
Install the extension for instant stats directly on GitHub, or use the web app for public GitHub repositories. It downloads the repo archive, runs tokei, and caches the results—delivering a breakdown faster than git clone.
Use OctoCounts
Surface | Use it for |
Analyze any public GitHub repository and share a permanent report. | |
Show SLOC directly inside GitHub's repository sidebar in Chrome, Edge, or Firefox. | |
See aggregate report totals, largest repos, language coverage, and source breakdown. | |
Comment SLOC changes on pull requests. | |
Run | |
Give agents and developer assistants access to SLOC reports. | |
Add a live SLOC badge that links to a permanent report page. | |
Copy product descriptions, launch posts, links, screenshots, and badges. | |
Original research: a pilot study on how test/doc/generated file filtering affects SLOC counts. |
Related MCP server: github-mcp
Preview
Why?
Sometimes you just want to know whether a repo is 2k lines, 200k lines, or a weekend-devouring monolith. GitHub already has the repo, the language stats, and the sidebar. OctoCounts fills in the missing numbers.
What it does
Adds a browser extension card to GitHub repo pages with files, total lines, code, comments, blanks, and language count
Provides a web app where you can paste any public GitHub repo URL
Resolves any branch, tag, or commit SHA — pins results to an exact commit so the cache is actually meaningful
Downloads the GitHub archive tarball instead of cloning (much faster, no git history overhead)
Counts files, lines, code, comments, and blanks per language via tokei
Caches reports by
owner + repo + commit + tokei version— repeat runs are instantQueues analysis jobs so concurrent requests don't bring the server to its knees
Exports reports as plain text, JSON, or a shareable PNG card
Stack
Layer | Tech |
Backend | Rust · Axum · Tokio · SQLx · Postgres · tokei |
Frontend | React · TypeScript · Vite · TanStack Query |
Infra | Cloudflare Pages (Pages Functions SSR) · GHCR image + sloc-infra dual VPS (VPS-A/VPS-B) + Cloudflare Tunnel · Postgres on VPS-B (Neon frozen snapshot as fallback) · Docker Compose (local dev only) |
Badges
Drop a live SLOC badge into any README:
<!-- Default branch — full SLOC summary -->
[](https://octocounts.com/github/:owner/:repo)
<!-- Specific branch -->
[](https://octocounts.com/github/:owner/:repo/tree/:branch)
<!-- Specific tag (immutable, cached forever) -->
[](https://octocounts.com/github/:owner/:repo/tree/:tag)
<!-- Specific commit SHA (immutable, cached forever) -->
[](https://octocounts.com/github/:owner/:repo/commit/:sha)Add ?lang=<language> to any of the above to get a per-language badge instead:
<!-- Lines of code for a single language -->
[](https://octocounts.com/github/:owner/:repo)
[](https://octocounts.com/github/:owner/:repo/tree/:branch)Language names are case-insensitive (rust, Rust, and RUST all work). If a language is not found in the report the badge shows —. While a fresh analysis is running the badge shows ··· — most badge CDNs will retry automatically.
Cache behaviour | Header |
Default branch / branch |
|
Tag / commit |
|
API
POST /api/analyze { "repoUrl": "...", "refName": "main" }
GET /api/jobs/:id
GET /api/reports/:id
GET /api/stats
GET /badge/:owner/:repo
GET /badge/:owner/:repo/branch/:branch
GET /badge/:owner/:repo/tag/:tag
GET /badge/:owner/:repo/commit/:shaAll badge routes accept an optional ?lang=<language> query parameter that switches the response from the full SLOC summary badge to a per-language shields.io-style badge.
Running it
See how-to-run-and-deploy.md for local development setup, host-native instructions, GitHub token configuration, and production deployment.
Growth report inventory
Popular public repositories are listed in data/popular-repos.txt. The scheduled workflow seed-popular-repos.yml refreshes that curated inventory weekly.
refresh-github-trending.yml discovers GitHub's daily Trending repositories, validates and publishes the current snapshot to /trending, and pre-generates their stable /github/:owner/:repo SLOC reports. Trending provenance stays separate from real /popular access counts; no date-stamped archive pages are generated.
License
Available Tools
2 toolsanalyze_repoB
Analyze a public GitHub repository with OctoCounts and return SLOC totals.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Optional branch, tag, or commit SHA. | |
| repo_url | Yes | GitHub URL or owner/repo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, yet it only discloses that repos must be public and that output is SLOC totals. It says nothing about rate limits, authentication, behavior on private/invalid URLs, or whether analysis is synchronous or queued.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the resource and the return value front-loaded. 'with OctoCounts' is slightly internal jargon but does not waste meaningful space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with fully documented schema and no output schema, the description adequately conveys purpose, scope (public repos), and the returned metric. The remaining gap is behavioral edge cases (private repos, errors), which is a minor omission given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — both repo_url ('GitHub URL or owner/repo') and ref ('branch, tag, or commit SHA') are fully documented in the schema. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Analyze) and resource (public GitHub repository) plus the output (SLOC totals), so the agent knows exactly what it gets back. It does not explicitly distinguish itself from the sibling compare_repos, but the singular 'repository' vs. comparison framing makes the difference inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no mention of compare_repos as the alternative for multi-repo analysis. The agent must infer usage purely from the name and the word 'public'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_reposB
Compare SLOC totals between two public GitHub repositories or refs.
| Name | Required | Description | Default |
|---|---|---|---|
| left_ref | No | Optional left branch, tag, or commit SHA. | |
| right_ref | No | Optional right branch, tag, or commit SHA. | |
| left_repo_url | Yes | Left GitHub URL or owner/repo. | |
| right_repo_url | Yes | Right GitHub URL or owner/repo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses one useful constraint ('public'), but says nothing about authentication, network/clone cost, rate limits, or what the comparison returns (a delta, a winner, per-ref breakdown). For a mutating nothing but potentially expensive fetch tool, this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the scope qualifier 'between two' and 'SLOC totals' are efficiently packed in. It is appropriately sized, though it stops short of any elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description would ideally hint at the return shape (e.g., comparison of SLOC deltas), and with no annotations it should note access constraints. It covers the what and the scope but leaves the agent guessing on results and side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each of the four parameters (left/right repo URL plus optional left/right ref) is described in the schema, so the baseline 3 applies. The description adds no syntax, default-ref, or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Compare), resource (GitHub repositories/refs), and the exact metric (SLOC totals), which cleanly separates it from the broad-sounding sibling analyze_repo. It is clear without naming the sibling or describing output form.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: call it when you want to contrast two repos or refs. There is no explicit when-to-use vs analyze_repo, no stated prerequisites, and no note on what happens with private repos beyond the word 'public'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
analyze_repo - First observed
compare_repos
TDQS
Scored across 2 tools
analyze_repo targets a single repository, while compare_repos targets two repositories or refs. Their boundaries are clear and non-overlapping, so an agent can easily choose the right tool.
Both tools follow a consistent verb_noun pattern: analyze_repo and compare_repos. Naming is predictable and uses snake_case throughout.
Two tools is thin for most MCP servers, though this server has a narrow SLOC-counting purpose. The pair covers single-repo analysis and comparison, but feels minimal.
The core domain of SLOC totals and comparisons is covered by the two tools. Minor gaps exist around deeper breakdowns or output options, but agents can accomplish the primary workflows.
Maintenance
Related MCP Connectors
GitHub MCP — wraps the GitHub public REST API (no auth required for public endpoints)
Revternal MCP — wraps the Revternal Developer Intelligence API
GitHub Private MCP Pack — access private repos, org data via OAuth.
Docker Hub MCP — wraps the Docker Hub v2 API (free, no auth required for public data)
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides four GitHub superpowers (repo-scorecard, compare-repos, commit-heatmap, trending-mcp) as MCP tools with React widgets for repository health analysis, comparison, commit activity visualization, and trending MCP servers.91 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables GitHub repository analysis including health score, comparison, commit heatmap, and trending MCP servers, with interactive widgets.91 npmMIT
- AlicenseNot gradedqualityBmaintenanceProvides GitHub open-source intelligence via MCP tools for discovering trends, evaluating repository health, comparing projects, and analyzing users, issues, and releases. It includes a CLI and can be used with AI assistants like Claude or Cursor.MIT
- FlicenseNot gradedqualityCmaintenanceEnables read-only browsing and querying of GitHub repositories, including listing repos, reading files, fetching READMEs, searching code, and viewing commit history via MCP.-