Skip to main content
Glama

CI npm License: Apache-2.0

SuperGrep

Retrieval + inference offload for AI coding agents. Help your agents find answers faster and at lower cost, while keeping the same answer quality. Particularly useful for cheaper models.

Try it live at infino.ai/supergrep - put a question to a real codebase and watch the same model answer it with and without SuperGrep, side by side, with the bill for each. Supergrep uses Infino for retrieval, which tightly integrates small language models with indexes for scalable agent retrieval on object storage.

SuperGrep: find and plain sql on your machine, search, semantic sql and ask in the Infino cloud, one index in both places

SuperGrep offers agents four tools over an Infino index, kept in two places. find and plain sql run on your machine, over the keyword index. Conversely, anything with meaning in it - search, a sql statement with a ranked search inside it, and ask - runs in the Infino cloud, so heavy compute and embeddings stays off your laptop. Infino also embeds inference models in its service, so several small models work the index at once and hand your agent the rows they found. Your agent keeps the reasoning, writes the answer, and decides when to use SuperGrep. No configuration.

tool

what it does

find

Every line containing an exact string, like grep -n: complete, unranked, with the repo-wide total and per-file counts. Tens of milliseconds from the index.

search

One ranked pass fusing exact keyword matching with semantic similarity, so it works whether or not you know the words. Hits carry the code, cited path:line.

sql

Read-only SQL over the index. The ranked searches are table-valued, so "which files have the most code about X" ranks and tallies in one query.

ask

A question that spans the repository, handed to small models on the Infino service that run the investigation against the index and return the rows they found, cited path:line, rather than prose. Several asks run at once.

Search functions live inside SQL, so one statement asks a whole question

What people use it for

  • Code review. The reviewer's question is "what else calls this, and what breaks if it changes". find answers with every caller and the counts in one call, and ask reads the paths that matter, so the review is about the change rather than about finding it.

  • Big explorations. "How does X work end to end" spawns fifty look-ups. With SuperGrep they run as fifty asks in parallel against the index, instead of fifty subagents reading the tree into your bill.

  • Several codebases at once. Every local tool takes a repository path, so one session roams across every repo it touches - the service, the client, the shared library - without a checkout per question.

  • Code, logs and issues together. Index the logs, the test output and the issue export beside the source, and "why did this integration test start failing?" is one question over all of them.

Related MCP server: mcplens

Your agent reaches for it on its own

Offered both, the model reaches for SuperGrep

The model's own file tools stayed available throughout. What it had from SuperGrep was the server's instructions, which say which tool fits which kind of question; the choice on each call was the model's. The calls that are not SuperGrep are mostly Read: the model opens a file after the index has told it which one, rather than instead of asking. That is the shape you want - the index does the finding, and the model still opens what it needs to quote.

Go beyond code - index your entire laptop or any corpus

SuperGrep looks across all the files a question needs, not just the source. Logs, test output, stack traces, CI output, configuration and docs go in beside the code - .log, .out, .err, .jsonl and .ndjson are chunked at record boundaries, so a stack trace stays with the message that explains it - and the same four tools run over all of it: find for an exact stack frame, search for a failure you can only describe, sql to count and rank across a run, ask for the question that spans several of them at once.

That matters most where a frontier model is weakest. A log is the pathological case for a context window - large, repetitive, mostly irrelevant, and paid for again on every turn it stays in the transcript. An index collapses it to the spans that matter before the model sees any of it. On a public benchmark of CI-failure diagnosis, scored by the benchmark's own judge, that puts SuperGrep second on the score and first by a distance on score per token of context:

LogDx-CI: diagnosis score, context handed to the model, and score per 1k tokens, by method

For files that don't fit on your laptop - write them out to Parquet files in object storage and point SuperGrep at them without ever loading them onto your laptop (how). You can search them together with your code or laptop files using the same tools.

What it saves

The same model, the same 36 questions about a 256,000-line codebase, with and without SuperGrep. Every answer was checked against the code. Fully correct means every claim held and the whole question was answered. The bill is everything you pay: your model, its subagents, and SuperGrep.

Real agent runs through the Claude Agent SDK, the same minimal prompt in every arm, on the infino engine repository, measured 2026-09-24. Thirty-six questions in five categories: aggregation (10), comprehension (6), by meaning (6), pinpoint (8), known file (6). The file-tools arm is stock Claude Code: Glob, Grep, Read, LS, Bash, and the Agent tool with the built-in Explore subagent. The SuperGrep arm is the same plus the four tools above. The judge is Opus 5.5.

The same model with file tools and with SuperGrep, on four Claude models: bill, fully correct answers, time

Your mileage will vary with the model.

  • Cheaper on every model. Haiku 41% off the total bill, Sonnet 58%, Opus 26%, Fable 14%.

  • Quality increases on the cheaper models. Haiku gets eight more fully correct answers with SuperGrep than without. On Sonnet, Opus and Fable the answers are level: the judge is itself a model, and graded four times the same answers came back with 19, 20, 17 and 23 claims it could not verify, so a difference under about six answers in 36 is noise, and those three are inside it.

  • No surprise bills. On about a third of the questions, Sonnet with file tools sends a subagent off to read through the repository. That one question then costs four to five times as much and takes four times as long. With SuperGrep it asks the index instead. Over the 36 questions that is $3.66 against $8.73 and 20 minutes against 41, with 20 fully correct answers against 18.

  • On your own code the gap is wider. These runs are on a public, open-source repository, because that is a test anyone can repeat - and the large models have seen it in training, which is a head start for reading files. On a private codebase the model has never seen, the index does more of the work, and the effect of SuperGrep is larger.

Where it wins, and where it does not

Fully correct answers by kind of question, all four models together

It wins on questions about the whole codebase: counts, rankings, every occurrence. Grep gives the first forty matches; an index gives the total. It loses on finding one named thing with a large model, which reads whole files well.

The cheapest model with SuperGrep against the most expensive without

Haiku with SuperGrep against Opus and Fable with file tools: bill, fully correct answers, time

23 correct answers to their 25, for $1.09 instead of $6.32 (Opus) or $19.36 (Fable).

Install

Two commands, once per machine. You need Claude Code and node 22 or newer, on macOS or Linux.

claude plugin marketplace add infino-ai/supergrep
claude plugin install code-context@infino-ai

That is the setup. Open Claude Code in any directory - a repository, a folder of logs, your notes, anything - and ask a question. The first question indexes the directory, and find, sql and read answer from then on. Nothing is configured per directory, no account exists yet, and nothing has left your machine.

search and ask are the cloud half - the embeddings and the retrieval loop run on the Infino platform, over a copy of the index kept there - and they need an account. The first time a question needs them, Claude tells you the one command and does not run it, because it is yours to run. Once, in a terminal:

npx -y @infino-ai/code-context login --platform https://api.supergrep.infino.ai

A free account, no credit card required. There is no form, no email, no password and no card, and nothing is created without your say-so: the command tells you that the contents of the directories you use search and ask in will be uploaded to Infino, asks, and only on your yes creates the account and stores its key at ~/.infino/key, mode 600, readable only by you. No config file ever holds a key or a path to one. Infino is SOC 2 Type 2 certified.

Keep that key. Because the free account asks for no email and no card, the key is the only thing that identifies you: it is how you get back in, and nothing else can. Back it up somewhere safe. When you add your details in the Infino console the same account gains a sign-in, and keys can be managed from there.

Restart the session and every directory you open has all four tools. Each one gets its own database on your account, named after the directory and loaded the first time you use search or ask there - a directory you only ever find in is never uploaded. One server answers for every directory a session touches: the tools take a path, so a session that spans several projects names the one it means. When the free credit runs out, ask says so and tells you how to add billing details and a card to the same account; find and plain sql keep working throughout.

If you already have an Infino account

Sign in once per machine instead. The key comes from a file or standard input, never from an argument - argv is readable by every process on the machine - and --yes is the same agreement the sign-up asks for, since a piped key leaves no terminal to ask on:

npx -y @infino-ai/code-context login --db https://api.supergrep.infino.ai --yes < keyfile

Other MCP clients, and one entry per repository

Cursor, Codex CLI, Gemini CLI, Windsurf and Cline take the same server over stdio. cx install in a repository writes its entry into .mcp.json there (--config for another client's file); with the account stored it names the repository's database and nothing else, --local-only writes the keyword-only entry, and install --platform https://api.supergrep.infino.ai is the sign-up and the entry in one for a machine with no account. The reference has every flag, and CONTRIBUTING how to build from source.

Indexing it yourself

The first question in a directory indexes it and the MCP server keeps it current, so most of the time you never run an index by hand. When you want to - a first pass over a huge tree, a CI step, a corpus that is not a git repository - index is the command. (cx below is npx -y @infino-ai/code-context, or cx itself after npm install -g @infino-ai/code-context.)

cx index                      # bring the index up to date; incremental, full on first run
cx index ~/notes              # index some other directory
cx index --full               # force a full rebuild
cx index --watch              # keep watching the tree and sync on every change
cx index --no-embed           # keyword index only, skip the vector stage
cx index --max-files 1000000  # raise the cap past the 500,000 default; over it, the index is
                              # partial and says so, with the value to pass to get all of it

The index is plain files under .infino/ in the directory you indexed. Keyword search is live within seconds of the first cx index; semantic and hybrid search light up as the vectors finish backfilling behind it. cx status says what the index holds and how fresh it is.

Signed in, cx index loads the platform copy in the same pass - the directory's own database on your account, so ask sees the same content as find. To load a database you name instead:

cx index --db https://api.supergrep.infino.ai/<database>

--embed-provider platform (the default) has the platform fill that table's vectors with its own model, server-side; local embeds on this machine and ships the vectors instead.

Indexing from object storage

For a corpus too big for your laptop - years of logs, a document dump, anything you already keep in a bucket as Parquet or JSON - leave it there, point the platform at the bucket, and it builds the index from there. Nothing is copied, nothing is downloaded to your machine, and no row passes through your laptop or through the API.

1. Grant read on your bucket to Infino's service account - we give you its address - on the prefix you want indexed: roles/storage.objectViewer on GCS, s3:GetObject + s3:ListBucket on S3. Read only: the platform writes nothing there.

2. Submit the job. One POST, naming the bucket and prefix, and it returns straight away - the build runs on the platform, not in the request:

curl -sS -X POST https://api.supergrep.infino.ai/v1/hydrate/<database> \
  -H "authorization: Bearer $(cat ~/.infino/key)" \
  -H 'content-type: application/json' \
  -d '{
        "table": "logs",
        "source": { "kind": "bucket", "bucket": "<your-bucket>", "prefix": "exports/logs/" },
        "fts":    [ { "column": "message" } ],
        "embed":  { "column": "embedding", "source": ["message"] }
      }'

(A database that lives in your own bucket can also read shards staged under its own _source/ prefix: "source": { "kind": "prefix", "prefix": "_source/logs/" }.)

{ "job": "hydrate/<customer>/<database>/logs", "state": "pending" }

Leave fts and embed out and the job reads a sample and picks the roles itself. columns narrows which source columns are carried; no_embed: true builds no vector column at all.

3. Follow it. The reply carries the state, how far it has got, the schema it settled on, and what it has cost so far:

curl -sS "https://api.supergrep.infino.ai/v1/hydrate/<database>?table=logs" \
  -H "authorization: Bearer $(cat ~/.infino/key)"

States are pending, running, cancelling, stopped, succeeded, failed. A job that stopped - a cancel, or an outage that outlasted its budget - resumes from its own checkpoint rather than starting over:

# resume where it left off
curl -sS -X POST https://api.supergrep.infino.ai/v1/hydrate/<database> -H "authorization: Bearer $(cat ~/.infino/key)" \
  -H 'content-type: application/json' \
  -d '{"table":"logs","source":{"kind":"bucket","bucket":"<your-bucket>","prefix":"exports/logs/"},"resume":true}'

# stop a running job at its next commit boundary
curl -sS -X DELETE "https://api.supergrep.infino.ai/v1/hydrate/<database>?table=logs" \
  -H "authorization: Bearer $(cat ~/.infino/key)"

By default a job that fails for good drops its half-built table, so a partial table is never served; "on_failure": "keep" keeps what was committed.

The table is then searchable like any other. ask runs over it, and one question can span it and your code at once.

Note: Hydrate API needs to be enabled per account. Contact support@infino.ai to enable.

Learn more

License

Apache-2.0

Available Tools

3 tools
reindexSync the code indexA

Bring the index up to date with the working tree. Incremental by default: only files that changed since the last index are re-chunked and re-embedded, and an unchanged tree is a fast no-op, so call this freely after edits. The server also auto-syncs in the background as queries arrive. On a repo that has never been indexed this builds the index from scratch, replying as soon as keyword search is live (seconds) while vectors backfill behind it. Pass full=true to force a rebuild from scratch. Returns what changed plus index status.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoForce a full rebuild instead of an incremental sync.
pathNoAbsolute path to the repository root to index. Defaults to the server's configured root; set it to target a specific repo when a session spans more than one.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so admirably. It states that the operation is incremental by default, only changed files are re-chunked/re-embedded, an unchanged tree is a fast no-op, and auto-sync already occurs in the background. It also covers cold-start behavior (build from scratch, keyword search live in seconds, vectors backfill) and what a full=true does. This is rich, honest behavioral context with no contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact but information-dense, with each sentence contributing distinct behavioral or usage detail. It front-loads the core purpose in the first clause, then systematically covers incremental behavior, auto-sync, cold-start, the full flag, and return value. There is no filler or repetition, making it an efficiently structured paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description covers all the essential aspects a calling agent needs: default behavior, when to call, how to force a rebuild, cold-start timing, return information, and the path parameter's role is covered by the schema. It is complete for a synchronization tool and leaves no significant ambiguity about invocation or expected results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both parameters (full and path) are described in the input schema. The description adds a bit by saying 'Pass full=true to force a rebuild from scratch,' but that essentially restates the schema's 'Force a full rebuild instead of an incremental sync.' Since the schema already documents meaning, the description adds marginal value beyond it, keeping the score at the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Bring the index up to date with the working tree.' It immediately distinguishes itself from sibling tools (search, sql) by focusing on index synchronization. The description also explains the core behavior (incremental by default, full rebuild option) and makes it unmistakable what the tool accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: 'call this freely after edits' and notes that the server auto-syncs in the background 'as queries arrive.' This tells the agent when to invoke the tool, though it doesn't explicitly name alternatives or state when not to call it. The guidance is strong enough that an agent knows to use this after modifications, and it gets a 4 rather than a 5 because it lacks explicit 'use X instead' exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sqlSQL over the code indexA

Whole-repo analytical questions that file tools cannot express at any budget: counts, rankings, GROUP BY across the codebase in one query, on table chunks(path, start_line, end_line, lang, content[, embedding]). Search functions are callable as table-valued relations, so one query can rank AND aggregate: bm25_search('chunks','content','terms', k) needs no embedding; hybrid_search('chunks','content','terms','embedding', {{q}}, k) and vector_search('chunks','embedding', {{q}}, k) take a {{name}} placeholder with an embed map: {"q":"query text"}. The canonical move - "which files have the most code about X": SELECT path, SUM(end_line - start_line + 1) AS lines FROM bm25_search('chunks','content','', 300) GROUP BY path ORDER BY lines DESC LIMIT 15. Build queries on bm25_search/hybrid_search so results are ranked by relevance to the topic, not on a raw scan of the whole table. Read-only, single statement. The result includes a 'usage' field - a one-line receipt (tokens returned, rows, session total). After you answer, end your reply by showing that 'usage' line to the user verbatim.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoAbsolute path to the repository root to query. Defaults to the server's configured root; set it to target a specific repo when a session spans more than one.
embedNoMap of placeholder name → query text, embedded server-side. E.g. {"q":"vector indexing"} fills {{q}}.
queryYesA single read-only SELECT or WITH statement. May use search table functions and {{name}} vector placeholders.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It explicitly discloses that the tool is read-only, supports a single statement, and returns a 'usage' field with a receipt. It does not cover error behavior or permissions, but the core behavioral traits an agent needs are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and somewhat long, but every component earns its place: purpose, search function usage, placeholder semantics, canonical example, read-only note, and result receipt. It is front-loaded with purpose and then provides practical details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool, the absence of an output schema, and the lack of annotations, the description is remarkably complete. It covers the table structure, callable search functions, vector placeholder syntax, read-only behavior, and the usage receipt, leaving an agent with enough information to construct and invoke valid queries.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds substantial meaning beyond the JSON schema: it explains the {{name}} placeholder mechanism, gives an embed map example, provides a canonical whole-repo query, and explains when embeddings are needed versus bm25-only search. This is far beyond baseline schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as SQL querying over the code index for whole-repo analytical questions including counts, rankings, and GROUP BY. It gives a canonical example and explicitly contrasts with 'file tools', making its distinct role easy to grasp.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to use SQL: for whole-repo analytical questions that file tools cannot express, and advises building queries on bm25_search/hybrid_search instead of raw scans. It does not explicitly compare against the sibling 'search' tool, but the guidance is clear enough for most routing decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedreindex
    • First observedsearch
    • First observedsql

TDQS

A4.3/5.0

Scored across 3 tools

Disambiguation4/5

The three tools are clearly distinct: search for finding code, sql for analytical queries, reindex for index maintenance. There is minor overlap between search and sql since sql can also perform ranked searches, but the descriptions make the intended use cases clear.

Naming Consistency4/5

Tool names are simple lowercase verbs: search, sql, reindex. This is consistent in style, though 'sql' is a noun rather than a verb_noun pattern, and 'reindex' is a verb. Minor deviation but predictable.

Tool Count4/5

Three tools is on the lean side for a code-context server, but each tool serves a distinct and substantial purpose: search, SQL analytics, and index maintenance. The count is appropriate for a focused utility, though a get/read tool could be expected.

Completeness3/5

The server covers search, analytical querying, and index maintenance well, but lacks a direct file-read or file-content retrieval tool. Agents can work around this via search hits and sql, but a dedicated read tool would round out the surface.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    A
    maintenance
    An MCP server that provides structural codebase indexing and surgical query tools to drastically reduce token usage through symbol-level searches and transitive impact analysis. It supports multiple languages and integrates with git to help AI agents understand code dependencies and the impact of changes in sub-millisecond time.
    69
    1,160
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A local MCP server that provides AI coding assistants with semantic search capabilities over codebases. It indexes code using local embeddings and exposes tools for efficient code retrieval, saving tokens and improving response quality.
    31 npm
    4
    MIT
  • A
    license
    A
    quality
    F
    maintenance
    MCP server for semantic code search that indexes your codebase and allows AI editors to search using natural language queries.
    9
    7 npm
    53
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that indexes reference repositories and provides tools for AI coding agents to retrieve lossless code context, enabling reasoning over codebases larger than the agent's context window.
    8
    2
    Apache 2.0