Skip to main content
Glama
README.md
# Intently

**Describe it badly. Find it anyway.**

Intently is an intent-first search engine for libraries of tools: AI models, apps, code libraries, MCP servers, or
anything else people look for by what it *does* rather than what it is *called*. People type vague, misspelled,
rambling requests; AI agents describe the problem they are stuck on. Intently works out what they mean, searches six
different ways at once, has an LLM check the shortlist, and answers with the tools that fit, why they fit, and where to
get them.

It ships as a small Python package with a web UI, a REST API, a Model Context Protocol (MCP) server for agents (stdio
and Streamable HTTP), and a one-command AWS Lambda deployment. It runs on a laptop CPU: no GPU, no vector database, no
search cluster.

> **Live demo: [ALFRED's AI TOOL BELT](https://alfreds.org/toolbelt)**, a library of about 4,900 AI tools, models
> and research projects built from 418 videos of the AI Search YouTube channel (2023 to 2026). Agents can use it
> directly: `claude mcp add --transport http alfreds-tool-belt https://alfreds.org/api/toolbelt/mcp`

## Contents

- [What it can do](#what-it-can-do)
- [Quick start](#quick-start)
- [Use it from AI agents](#use-it-from-ai-agents)
- [How it works](#how-it-works)
- [Why this architecture](#why-this-architecture)
- [Results](#results)
- [Benefits](#benefits)
- [Drawbacks and limits](#drawbacks-and-limits)
- [Build your own library](#build-your-own-library)
- [Configuration](#configuration)
- [Deployment](#deployment)
- [Evaluation](#evaluation)
- [Project layout](#project-layout)

## What it can do

| You type | Intently understands | And returns |
|---|---|---|
| `that thing where you give it a song and it spits out notes for every instrument` | audio-to-MIDI transcription with separate instrument tracks | the tools that do it, each with a fit grade and a one-line reason |
| `hunyan 3d` | the Hunyuan3D model family (misspelled) | every version (2.1, 2.5, 3.0), newest mentions first |
| `free image generator that runs on my 8gb laptop gpu` | image generation + constraints: free, local, 8 GB VRAM or less | tools that meet the constraints ranked above ones that don't |
| an agent's task: `turn a raw two-person interview into captioned vertical clips, open source, local` | four capabilities: noise cleanup, transcription with speaker labels, clip cutting, captions | tools grouped by capability, with links, install commands and fit grades |

Every result carries evidence: a deep link to the exact video moment where the tool is shown, or the project page for
catalog entries, plus official links (GitHub, Hugging Face, project sites) matched automatically.

Without any LLM keys it still works in offline mode (hybrid retrieval only), just with less understanding of vague
requests.

This is what an agent receives from `find_tools` on the live demo (markdown format, trimmed):

```markdown
# Tools for: Create a 3D model of a product using a few phone photos

## 1. 3D reconstruction from photos
- **AnyRecon** (tool app) | fit: exact | Open-source, local, free; handles multiple unaligned smartphone photos
  AnyRecon is a lightweight 3D reconstruction tool that takes a series of unaligned photos taken at random angles ...
  open source: yes · runs locally: yes · pricing: free · availability: released
  Links: https://yutian10.github.io/AnyRecon/
  Seen in: Robot girlfriends, recursive AI agents, full AI research, Happy Horse: AI NEWS (2026-05-03) https://www.youtube.com/watch?v=7r_WJ9xpne0&t=1022s
- **ShapeR** (model · Meta) | fit: exact | Open-source, runs locally, free; reconstructs from photo sequences
  ...
```

## Quick start

Requires Python 3.10+.

```bash
git clone https://github.com/Alphaxalchemy/intently
cd intently
pip install -e .
```

Build the example library (29 well-known open-source AI tools; no LLM or API key needed):

```bash
intently ingest catalog examples/catalog
intently build
```

The first build downloads the embedding model (about 90 MB). Then:

```bash
intently serve                    # web UI, API and MCP at http://127.0.0.1:8423
intently search "pull the vocals out of a song"
intently find-tools "transcribe a two person podcast with speaker labels, locally"
```

For smart mode (query understanding and LLM reranking), copy `.env.example` to `.env` and add at least one key. Free
tiers work: [Groq](https://console.groq.com) or [Gemini](https://aistudio.google.com). Anthropic, OpenAI (or any
OpenAI-compatible endpoint), Kimi and DeepSeek are also supported.

## Use it from AI agents

Intently is designed to be called by an agent in the middle of a task: "before I build this, is there a tool for it?"

**MCP, remote (Streamable HTTP):**

```bash
claude mcp add --transport http intently http://127.0.0.1:8423/mcp
```

**MCP, local (stdio):**

```bash
claude mcp add intently -- intently mcp
```

Any MCP client works the same way: point it at `/mcp` over HTTP, or run `intently mcp` as a stdio server.

| MCP tool | Use it when | Returns |
|---|---|---|
| `find_tools(task)` | the agent has a goal or problem and wants existing tools for it | capabilities the task needs, each with tools: fit (exact / strong / related / unverified), why, summary, open source / local / pricing / license, links, install and usage hints, where it was shown |
| `search_tools(query)` | the agent needs one specific tool it can describe but not name | matching tools with fit grades, links and mentions |
| `tool_details(tool_id)` | the agent wants everything about one result | aliases, links, versions, related tools, every mention |

**REST:**

```bash
curl "http://127.0.0.1:8423/api/find-tools?task=remove+background+noise+from+podcast+audio&format=markdown"
curl "http://127.0.0.1:8423/api/search?q=wisper+but+faster"
curl "http://127.0.0.1:8423/api/entity/fasterwhisper"
```

Every instance also serves `/llms.txt` (a plain-text guide for agents) and `/docs` (OpenAPI). [AGENTS.md](AGENTS.md)
is the full guide for agents, both for using Intently and for working on this code base, and
[skills/tool-belt/SKILL.md](skills/tool-belt/SKILL.md) is a ready-made agent skill that tells an agent when and how to
ask the live Tool Belt for tools.

## How it works

```mermaid
flowchart LR
  subgraph Index["Index time (once per new content)"]
    S1[Video transcripts] --> N[Normalize: documents, sections, passages]
    S2[Catalog entries YAML/JSON] --> N
    N --> E[LLM card per section: names and aliases, summary, capabilities, attributes, 6 search phrasings, relations]
    E --> R[Entity registry: aliases, versions, relations, official links]
    E --> B[BM25: cards and raw passages]
    E --> D[Dense vectors: cards, synthetic questions, raw passages]
  end
  subgraph Query["Query time"]
    Q[Request] --> A[1. Understand: intent, rewrites, hypothetical answer, names, constraints]
    A --> C[2. Retrieve: 6 channels x every query variant]
    C --> F[3. Fuse: weighted reciprocal rank fusion + constraint boosts]
    F --> K[4. Rerank: LLM grades top candidates 0-3]
    K --> G[5. Group by tool, attach links and evidence]
  end
  R -.-> C
  B -.-> C
  D -.-> C
```

**Index time.** Each source is normalized into documents (a video, a catalog entry), sections (a chapter, an entry)
and passages (about 220 words). For video transcripts, chapter start times are recovered from the description
timestamps, and caption fragments that spill across chapter boundaries are moved back. An LLM then writes one
structured card per section: canonical names plus aliases (including speech-recognition misspellings such as "RD" for
"Ardy"), a plain summary, capabilities in the words users search with, attributes (open source, runs locally,
pricing, license, minimum VRAM, availability), six realistic search queries and relations to other tools. Catalog
entries already contain most of this, so they need no LLM. From the cards Intently builds:

- an **entity registry**, a deliberately light graph: aliases merged, version families grouped ("Wan 2.1, 2.2, 2.5"),
  relations counted, official links matched from video descriptions;
- **BM25** indexes over card fields and over raw passages, for exact names and version numbers;
- **dense vector** indexes over cards, over every synthetic question, and over raw passages, computed locally with
  [`MongoDB/mdbr-leaf-ir`](https://huggingface.co/MongoDB/mdbr-leaf-ir) (23M parameters) through ONNX Runtime.

**Query time.**

1. **Understand** (one fast LLM call, about 1 s). The request becomes an intent sentence, three rewrites in the
   vocabulary a product page or transcript would use, a hypothetical answer (HyDE), corrected product names and
   constraints (open source, local, free, VRAM, approximate dates). Without an LLM, a heuristic repairs typos against
   the corpus vocabulary and extracts constraints with patterns.
2. **Retrieve** (about 100 ms, all local). Six channels run for every query variant: dense cards, dense synthetic
   questions, dense raw passages, BM25 cards, BM25 passages, and entity lookup (fuzzy alias match, then every version
   of the named thing).
3. **Fuse.** Weighted reciprocal rank fusion (k = 20, weights grid-tuned), then soft boosts and penalties for the
   stated constraints and for intro, outro and sponsor segments.
4. **Rerank** (one or two LLM calls, 1 to 2 s). A listwise call grades the top 24 candidates 0 to 3 against the
   intent, seeing each candidate's summary and its best-matching raw excerpt. A second pool is graded only if no exact
   match was found. **Recall safety:** anything the reranker did not confirm still follows as "unverified", so a lazy or
   wrong reranker cannot hide the right answer.
5. **Group** sections by their primary tool and attach deep links, official links and evidence.

**`find_tools`** adds a planning step: one LLM call splits the agent's task into one to four capabilities and writes
search variants for each. Each capability then goes through retrieval and reranking in parallel, on different
rate-limit buckets, under a latency deadline. A capability whose rerank is late keeps its fused ranking, marked
unverified, instead of stalling the agent.

Every stage degrades gracefully, and every response says which stages ran and which model did the work.

## Why this architecture

The problem has a specific shape. Queries are short or rambling descriptions of *one* item. The words in the query
often share nothing with the words in the source ("movement animation" vs "text-to-motion"). Names are misspelled
twice over, by users and by speech recognition. And the library is full of near-duplicates (versions, forks, rivals).
Each popular approach solves part of this:

| Approach | What it gets right | Why it is not enough on its own |
|---|---|---|
| Keyword search (BM25, SQLite FTS) | exact names and version numbers | cannot bridge vocabulary: 0.01 hit@3 on queries sharing no words with the answer |
| Plain vector RAG over raw text | paraphrases | raw transcripts describe tools differently from how people ask: 0.21 hit@3 on the same queries |
| GraphRAG, LightRAG | multi-hop questions, corpus-wide summaries | on single-item lookups independent benchmarks show plain retrieval plus reranking matching or beating it; indexing needs 10 to 15 thousand LLM calls for 2M words; entities merge by exact name, so misspellings become separate nodes; no keyword channel for exact names |
| RAG-Anything | parsing PDFs, images, tables, equations | nothing to parse in text transcripts; heavy install (2 to 4 GB) and a hard dependency on older LightRAG |
| Cross-encoder reranker only | cheap precision | only reorders what retrieval found; cannot reason about constraints or unseen 2025-2026 tool names |
| Fine-tuned embeddings | domain fit | needs labeled pairs and retraining whenever the library changes |

Intently combines the pieces that address each failure, following published evidence (details and sources in
[docs/research](docs/research)):

- **Write the index in the users' language at index time.** LLM-written cards and synthetic questions close the
  vocabulary gap from the document side (the idea behind Anthropic's Contextual Retrieval, doc2query and QuOTE). On
  our hardest queries, dense search over cards alone found the right tool in the top 3 about 50% more often than
  dense search over raw text (0.32 vs 0.21).
- **Rewrite the query too.** Intent sentences, rewrites and a hypothetical answer (HyDE) close the gap from the query
  side; this lifted hybrid retrieval from 0.27 to 0.49 hit@3 on the hard set.
- **Keep several signals and fuse them.** Dense, lexical and entity channels fail on different queries. Reciprocal
  rank fusion is robust and needs no score calibration; its weights were tuned on a dev split and accepted only
  because they also improved the held-out split.
- **Let an LLM make the final call, but never let it delete.** Listwise LLM reranking is where vague intent and
  constraints are actually resolved (0.49 to 0.73 hit@3 on the hard set). Recall safety keeps its mistakes visible.
- **Use a graph only where a graph helps.** The entity registry handles aliases, versions, relations and per-tool
  pages without GraphRAG's indexing cost or its blending of similar tools.
- **Stay small.** A 23M-parameter embedder on CPU, numpy arrays instead of a vector database, and at most a few
  thousand tokens of LLM calls per query: cheap enough for free API tiers and a serverless function.

## Results

Measured on the live demo's library (4,973 chapters from 418 YouTube AI-news videos), on a held-out test split. Two
query sets: **easy** (101 lazy, typo-ridden, rambling or half-remembered queries written from raw chapter text by a
different model family than the one that wrote the cards, plus real user queries) and **hard** (71 queries
machine-checked to share **zero** content words with the target chapter). "Hit@k" counts the right tool anywhere in the
top k, in any video.

| System | Hard hit@1 | Hard hit@3 | Hard hit@10 | Easy hit@1 | Easy hit@3 | Easy hit@10 |
|---|---|---|---|---|---|---|
| BM25 over raw transcripts | 0.00 | 0.01 | 0.03 | 0.66 | 0.84 | 0.93 |
| Dense over raw transcripts | 0.11 | 0.21 | 0.45 | 0.65 | 0.78 | 0.93 |
| Dense over LLM cards | 0.24 | 0.32 | 0.59 | 0.72 | 0.87 | 0.96 |
| Hybrid, no LLM (offline mode) | 0.11 | 0.27 | 0.55 | 0.81 | 0.91 | 0.96 |
| Hybrid + LLM query understanding | 0.34 | 0.49 | 0.82 | 0.81 | 0.92 | 0.96 |
| **Full pipeline (+ LLM rerank)** | **0.55** | **0.73** | **0.85** | 0.80 | 0.90 | **0.98** |

An LLM judge that saw only raw captions (never the cards) graded the top 3 of each system 0 to 3:

| System | Hard nDCG@3 | Hard: top result irrelevant | Easy nDCG@3 | Easy: top result irrelevant |
|---|---|---|---|---|
| Keyword search baseline (SQLite FTS5) | 0.13 | 66% | 0.68 | 9% |
| Hybrid, no LLM | 0.46 | 22% | 0.82 | 5% |
| **Full pipeline** | **0.86** | **3%** | **0.90** | **1%** |

Latency: about 2 to 3 s per smart search or `find_tools` call with Groq's free tier, about 0.1 s in offline mode.
The engine loads this index in about 2 s and uses about 400 MB of RAM.

## Benefits

- **Finds things that keyword and plain vector search cannot**, especially when the request shares no words with the
  answer (0.73 vs 0.01 to 0.27 hit@3 above).
- **Cheap to run.** CPU only. The whole index is flat files (about 180 MB for 5,000 sections) loaded into memory. Query
  cost is a few thousand LLM tokens, which free tiers cover for personal and small-team use.
- **Degrades instead of failing.** No LLM keys, a rate-limited vendor or a lazy reranker all produce weaker but still
  useful results, and the response says which stages ran.
- **Explains itself.** Each result carries a fit grade, a reason, the understood intent and constraints, and evidence
  links.
- **Agent-native.** MCP over stdio and stateless HTTP (works behind proxies and on serverless), compact markdown
  output, and a tool (`find_tools`) shaped around how agents actually ask.
- **Easy to grow.** Add a YAML entry or a transcript, run `ingest`, `enrich` and `build`; enrichment and embeddings are
  cached, so only new content costs anything.

## Drawbacks and limits

- **Index-time LLM cost.** About one call per six sections (the demo library took about 1,000 calls). Changing the card
  format means re-enriching.
- **Best quality needs an LLM at query time.** Offline mode is noticeably weaker on vague requests (0.27 vs 0.73 hit@3
  on the hard set). Free tiers cap daily volume, and beyond the cap a public instance falls back to offline ranking.
- **Latency.** 2 to 3 seconds in smart mode, versus about 100 ms offline.
- **Exact in-memory search.** Right for thousands to roughly a hundred thousand sections; much larger libraries need an
  approximate nearest-neighbor index and more memory.
- **Tuned for tools.** The default profile (categories, VRAM and license attributes, prompts) targets AI and software
  tools. Other domains need a new profile, and possibly different card fields and re-tuned fusion weights.
- **LLM-written metadata can be wrong.** Cards are instructed to say "unknown" rather than guess, raw passages remain
  a separate signal, and official links are matched heuristically. Check before relying on an attribute.
- **Evaluated on one library.** The numbers above come from one corpus and synthetic plus real queries; expect
  different numbers elsewhere. The evaluation harness is included so you can measure your own.
- **English-centric heuristics.** Typo repair and stemming assume English; the LLM stages handle other languages.

## Build your own library

Two source types ship with Intently. Both can live in the same library.

**Catalog entries** (YAML or JSON, one file or a folder). Only `name` and `summary` are required:

```yaml
- name: Basic Pitch
  developer: Spotify
  type: model
  url: https://github.com/spotify/basic-pitch
  summary: Lightweight neural network that transcribes audio of singing or instruments into MIDI notes.
  capabilities: [audio to MIDI, polyphonic note transcription]
  categories: [music_audio]
  open_source: yes
  runs_locally: yes
  pricing: free
  license: Apache-2.0
  install: pip install basic-pitch
  usage: basic-pitch <output-dir> <audio-file>
  alternatives: [Melodyne]
```

See [intently/sources/catalog.py](intently/sources/catalog.py) for every field and
[examples/catalog/ai-tools.yaml](examples/catalog/ai-tools.yaml) for 29 examples.

**Video transcripts** (Markdown, one `*_Transcript.md` per video, with a description block and `###` chapter
headings), the format written by common YouTube transcript downloaders. See
[intently/sources/transcripts.py](intently/sources/transcripts.py).

```bash
intently ingest catalog path/to/catalog        # repeatable; --name keeps several catalogs apart
intently ingest transcripts path/to/transcripts
intently enrich                                # LLM cards for new sections (cached, needs a key)
intently build                                 # indexes; only new text is embedded
```

**Profiles** hold everything corpus-specific: prompt wording, the category list and labels, example queries, UI text
and MCP instructions. Copy [intently/profiles/ai-tools.json](intently/profiles/ai-tools.json), edit it, and point
`INTENTLY_PROFILE` at it before building. The index keeps a copy of the profile it was built with.

## Configuration

Everything is an environment variable (a `.env` file in the working directory is read automatically).

| Variable | Default | Purpose |
|---|---|---|
| `INTENTLY_DATA_DIR` | `./intently-data` | corpus, cards, index, caches, model |
| `INTENTLY_PROFILE` | packaged `ai-tools` profile | corpus profile JSON |
| `GROQ_API_KEY`, `GEMINI_API_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY` (+ `OPENAI_BASE_URL`), `MOONSHOT_API_KEY`, `DEEPSEEK_API_KEY` | none | LLM vendors; any subset |
| `INTENTLY_VENDORS` | all with keys | restrict vendors, e.g. `groq,gemini` |
| `INTENTLY_CHAINS` | built in | JSON file overriding the model fallback chains (`query`, `rerank`, `bulk`) |
| `INTENTLY_EMBED_MODEL` | `MongoDB/mdbr-leaf-ir` | any sentence-transformers model that publishes `onnx/model.onnx` (rebuild after changing) |
| `INTENTLY_PUBLIC_URL` (or `_PUBLIC_API_URL`, `_PUBLIC_MCP_URL`) | none | the URLs agents should use, shown in the UI, `llms.txt` and tool output |
| `INTENTLY_RATE_LIMIT_PER_MIN`, `INTENTLY_RATE_LIMIT_PER_DAY` | off | per-client limits on search, find-tools, answer and MCP tool calls |
| `INTENTLY_SMART_DAILY_LIMIT` | off | LLM-backed requests per process per day before offline mode |
| `INTENTLY_PROXY_SECRET` | off | accept only requests from a reverse proxy that sends `X-Intently-Proxy-Secret`, and trust its `X-Forwarded-For` |
| `INTENTLY_FIND_TOOLS_DEADLINE` | `9` | seconds `find_tools` waits for reranking |

## Deployment

- **Any machine:** `intently serve --host 0.0.0.0 --port 8423` (FastAPI + uvicorn). Set the rate limits above for a
  public instance.
- **AWS Lambda:** [deploy/aws-lambda](deploy/aws-lambda) builds a Linux package without Docker and deploys a
  CloudFormation stack: a Function URL, the index in S3 (copied to `/tmp` on cold start, about 2.5 s), API keys in SSM
  Parameter Store, and a warm-up schedule. At low traffic the cost is close to zero. This is how the live demo runs,
  behind a Next.js route on alfreds.org that forwards `/api/toolbelt/*` with a shared secret.
- **Inside an existing website:** [deploy/static_page.py](deploy/static_page.py) renders the web UI as one static
  HTML file that calls the API under any path prefix, so a site can host the page and proxy the API (how
  alfreds.org/toolbelt is served).

## Evaluation

Put query files in `<data dir>/eval/queries_<set>.json`:

```json
[{"query": "that thing that turns a song into notes", "answer_entity": "Basic Pitch", "target": "tool-basic-pitch#0", "style": "vague_capability"}]
```

```bash
intently eval                                       # ablation ladder, hit@k and MRR per system
python -m intently.eval.judge --systems=hybrid,full # LLM-graded nDCG (use a different model family than the reranker)
python -m intently.eval.tune --llm                  # grid-tune fusion weights on dev, accept only if test improves too
```

Files named `queries_hard*.json` are reported separately. Re-run the evaluation after any change to retrieval.

## Project layout

| Path | What it does |
|---|---|
| `intently/sources/` | `transcripts.py`, `catalog.py`: raw material to normalized documents, sections and passages |
| `intently/enrich.py` | LLM cards (cached per group in `cards/`) |
| `intently/build_index.py` | entity registry, link matching, BM25 and dense indexes |
| `intently/engine.py` | query analysis, retrieval, fusion, rerank, grouping |
| `intently/toolbelt.py` | `find_tools`: task planning, parallel per-capability search, markdown output |
| `intently/mcp.py` | dependency-free MCP server (stdio and HTTP) |
| `intently/server.py` | FastAPI app: UI, REST API, `/mcp`, `/llms.txt`, rate limits |
| `intently/llm.py`, `keys.py` | vendor fallback chains, rate-limit cooldowns, response cache |
| `intently/embed.py`, `textindex.py` | ONNX embeddings, BM25, cosine search |
| `intently/ui/index.html` | the web UI (single file) |
| `intently/aws_lambda.py`, `deploy/aws-lambda/` | Lambda entry point, CloudFormation template, deploy script |
| `intently/eval/` | evaluation, LLM judge, weight tuning |
| `docs/research/` | the design research behind the architecture |
| `tests/` | offline test suite (`pytest`) |

## License

MIT, see [LICENSE](LICENSE). Created by Alfred Kolakkal.