Skip to main content
Glama

OpenAI Cookbook KG-RAG

A local retrieval toolkit that turns OpenAI Cookbook notebooks and Markdown into a searchable vector index and knowledge graph, then exposes the results through a CLI or MCP server.

The OpenAI Cookbook is an input corpus, not bundled source. This repository ships the ingestion, graph, retrieval, evaluation, CLI, and MCP implementation; cloned cookbooks, generated chunks, embeddings, caches, and SQLite databases remain local. Gemini is used for ingest-time embeddings and entity extraction, so live ingestion requires a Gemini API key. The smoke suite itself makes no live API calls.

Install and verify

Python 3.11 or newer is recommended. On Termux, --system-site-packages reuses native packages such as NumPy instead of attempting incompatible manylinux wheels.

python -m venv --system-site-packages .venv
. .venv/bin/activate
python -m pip install -r requirements.txt
python smoke_test.py

In a fresh clone, data-integrity checks are reported as skipped because generated indexes are intentionally absent. All unit and source-level smoke checks should pass.

Related MCP server: docs-search-engine

Use it

Set GEMINI_API_KEY in the environment or in a gitignored .env.local, then ingest a local checkout of the Cookbook or another notebook/Markdown directory:

git clone https://github.com/openai/openai-cookbook.git openai-cookbooks
python openai-cookbook-agent.py --ingest --source ./openai-cookbooks --limit 10
python openai-cookbook-agent.py --build-kg
python openai-cookbook-agent.py --query "how do I stream responses?" --no-interactive

Run python openai-cookbook-agent.py --help for the complete CLI. Generated state is written beneath sources/, embeddings/, kg/, and .cache/; all are ignored.

MCP configuration

After installation and ingestion, add the server to your MCP client, replacing the placeholder with the checkout's absolute path:

{
  "mcpServers": {
    "openai-cookbook": {
      "command": "/absolute/path/to/openai-cookbook-kg-rag/.venv/bin/python",
      "args": [
        "/absolute/path/to/openai-cookbook-kg-rag/openai-cookbook-agent.py",
        "--serve-mcp"
      ],
      "env": {
        "GEMINI_API_KEY": "set-this-in-your-private-client-config"
      }
    }
  }
}

The server provides tools for semantic search, code exemplars, knowledge-graph statistics, and cached retrieval.

License

Apache-2.0. The separately cloned OpenAI Cookbook retains its own license and is not redistributed by this repository.

Available Tools

4 tools
find_exemplarA

Find canonical code examples for an OpenAI API pattern. Searches Jupyter code cells specifically. Use for: 'show me how to do X', 'canonical example of Y'.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_kNoNumber of examples (default: 3)
patternYesPattern or concept (e.g. 'streaming response', 'function calling', 'structured output')

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It discloses the important scope limitation that it searches Jupyter code cells specifically, which is useful. However, it does not describe result ranking, return format, or behavior when no examples are found, leaving noticeable gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler: it states the core purpose, specifies the corpus scope, and gives direct usage examples. Each sentence earns its place and important constraints are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter search tool with no output schema, the description provides enough context to call it correctly: it names the input pattern, gives usage examples, and states the search domain. It could mention result count or ordering, but the schema already covers top_k, so the definition is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the input schema. The description repeats the notion of a pattern but adds no additional parameter-level semantics beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool finds canonical code examples for an OpenAI API pattern, and adds the specific scope of searching Jupyter code cells. This distinguishes it from sibling tools like search_kg and get_doc, which likely search broader knowledge or documentation sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit use-case examples ('show me how to do X', 'canonical example of Y'), giving the agent clear context for when to invoke this tool. However, it does not explicitly name alternatives or state when not to use it, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_docA

Retrieve a specific Cookbook document or chunk by its ID or partial source file path. Use after search_kg to read the full content.

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesChunk ID or partial source file path

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. 'Retrieve' implies a non-destructive read, and 'full content' clarifies the expected scope, but it does not disclose output format, error behavior, or unknown-ID handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first states the action and resource, the second gives the workflow context. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool, the description tells the agent what to pass, when to call it, and what to expect ('full content'). The absence of an output schema and any return-format hints is a minor gap, but the core call is well specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: doc_id is already documented as 'Chunk ID or partial source file path'. The description restates this same meaning without adding format examples, constraints, or clarification of what 'partial' means in practice.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Retrieve'), a clear resource ('Cookbook document or chunk'), and the lookup key ('ID or partial source file path'). It also differentiates itself from search_kg by framing this as the follow-up full-content reader.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Use after search_kg to read the full content' explicitly identifies the workflow context and intended sequencing. However, it does not mention exclusions or alternative tools like find_exemplar, so it stops short of full when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kg_statsA

Return corpus statistics: total chunks, KG nodes/edges, entity type breakdown, most-connected nodes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It does disclose what the tool returns ('total chunks, KG nodes/edges, entity type breakdown, most-connected nodes'), which conveys the behavior of producing aggregate counts and summaries. However, it does not disclose performance considerations, data freshness, or what 'most-connected' threshold is used, so transparency is reasonable but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the action ('Return corpus statistics') and then lists the specific statistics. Every item appears relevant and there is no filler or redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter statistics tool, the description covers the main use case and expected output categories. Without an output schema, the description should ideally clarify the response structure, and it doesn't fully specify the format of 'entity type breakdown' or 'most-connected nodes', but the enumeration of returned stats provides most of the needed context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, so there is no parameter meaning to add. The description appropriately explains what the tool does given its no-argument signature, making it clear that it is a zero-configuration aggregate query.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Return') and identifies a precise resource ('corpus statistics') with a clear list of content categories. It distinguishes itself from siblings like search_kg or get_doc by indicating it is a summary/aggregate operation rather than a retrieval or search function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when the agent needs corpus-level statistics rather than specific document or entity retrieval. However, it provides no explicit exclusions or comparison to siblings, so the usage guidance relies on inference from the listed return contents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_kgA

Search the OpenAI Cookbook knowledge base using hybrid vector + KG-walk retrieval. Returns ranked passages with source references. Best for: API usage patterns, model comparisons, code examples.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural language question about the OpenAI API or Cookbook
top_kNoNumber of results to return (default: 5)

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. It mentions the retrieval mechanism and the return type, but it does not disclose any side effects, permissions, rate limits, error handling, or limitations. For a search operation this is likely read-only, but that is not explicitly confirmed, leaving a significant transparency gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core action. The second sentence efficiently provides usage context. No filler or redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should explain return values thoroughly. It states 'ranked passages with source references' but does not specify result structure, ranking details, or behavior when no results are found. The description is adequate for basic usage but leaves gaps an agent might need to handle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters (query and top_k) having clear descriptions. The tool description itself adds no additional parameter meaning, so it falls at the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search), the resource (OpenAI Cookbook knowledge base), and the specific retrieval method (hybrid vector + KG-walk). It also indicates the output type (ranked passages with source references). This distinguishes it from siblings like get_doc or kg_stats by implying a general search over content, though it doesn't explicitly name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Best for: API usage patterns, model comparisons, code examples' provides clear contextual guidance on when to use this tool. However, it does not explicitly state when not to use it or mention alternative tools, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedfind_exemplar
    • First observedget_doc
    • First observedkg_stats
    • First observedsearch_kg

TDQS

A4/5.0

Scored across 4 tools

Disambiguation4/5

search_kg and find_exemplar both handle code-example lookups, which could cause misselection, but their descriptions clarify the distinction (general retrieval vs. canonical Jupyter-cell examples). get_doc and kg_stats are clearly separate.

Naming Consistency4/5

Three tools follow a verb_noun pattern (search_kg, get_doc, find_exemplar), but kg_stats is a noun phrase, a minor deviation. Style is otherwise consistent (lowercase underscores).

Tool Count5/5

Four tools fit the server's read-only knowledge-base purpose well: search, retrieve, specialized search, and stats. Each earns its place without redundancy or bloat.

Completeness5/5

The domain is querying the OpenAI Cookbook, and the core workflow (search, retrieve full content, find code examples) is fully covered. kg_stats adds useful corpus metadata, and no obvious dead ends remain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers