Dine-Discover-AI MCP Server
Serves as an alternative LLM provider for the agent: an OPENAI_API_KEY can be supplied in place of the Groq key to run the tool-routing model (gpt-oss-20b) that decides which retrieval tools to call, and the stronger reasoning model (gpt-oss-120b) that writes the grounded, cited final answer and optionally acts as the listwise reranker and LLM judge in evaluation.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Dine-Discover-AI MCP ServerRecommend a cozy Italian restaurant in Los Angeles."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Dine-Discover-AI
Dine-Discover-AI is a conversational RAG agent that acts as an expert guide to California's restaurant scene. Ask in plain language — "cheap tacos near the beach", "a calm zen place for sushi", "what did reviewers say about The Gilded Artichoke?", "show me a margherita pizza" — and the agent retrieves the relevant restaurants, reviews and photos from a vector database, then writes an answer grounded only in what it retrieved.
It is built on the Model Context Protocol (MCP): retrieval is exposed as MCP tools, and a ReAct agent decides which tools to call.
Features
Hybrid retrieval (RAG): dense vector search (Jina embeddings in ChromaDB) plus BM25 keyword search, fused with weighted Reciprocal Rank Fusion (weights tuned on a dev split).
Cross-encoder reranking:
jina-reranker-v2-base-multilingualreorders the top 20 candidates (about 0.8 s). The oldergpt-oss-120blistwise reranker is still available withRERANKER=llm.Grounded generation with citations: a fast model (
gpt-oss-20b) routes tool calls, and a stronger model (gpt-oss-120b) writes the final answer from the retrieved context only, citing sources as[n].RAG trace panel: under each answer, the UI shows which MCP tools were called, their arguments and latency, the retrieval method, and the numbered sources the answer cites.
Metadata filters: location, maximum price ($–$$$$) and minimum rating. Price and rating filters are applied inside Chroma's
whereclause.Multimodal search:
jina-clip-v2puts text and images in one vector space. You can search for photos by description, or upload a food photo to find similar dishes.Semantic fallbacks: misspelled or descriptive restaurant names still resolve, via embedding search.
Measured quality, end to end: retrieval (Recall@k, MRR, nDCG@10 with bootstrap CIs), agent routing (tool and argument accuracy), and answers (faithfulness, hallucinated names, citation validity, declining out-of-corpus requests). See Evaluation results.
Related MCP server: Agentic Commerce MCP Demo
Architecture
┌──────────────── offline: python src/build_index.py ────────────────┐
data/processed/*.json │ chunk → embed (Jina v3 / jina-clip-v2) → ChromaDB (data/chroma/) │
data/raw/*.txt, *.png └────────────────────────────────────────────────────────────────────┘
Browser ──► Gradio UI (src/app.py) ── MCP host + ReAct loop
│
├──► gpt-oss-20b (Groq) — decides which retrieval tools to call
├──► gpt-oss-120b (Groq) — writes the grounded final answer
│
└──► MCP server (src/server.py, stdio subprocess)
├─ recommend_by_vibe → HybridRetriever (dense + BM25 → weighted RRF → Jina rerank)
├─ search_knowledge_base → dense search over culinary-map prose + reviews
├─ search_images → jina-clip-v2 text→image search
├─ get_restaurant_info → name match, semantic fallback
├─ get_review → itemId join, semantic fallback, photo URLs
└─ resource: culinary-map://californiaRetrieval pipeline (src/retrieval.py)
Dense (restaurants): the query is embedded with
jina-embeddings-v3(task=retrieval.query) and matched by cosine similarity against one document per restaurant, which were embedded withtask=retrieval.passage.Dense (culinary map): the same query is matched against the prose paragraph written about each restaurant, and hits are mapped back to the restaurant by
itemId.Sparse: BM25 over each restaurant's document plus its prose paragraph.
Fusion: weighted Reciprocal Rank Fusion,
score = Σ wᵢ/(60 + rank), over the three ranked lists. Weights(0.25, 0.25, 2.0)for (dense restaurants, dense culinary map, BM25) were chosen by grid search on the Set B+C dev split (python src/evaluate_retrieval.py --tune, saved ineval/fusion_tuning.json). BM25 alone scored highest on that split, but it was not chosen, because Sets B and C are written from the record text and so favour exact word overlap; the chosen weights are the best setting that keeps all three retrievers.Filters: location,
max_price,min_rating.Rerank: the Jina cross-encoder scores the top 20 against each restaurant's record plus its prose paragraph. With
RERANKER=llm,gpt-oss-120breranks the list instead. If the rerank call fails, the fused order is kept.
Vector index (src/build_index.py)
Chunking is entity-level: one chunk per restaurant record, per culinary-map paragraph and per review. Every source item is a single short paragraph, so it is never split further. Other strategies (merged record + prose chunks, per-field vectors) have not been evaluated yet.
Collection | Chunks | Content | Embedding model |
| 210 | one document per structured record + filter metadata |
|
| 210 | one prose paragraph per restaurant (by itemId) |
|
| 10 | review text + photo captions |
|
| 118 | 109 recipe photos + 9 review photos |
|
Indexing is incremental: every chunk stores a content hash, so a rerun only re-embeds chunks that changed. Adding, editing or deleting a restaurant with restaurant_data_management.py updates the index straight away.
Project Structure
Dine-Discover-AI/
├── .env # API keys (not committed)
├── requirements.txt
├── src/
│ ├── app.py # Gradio UI + ReAct agent (MCP host)
│ ├── server.py # MCP server: RAG tools + resource
│ ├── client.py # MCP smoke test for all tools (no chat LLM)
│ ├── config.py # env, model names, paths
│ ├── embeddings.py # Jina embeddings client (text + multimodal)
│ ├── build_index.py # ingest → chunk → embed → ChromaDB
│ ├── retrieval.py # hybrid retriever, RRF, LLM reranker
│ ├── evaluate_retrieval.py # retrieval metrics
│ ├── evaluate_answers.py # answer quality: LLM judge + deterministic checks
│ ├── evaluate_agent.py # tool-routing and argument-extraction accuracy
│ └── restaurant_data_management.py # CRUD CLI with LLM extraction + index sync
├── eval/ # evaluation queries and results
└── data/
├── raw/ # culinary map text, recipe photos
├── processed/ # structured restaurants, reviews, recipes
└── chroma/ # persisted vector index (generated)Setup & Installation
Clone the repository:
git clone https://github.com/FenilKaneria/Dine-Discover-AI.git cd Dine-Discover-AICreate a virtual environment (recommended):
python -m venv .venv .venv\Scripts\activate # Windows source .venv/bin/activate # macOS / LinuxInstall dependencies:
pip install -r requirements.txtAdd API keys to a
.envfile in the project root:GROQ_API_KEY=gsk_... JINA_API_KEY=jina_...OPENAI_API_KEYis also accepted in place ofGROQ_API_KEY. The Groq endpoint is fixed insrc/config.py, so no base URL is needed.Do not wrap values in quotes, and never leave a quote unterminated.
python-dotenvsilently skips a malformed line and the line after it.Build the vector index (one-off; takes about a minute and a half on the Jina free tier):
python src/build_index.py # incremental python src/build_index.py --rebuild # re-embed everything
Usage
python src/app.pyOpen http://127.0.0.1:7860. Photos found by the agent show up in the chat. Open "How this answer was built (RAG trace)" to see the MCP tool calls and sources behind the last answer, and "Find similar dishes from a photo" to search by image.
Optional environment overrides:
Variable | Default | Effect |
|
| Tool-routing model and Set B/C query generator. |
|
| Answer generation, LLM reranker, LLM judge. |
|
| Text embedding model. |
|
| Multimodal embedding model. |
|
|
|
|
| Jina reranker model. |
|
| Set to |
|
| Set to |
|
| Set to |
MCP smoke test (spawns the server over stdio, checks the 5 tools and the resource, then calls each tool once):
python src/client.pyRestaurant database CLI:
python src/restaurant_data_management.py # interactive CRUD (add uses the LLM; syncs the index)
python src/restaurant_data_management.py --test # offline unit testsEvaluation
A RAG agent can fail at three stages, so each one is measured separately:
Stage | Question | Script |
Retrieval | Are the right restaurants in the top results? |
|
Agent routing | Does the router pick the right MCP tool and extract the right filters? |
|
Generation | Is the answer grounded, cited, and does it decline when nothing matches? |
|
python src/evaluate_retrieval.py --tune # grid-search fusion weights on the dev split
python src/evaluate_retrieval.py # retrieval metrics → eval/results.{json,md}
python src/evaluate_agent.py # routing accuracy → eval/agent_routing.{json,md}
python src/evaluate_answers.py # answer quality → eval/answer_quality.{json,md}Query sets
Set A: rule-labelled requests (40 queries). Hand-written requests such as "cheap mexican food". A restaurant counts as relevant when its structured fields match a rule (cuisine/dish/vibe regex, location, price, rating). Many restaurants are relevant per query, so Precision@5 and nDCG@10 are the useful columns.
Set B: known-item paraphrases.
gpt-oss-20bwrites one request per restaurant without using its name; the target is the only relevant item.Set C: vague known-item. Like Set B, but with no names, places, dishes or drinks.
Set D: hand-written, human-style known-item (25 queries) plus 15 out-of-corpus requests (
eval/queries_d.json). These are not generated from the record text, so they check that the numbers hold up beyond LLM-written queries. The out-of-corpus requests ("Chicago deep-dish pizza", "kosher deli") have no answer in the data. The only correct response is to say so.Dev/test split. Sets B and C are split 50/50 (seed 42). Fusion weights are tuned on dev only, and every table reports the test half.
Confidence intervals. Bracketed values are 95% bootstrap intervals (1000 resamples). Differences between systems are paired bootstrap deltas on the same queries; if the interval contains 0, the difference may be noise.
Agent routing (30 messages,
eval/queries_agent.json). Name lookups, reviews, vibe searches with and without filters, open questions, image requests and small talk. Checks the first tool call and the extracted arguments, for example "under $$" →max_price=2, and no invented filters.Answer quality (20 in-corpus + 10 out-of-corpus). Each query goes through the production tool, context builder and answer prompt.
gpt-oss-120bjudges faithfulness, relevance and whether the answer declined. Two checks need no LLM: restaurant names in the answer that are not in the retrieved context, and whether each[n]citation points to a real source.
Evaluation results
Measured on 2026-09-27 over the 210-restaurant corpus. Raw numbers are in eval/. Latencies are end-to-end from a laptop and include the Jina API call that embeds the query (about 0.75 s).
Retrieval: nDCG@10 on each set [95% CI]
System | Set A (40) | Set B test (103) | Set C test (104) | Set D (25) | p50 latency |
keyword (original) | 0.336 | 0.062 | 0.054 | 0.126 | <2 ms |
BM25 | 0.767 [0.70–0.84] | 0.986 | 0.936 | 0.958 | <1 ms |
dense (Jina v3 + Chroma) | 0.558 | 0.906 | 0.512 | 0.896 | 0.8 s |
hybrid, equal weights (before) | 0.753 | 0.951 | 0.742 | 0.972 | 0.8 s |
hybrid, tuned weights | 0.782 | 0.978 | 0.897 | 0.968 | 0.8 s |
hybrid + Jina rerank (production) | 0.817 [0.75–0.88] | 0.980 | 0.957 | 1.000 | 1.5–1.7 s |
Known-item Hit@1 for the production system: Set B 0.961, Set C 0.904, Set D 1.000. Text-to-image (jina-clip-v2, 109 queries over 118 images): Recall@1 0.826, Recall@5 0.982, MRR 0.899. Full per-metric tables are in eval/results.md.
Paired deltas, production minus BM25 (nDCG@10): Set A +0.051 (CI 0.001 to 0.103), Set B −0.006 (−0.018 to 0.002), Set C +0.021 (−0.010 to 0.054), Set D +0.042 (0.000 to 0.086).
Agent routing (30 messages, router gpt-oss-20b)
Tool accuracy | Argument accuracy |
100% (30/30) | 100% (20/20) |
The first run scored 90% on arguments. The router set location="beach" for "near the beach" and ignored an explicit "$". The recommend_by_vibe docstring now says what counts as a location filter and how $ maps to max_price. That fix was made against these same 30 messages, so 100% is an optimistic figure. A held-out routing set is still to do.
Answer quality (20 in-corpus + 10 out-of-corpus, generator and judge gpt-oss-120b)
Metric | First run | After prompt fix |
Faithfulness, mean (1–5) | 4.53 | 5.00 |
Fully faithful answers | 57% | 100% |
Answers with judge-flagged unsupported claims | 13 / 30 | 0 / 30 |
Answers with hallucinated restaurant names | 0 / 30 | 0 / 30 |
Valid | 100% | 100% |
In-corpus answers that cite sources | 100% | 95% |
Answer relevance, in-corpus (1–5) | 5.00 | 4.95 |
Out-of-corpus requests correctly declined | 90% | 80% |
In-corpus requests wrongly declined | 5% | 10% |
The first run never invented a restaurant. Its faithfulness losses were small embellishments: adjectives or dish preparations missing from the context, often the user's own wording repeated back as a fact about the restaurant. The answer prompt now forbids that. The trade-off is a more cautious model: a few more answers say "not mentioned". (Citation counts treat 【n】, which gpt-oss sometimes writes, as [n]; the app normalizes it the same way.) With 10 out-of-corpus queries, 90% → 80% is a single query.
Is it good enough?
Target | Result | Met? |
Production retrieval ≥ BM25 on every set | better on A and D; equal within noise on B and C | yes |
Rerank latency p50 < 1.5 s end-to-end | 1.48–1.70 s (was 17.6 s with the LLM reranker) | borderline |
Faithfulness ≥ 4.5 and ≥ 85% fully faithful | 5.00, 100% | yes |
No hallucinated restaurant names | 0 / 30 | yes |
Tool-routing accuracy ≥ 90% | 100% (tuned on the test set, see above) | yes, with caveat |
Out-of-corpus requests declined ≥ 80% | 80% | just |
What the numbers say
BM25 is still the strongest single retriever on this corpus. The records are short and full of distinctive words, and Sets B and C are written from those records. Tuning on the dev split picked BM25 alone as the best setting. The chosen weights keep dense retrieval, which costs about 0.04 nDCG@10 on Set C (significant). In return, dense covers queries that share no words with the record. On the hand-written Set D, hybrid and BM25 are level.
Tuning the fusion weights fixed most of the vague-query gap. On Set C, the tuned hybrid beats the equal-weight hybrid by +0.156 nDCG@10 (CI 0.110 to 0.210).
The cross-encoder reranker is where the quality comes from. It lifts Set C Hit@1 from 0.789 (tuned hybrid) to 0.904 and gives a perfect score on Set D. It is about 10× faster than the LLM reranker, and it does not use Groq quota.
Generation is grounded but cautious. No invented restaurants in either run. The remaining weak spot is judgement at the edges: about 1 in 10 answers declines a request the data could serve, and about 1 in 5 out-of-corpus requests still gets a "closest match" instead of a clear "not available".
Not yet measured / next steps
Held-out agent-routing set (the current one was used to fix the docstring).
Larger answer-quality sample with confidence intervals; 30 answers is enough to catch big problems, not to separate close variants.
Multi-turn conversations (follow-ups such as "cheaper than that?").
Chunking ablations: record and prose merged into one vector, and per-field vectors.
hybrid + LLM rerankon the new splits (python src/evaluate_retrieval.py --llm-rerank), for comparison with the Jina reranker.
Data Notes
structured_restaurant_data.json: 210 restaurant records (name,location,type,food_style,rating,price_range,signatures,vibe,environment,shortcomings,itemId). Only 175 names are unique, because some chains appear in several cities;itemIdtells them apart.California-Culinary-Map.txt: one prose paragraph per restaurant, in the same order as the JSON records. The index verifies each link with a name check.augmented_user_review.json: 10 reviews, linked byitemId, with photo URLs and captions.augmented_food_recipe.json+synthetic_recipe_images/: 109 recipes with photos, searchable throughsearch_images.
This server cannot be deployed
Maintenance
Related MCP Connectors
AI-native restaurant discovery: verified/menu-indexed/discovered tiers + signed allergy-safety data.
Agent-native registry: 168k+ real restaurants in LA, Hong Kong & Tokyo. Unranked, honest signals.
Bounded tools for rendering, extraction, RAG, enrichment, local discovery and review analysis.
- daxoomOAuthcom.daxoom
Owner-verified local business data for AI agents: profiles, hours, prices, with provenance.
Related MCP Servers
- AlicenseAqualityFmaintenanceAn AI-powered server that helps users discover and book restaurants based on location, cuisine preferences, mood, and event type, with integration to Google Maps Places API for accurate recommendations.516MIT
- AlicenseNot gradedqualityDmaintenanceEnables interactive restaurant discovery and ordering through a synthetic commerce flow with rich HTML UI. Demonstrates agentic commerce UX with tools to find restaurants, view menus, place mock orders, and generate receipts.4MIT
- FlicenseNot gradedqualityNot gradedmaintenanceAn AI-native restaurant discovery service that enables searching and receiving natural language recommendations for over 2,200 restaurants across 15+ US cities. It provides tools for accessing detailed restaurant info, curated lists, and cuisine-specific searches through the Model Context Protocol.-
- FlicenseDqualityDmaintenanceEnables interaction with a restaurant FastAPI backend for menu management, order processing, customer management, and analytics through natural language.241-