Skip to main content
Glama
CarmitHaas

Customer Service Data Analyst MCP Server

by CarmitHaas
README.md
# Customer Service Data Analyst Agent

![Python](https://img.shields.io/badge/python-3.11%2B-3776AB?logo=python&logoColor=white)
![LangGraph](https://img.shields.io/badge/LangGraph-ReAct-5e35b1)
![FastMCP](https://img.shields.io/badge/FastMCP-server-0e7c66)
![Nebius](https://img.shields.io/badge/LLM-Nebius%20Token%20Factory-1565c0)
![License: MIT](https://img.shields.io/badge/license-MIT-2e7d32)

A LangGraph **ReAct agent** that answers questions about the
[Bitext customer-support dataset](https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset)
(26,872 tagged support messages across 11 categories and 27 intents). It routes each
question, calls typed tools over the data, remembers the conversation and a per-user
profile across restarts, and exposes its tools over the Model Context Protocol.

It handles three kinds of question:

| Type | Example | What happens |
|---|---|---|
| **Structured** | "How many refund requests are there?" | chains tools → **997 (3.71%)** |
| **Unstructured** | "Summarize the FEEDBACK category." | samples rows → a grounded summary |
| **Out-of-scope** | "Who is the president of France?" | politely declined, never answered from general knowledge |

> Built by **Carmit Shaemesh Haas** for Nebius Academy Assignment 3.

---

## Demo

The CLI prints every step of the agent's reasoning. Below, it answers a question and then
resolves a follow-up ("what about cancellations?") — noticing a wrong filter and retrying:

![CLI demo](docs/cli_demo.gif)

The Streamlit UI shows the same reasoning in a chat, with a session switcher and the live
user profile in the sidebar.

| A structured question | The query recommender | An out-of-scope decline |
|---|---|---|
| ![structured](docs/streamlit_chat.png) | ![recommender](docs/streamlit_recommender.png) | ![decline](docs/streamlit_decline.png) |

---

## Architecture

The agent is a LangGraph graph. A dedicated **router** classifies every question *before*
any tool is chosen; out-of-scope questions are refused structurally (they never reach the
generation model's general knowledge). In-scope questions enter the ReAct loop, and a
profile-update step distills what it learned about the user.

![architecture](docs/architecture.png)

The compiled LangGraph itself (auto-rendered from the code):

![agent graph](docs/langgraph_flow.png)

The editable source for the system diagram is
[`docs/architecture.drawio`](docs/architecture.drawio).

**Pieces:**

- **Router** (`src/cs_agent/agent/router.py`) — labels a question `structured`,
  `unstructured`, `out_of_scope`, or `recommend`, using the small model with typed
  structured output (and a plain-text fallback).
- **Tools** (`src/cs_agent/tools/`) — five Pydantic-typed tools (`list_categories`,
  `list_intents`, `filter_records`, `count_records`, `summarize_category`) implemented as
  **pure functions** over a pandas DataFrame. The agent and the MCP server both call these
  same functions, so they can never drift apart.
- **Memory** — two kinds:
  - *Episodic*: a LangGraph **SqliteSaver** checkpoint per `--session`, so a conversation
    resumes after a restart and follow-ups ("what about refunds?") resolve.
  - *Semantic*: a per-user profile in `profiles/<user>.md` (name, interests, preferences),
    distilled after each answered turn and injected into the prompt.
- **Guardrails** — a `decline` node for out-of-scope questions and a graceful **fallback**
  after `MAX_ITERATIONS` (12) so the loop never spins forever.
- **MCP** — a FastMCP server (`mcp_server/server.py`) exposes the same five tools to any
  MCP client.

### Model choice

Both models run on **Nebius Token Factory** (OpenAI-compatible). The agent uses two, on
purpose:

| Role | Model | Why |
|---|---|---|
| Generation, tool calling, summaries, recommendations | `meta-llama/Llama-3.3-70B-Instruct` | reliable OpenAI-style function calling and grounded writing |
| Routing + profile distillation | `Qwen/Qwen3-30B-A3B-Instruct-2507` | a Mixture-of-Experts model with ~3B *active* parameters: much cheaper and faster than the 70B, and strong at short classification and merge tasks |

Routing and profile-merging are easy, high-volume jobs, so they go to the small fast model;
the heavier reasoning and writing go to the large one. Both IDs live in
`src/cs_agent/config.py` and can be overridden via `.env`.

---

## Quickstart (clone to running in ~5 minutes)

**Prerequisites:** Python 3.11+, a [Nebius Token Factory](https://tokenfactory.nebius.com)
API key, and [`uv`](https://docs.astral.sh/uv/) (recommended) or `pip`.

```bash
# 1. clone
git clone https://github.com/CarmitHaas/customer-service-agent-carmit-haas.git
cd customer-service-agent-carmit-haas

# 2. install (creates a venv and installs the package + deps)
uv sync
#   --- or with pip ---
# python -m venv .venv && source .venv/bin/activate
# pip install -e .

# 3. add your API key
cp .env.example .env
# edit .env and set NEBIUS_API_KEY=...

# 4. run the CLI
uv run python main.py --session demo --user carmit
```

On first run the dataset (~27k rows) is downloaded from Hugging Face once and cached to
`data/bitext.parquet`, so later runs start instantly and work offline.

---

## Using the CLI

```bash
uv run python main.py --session demo --user carmit
```

`--session` names the conversation (resume it later with the same value); `--user` selects
the persistent profile. Every tool call and observation is printed as it happens. Try:

```
How many refund requests are there?
What categories exist in the dataset?
What is the distribution of intents in the ACCOUNT category?
Summarize the FEEDBACK category.
What should I query next?          # the recommender: suggests, you confirm, it runs
What do you remember about me?      # answered from your profile
Who is the president of France?     # politely declined
```

To see memory survive a restart: ask something, `exit`, relaunch with the same
`--session`, and ask a follow-up like "what about shipping?".

## Using the Streamlit app

```bash
uv run streamlit run src/cs_agent/ui/streamlit_app.py
```

Chat in the browser; the reasoning steps appear in a collapsible panel and the sidebar has
the session switcher and the live profile.

---

## MCP server

Start the server (stdio transport):

```bash
uv run python mcp_server/server.py
```

Connect a client and call a tool. A runnable example is in
[`mcp_server/client_demo.py`](mcp_server/client_demo.py):

```python
import asyncio
from fastmcp import Client

async def main():
    async with Client("mcp_server/server.py") as client:
        tools = await client.list_tools()
        print([t.name for t in tools])
        result = await client.call_tool("count_records", {"intent": "get_refund"})
        print(result.data)   # {'count': 997, 'total': 26872, 'pct': 3.71, ...}

asyncio.run(main())
```

Run it directly:

```bash
uv run python mcp_server/client_demo.py
```

---

## Project layout

```
customer-service-agent-carmit-haas/
├── main.py                       # CLI entry point
├── src/cs_agent/
│   ├── config.py                 # endpoint, model IDs, paths, MAX_ITERATIONS
│   ├── data.py                   # cached dataset loader
│   ├── tools/
│   │   ├── schemas.py            # Pydantic input/return models + tool descriptions
│   │   └── analytics.py          # pure analysis functions (single source of truth)
│   ├── agent/
│   │   ├── state.py              # graph state
│   │   ├── llm.py                # Nebius model factories
│   │   ├── router.py             # query router node
│   │   ├── tool_bindings.py      # tools as LangChain @tool
│   │   ├── graph.py              # the LangGraph wiring
│   │   ├── profile.py            # per-user profile
│   │   └── persistence.py        # SqliteSaver checkpointer
│   └── ui/streamlit_app.py       # Streamlit chat (Bonus A)
├── mcp_server/
│   ├── server.py                 # FastMCP server (Task 3)
│   └── client_demo.py            # minimal MCP client
├── tests/test_analytics.py       # tool tests (no API key needed)
└── docs/                         # diagrams + screenshots
```

## Tests

```bash
uv run pytest
```

The tests cover the pure analysis tools against known dataset facts and need no API key.

---

## Notes

- Out-of-scope refusal is enforced **structurally** (a dedicated `decline` node), not just
  by a prompt instruction, so the model can't be talked into answering off-topic questions.
- The recommender proposes with a *no-tools* model, so it can suggest but never execute; a
  `pending_suggestion` flag makes the suggest → refine → confirm loop deterministic.

## License

MIT — see [LICENSE](LICENSE).

TDQS

A4.3/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a unique purpose: counting without rows, sampling with total count, listing categories, listing intents (optionally with distribution), and returning representative messages for summarization. No two tools perform the same function.

Naming Consistency5/5

All tool names follow the consistent verb_noun pattern using snake_case (e.g., count_records, filter_records). There are no deviations or mixed conventions.

Tool Count5/5

Five tools is well-scoped for the server's purpose of analyzing customer service data. Each tool serves a distinct, necessary function without redundancy or bloat.

Completeness4/5

The set covers discovery (categories, intents), counting, sampling, and summarization. A minor gap is lack of comparative or aggregate statistics across categories, but the core analytical workflow is supported.

Maintenance

ActivitySlowing
ResponsivenessNo issues