Skip to main content
Glama
aleehamza25

Voice AI Website Analyzer MCP Server

by aleehamza25

Voice AI Website Analyzer MCP Server

An MCP (Model Context Protocol) server designed for Voice AI in GoHighLevel (GHL) that fetches and analyzes website content to provide business context to AI agents.

Features

  • Smart Website Crawling: Fetches homepage + up to 4 important pages (About, Services, Contact, etc.)

  • Business Details Extraction: Automatically extracts:

    • Business name and description

    • Contact information (phone, email, address)

    • Business hours

    • Services offered

  • Text Content Analysis: Provides comprehensive text summaries of each page

  • AI-Ready Output: Returns formatted text description perfect for AI agent context

Related MCP server: Crawl4AI RAG MCP Server

Installation

  1. Clone this repository:

git clone <your-repo-url>
cd Voice_MCP
  1. Install dependencies:

npm install
  1. Build the project:

npm run build

Usage

Running Locally

The MCP server runs on stdio transport:

npm start

Configuring in Claude Desktop or MCP Client

Add to your MCP client configuration (e.g., claude_desktop_config.json):

{
  "mcpServers": {
    "voice-ai-website-analyzer": {
      "command": "node",
      "args": ["d:\\Voice_MCP\\dist\\index.js"]
    }
  }
}

Using the Tool

Once configured, you can use the analyze_website tool:

analyze_website({ url: "https://example.com" })

The tool will:

  1. Fetch the homepage

  2. Identify and fetch up to 4 important pages (About, Services, Contact, etc.)

  3. Extract business details from all pages

  4. Return a comprehensive text analysis

Example Output

BUSINESS WEBSITE ANALYSIS
==================================================

Website: https://example.com
Pages Analyzed: 5

BUSINESS DETAILS
--------------------------------------------------
Business Name: Example Business Inc.
Description: We provide excellent services to our customers
Phone: (555) 123-4567
Email: info@example.com
Address: 123 Main Street, City, State 12345
Business Hours: Monday-Friday 9AM-5PM

SERVICES OFFERED
--------------------------------------------------
1. Web Development
2. Mobile App Development
3. Consulting Services
4. Technical Support

PAGE SUMMARIES
--------------------------------------------------

Page 1: Home - Example Business
URL: https://example.com
Content Preview: Welcome to Example Business...

Page 2: About Us
URL: https://example.com/about
Content Preview: Learn more about our company...

Deployment on Vercel

Option 1: Deploy via Vercel CLI

  1. Install Vercel CLI:

npm i -g vercel
  1. Deploy:

vercel

Option 2: Deploy via GitHub

  1. Push your code to GitHub

  2. Import the repository in Vercel dashboard

  3. Vercel will auto-detect the project and deploy

Important Note About Vercel Deployment

⚠️ MCP servers typically run on stdio transport and are designed to be run locally or on long-running servers. Vercel is optimized for serverless functions with HTTP endpoints.

For production use with GHL Voice AI, consider:

  1. Hosting on a VPS (Digital Ocean, AWS EC2, etc.) where the MCP server can run continuously

  2. Converting to HTTP API if you need serverless deployment

  3. Using Vercel for API endpoints and wrapping the MCP functionality in HTTP handlers

Converting to HTTP API (for Vercel)

If you need to deploy on Vercel, you'll want to create API endpoints instead. Let me know if you need help converting this to an HTTP API format.

Configuration

The server is configured to:

  • Fetch maximum of 5 pages total (1 homepage + 4 additional)

  • Extract text content (up to 5000 characters per page)

  • Identify important pages using keywords: about, services, contact, products, portfolio, team

  • Extract common business information patterns

Development

Project Structure

Voice_MCP/
├── src/
│   └── index.ts          # Main MCP server implementation
├── dist/                 # Compiled JavaScript (generated)
├── package.json
├── tsconfig.json
├── vercel.json
└── README.md

Building

npm run build

Testing Locally

After building, run:

node dist/index.js

The server will start and wait for MCP protocol messages on stdin.

Integration with GHL Voice AI

When integrated with GHL Voice AI:

  1. The AI agent receives the website URL from user input during conversation

  2. The agent calls the analyze_website tool with the URL

  3. The MCP server fetches and analyzes the website

  4. The business context is returned to the AI agent

  5. The AI agent uses this context to provide tailored responses about the business

Requirements

  • Node.js 18 or higher

  • TypeScript 5.x

Dependencies

  • @modelcontextprotocol/sdk: MCP protocol implementation

  • cheerio: HTML parsing and manipulation

  • node-fetch: HTTP requests

License

MIT

Support

For issues or questions, please open an issue in the repository.

Available Tools

1 tool
analyze_websiteA

Fetches and analyzes a website to extract business context. Crawls the homepage and up to 4 additional important pages (About, Services, Contact, etc.). Extracts text content, business details (hours, contact info, services), and URLs. Returns a comprehensive text description of the business.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe website URL to analyze (e.g., https://example.com)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses crawling behavior (homepage + up to 4 additional pages), extraction of text and business details, and output format. This provides good behavioral insight, though it doesn't mention potential issues like login walls or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, concise, and front-loaded with the primary action. Every sentence adds value—crawl scope, extracted details, and output summary. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single parameter, lack of output schema, and no siblings, the description is complete. It explains what the tool does, how it behaves (crawling strategy), and what it returns, covering all necessary contextual information for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage with a clear description of the 'url' parameter. The tool description adds no additional meaning beyond what the schema states, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Fetches and analyzes a website') and the purpose ('to extract business context'), with specific details about crawling up to 5 pages and extracting content types. No siblings exist, so no differentiation needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for analyzing websites to get business context, but provides no explicit guidance on when to use this tool versus alternatives. Since no sibling tools exist, the lack of alternatives is acceptable, but it doesn't address when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedanalyze_website

TDQS

A4.2/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility for confusion or overlap. The tool's purpose is singular and clear.

Naming Consistency5/5

The single tool uses a clear verb_noun pattern ('analyze_website'), which is consistent and descriptive. With one tool, naming consistency is trivially maintained.

Tool Count3/5

A single tool is on the lower end for a server claiming to be a 'website analyzer,' but it covers the core functionality of crawling and extracting business context. It feels thin but is arguably appropriate for a focused tool.

Completeness4/5

The tool comprehensively analyzes a website by crawling key pages and extracting business details. Minor gaps might include handling of dynamic content or pagination, but the main workflow is covered.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables web content scanning and analysis by fetching, analyzing, and extracting information from web pages using tools like page fetching, link extraction, site crawling, and more.
    6
    13
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to crawl websites, extract and store web content with semantic search capabilities using vector embeddings, and retrieve information through natural language queries with tag-based filtering and intelligent content cleaning.
    -