Scholar MCP
Enables searching for academic papers on Google Scholar by keywords and author names, with features for paginated results and filtering by specific year ranges.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Scholar MCPfind recent papers on large language models from 2024"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Scholar MCP
Local MCP server that searches Google Scholar. Scrapes results with requests + BeautifulSoup -- no API keys, no paid services.
Tools
search_papers_by_topic-- search by keywords, optional year range, paginatedget_author_papers-- find papers by author name, paginated
Related MCP server: semantic-scholar-mcp
Install
Clone and install:
git clone https://github.com/ProPriyam/Scholar-MCP.git
cd Scholar-MCP
pip install -e .Or run directly without cloning (needs uv):
uvx --from git+https://github.com/ProPriyam/Scholar-MCP scholar-mcpClient setup
All configs use python -m scholar_mcp.server to start the server. This avoids PATH issues that pip install can cause on Windows.
VS Code
Add to .vscode/mcp.json:
{
"servers": {
"scholarMcp": {
"type": "stdio",
"command": "python",
"args": ["-m", "scholar_mcp.server"],
"env": {
"PYTHONUNBUFFERED": "1"
}
}
}
}OpenCode
Add to opencode.json in your project root:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"scholar_mcp": {
"type": "local",
"command": ["python", "-m", "scholar_mcp.server"],
"enabled": true,
"environment": {
"PYTHONUNBUFFERED": "1"
}
}
}
}Claude Code
claude mcp add --transport stdio --scope project scholar-mcp -- python -m scholar_mcp.serverConfiguration
All optional. Set as environment variables.
Variable | Default | Description |
| Chrome-like UA | User-Agent header for requests |
|
| HTTP timeout in seconds |
|
| Minimum delay between requests (seconds) |
|
| Retry attempts on failure |
|
| Backoff multiplier between retries |
| none | HTTP/HTTPS proxy URL |
|
| Max results per request |
Notes
This scrapes Google Scholar HTML. It can break if Google changes their markup or blocks requests.
Available Tools
2 toolsget_author_papersB
Get papers filtered by author name from Google Scholar search.
| Name | Required | Description | Default |
|---|---|---|---|
| author | Yes | ||
| page_size | No | ||
| cursor | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It successfully identifies Google Scholar as the external data source, but fails to explain pagination behavior (how cursor/page_size interact), rate limiting, or result ordering. It provides minimal viable context but lacks operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and efficient, avoiding redundancy. However, given the 0% schema coverage and presence of pagination parameters, the extreme brevity becomes a liability rather than a virtue, as necessary parameter documentation is sacrificed for conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core retrieval scenario adequately, and the existence of an output schema absolves it from explaining return values. However, with zero schema descriptions and pagination complexity, the failure to explain cursor-based pagination mechanics leaves a significant gap in contextual completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, requiring the description to compensate. While it implies the 'author' parameter through 'filtered by author name,' it completely omits explanation of the pagination parameters 'cursor' and 'page_size,' leaving critical implementation details undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (papers), the filtering mechanism (by author name), and the data source (Google Scholar). It implicitly distinguishes from sibling tool 'search_papers_by_topic' by specifying 'author name' as the filter criterion, though it could be stronger with an explicit comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through 'filtered by author name,' suggesting when to use the tool (when seeking papers by a specific author). However, it lacks explicit guidance on when NOT to use it or direct comparison to the sibling 'search_papers_by_topic' for topic-based queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papers_by_topicC
Search Google Scholar papers by topic keywords.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| page_size | No | ||
| cursor | No | ||
| year_min | No | ||
| year_max | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure but only identifies the data source ('Google Scholar'). It fails to mention pagination behavior, rate limits, authentication requirements, result ordering, or whether results are cached.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence contains no redundant words and immediately communicates the core function. However, the extreme brevity contributes to under-specification given the tool's complexity (5 parameters with pagination and filtering).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having 5 parameters (including pagination and date filtering) and an output schema, the description is minimal. It omits critical context for effective use such as pagination mechanics, date filter applicability, and how results are ranked or formatted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, requiring the description to compensate. It only implicitly hints at the 'query' parameter via 'topic keywords' but provides no explanation for 'cursor' (pagination), 'page_size', or the year range filters, leaving significant semantic gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (Search), resource (Google Scholar papers), and method (by topic keywords). The phrase 'by topic keywords' implicitly distinguishes this tool from the sibling 'get_author_papers', though it does not explicitly name the alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus 'get_author_papers', nor does it explain pagination strategy (cursor parameter) or when to apply date filters (year_min/year_max). No prerequisites or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools have clearly distinct purposes: one filters papers by author name, while the other searches by topic keywords. There is no overlap in functionality, making it easy for an agent to choose the correct tool based on the query.
Both tools follow a verb_noun pattern (get_author_papers, search_papers_by_topic), which is consistent and readable. The minor deviation is in the prepositional phrase 'by_topic' vs. no preposition in the first name, but this does not hinder clarity.
With only 2 tools, the server feels thin and under-scoped for a Google Scholar integration. Key operations like retrieving paper details, citations, or author profiles are missing, limiting the server's utility for comprehensive scholarly research.
The toolset is severely incomplete for a Google Scholar domain. It lacks basic CRUD operations such as getting paper metadata, author information, or citation counts, and there are no update or delete capabilities, leaving significant gaps for agent workflows.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Academic research MCP server for paper search, citation checks, graphs, and deep research.
An MCP server that gives your AI access to the source code and docs of all public github repos
Driflyte MCP server which lets AI assistants query topic-specific knowledge from web and GitHub.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables academic research by searching Google Scholar, fetching paper content, and converting web pages to clean Markdown.1MIT
- FlicenseAqualityDmaintenanceAn MCP server that provides access to Semantic Scholar's academic paper database, enabling paper search, detailed retrieval, author info, and citation export.410
- AlicenseNot gradedqualityCmaintenanceAn MCP server for searching Google Scholar, enabling paper search, author lookup, citation tracking, and BibTeX export for AI assistants and automation workflows.2MIT
- AlicenseAqualityDmaintenanceAI-powered academic paper search MCP server with relevance scoring, summarization, author search, and credit checking. Enables users to search papers naturally and get scored results.350MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ProPriyam/Scholar-MCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server