papervpn-mcp
This MCP server lets an AI agent download paywalled academic papers through the Hainan University WebVPN using your existing authenticated session.
download_paper: Download a single paper PDF from a WebVPN URL, publisher URL, or DOI, optionally specifying a filename or output directory.
download_papers: Download multiple papers sequentially with polite delays, with a configurable max per run (default 10) and output directory.
check_session: Verify whether the current WebVPN session is still authenticated, so you know when re-login is needed.
Enables downloading papers from Elsevier ScienceDirect through the institution's WebVPN, handling pdfDownload/pdfft endpoints and JavaScript challenges via a real browser.
Enables downloading papers from IEEE Xplore through the institution's WebVPN, including IEEE /document/ URLs and stamp/stamp.jsp PDF endpoints.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@papervpn-mcpdownload the paper with DOI 10.1016/j.jare.2020.03.005 to ./papers"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
papervpn
Fetch papers you are institutionally entitled to read — IEEE Xplore,
Elsevier ScienceDirect, and other publishers Hainan University subscribes to —
through the university's WebVPN (webvpn.hainanu.edu.cn).
It works as a CLI, as an MCP server for AI agents, and as a bundled skill for opencode / Codex / Claude Code / the DeepSeek harness.
It downloads one paper at a time, with polite delays, and never bypasses a paywall the institution does not have — it only automates access the user already has.
How it works
The Hainan WebVPN is a 网瑞达 / Wengine gateway. A proxied URL looks like
https://webvpn.hainanu.edu.cn/https/<TOKEN>/<original-path>where TOKEN = IV(16 bytes) + AES-CFB(hostname) and is deterministic per host.
papervpn:
reuses your browser's authenticated
wengine_vpn_ticketsession;builds/reuses the per-host token (and can regenerate tokens for new hosts after reading
wrdvpnKey/wrdvpnIVfrom the portal's/user/info);rewrites the publisher page into its real PDF endpoint — IEEE
stamp/stamp.jsp→stampPDF/getPDF.jsp, ScienceDirect"pdfDownload"→pdfft(clearing the JS challenge with a real Chrome window);saves the PDF, de-duplicates, and throttles.
Related MCP server: PDF Indexer MCP Server
Requirements
Python ≥ 3.10,
requests,cryptography.Optional:
playwright(ScienceDirect),secretstorage(renew the session from the browser keyring).An existing, logged-in WebVPN session in a local browser.
Install
pip install --user --break-system-packages -e .
# or just run the launchers in skill/papervpn/scripts/Install the agent skill. By default it goes to ~/.agents/skills, the shared
directory that opencode, Codex, Claude Code and dsh all read — one copy covers
them all:
./install_skill.sh # -> ~/.agents/skills (default)
./install_skill.sh -a codex,dsh # or pick specific agents (multi-select)
./install_skill.sh --all # install to every known agent
./install_skill.sh --list # show install status
./install_skill.sh -r --all # remove from everywhereQuick start
papervpn check # authenticated: True ?
papervpn login --from-chrome # import the live ticket from your browser
papervpn download "10.1016/j.jare.2020.03.005" # saves into the current directory
papervpn download "https://ieeexplore.ieee.org/document/11376648" -o some/dirAccepted inputs: WebVPN URLs, publisher URLs, IEEE /document/<arnumber>
URLs, DOIs, and ScienceDirect PIIs.
CLI
Command | Purpose |
| Validate the session; prints the token/key status |
| Authenticate: |
| Download PDFs ( |
| Read |
| Inspect / extend the per-host token cache |
Download politeness: serial, 3–8 s random delay, de-duplicated history,
--max 10 per run, hard cap 50 items.
When the session expires
wengine_vpn_ticket is short-lived. When check reports
authenticated: False, renew it — no password needed if the browser is still
logged in:
papervpn login --from-chromeThis reads the live ticket from Chrome/Chromium (keyring via secretstorage)
or Firefox (plaintext) and saves it to ~/.papervpn/cookies.json. Fallbacks:
papervpn login --cookie "<browser 'Copy as cURL'>"
papervpn login --cookie "wengine_vpn_ticket=...; route=..."
papervpn login # open a browser and log in by handMCP server
papervpn-mcp speaks stdio MCP and exposes download_paper,
download_papers, and check_session.
// opencode ~/.config/opencode/opencode.jsonc
"mcp": { "papervpn": { "type": "local", "command": ["papervpn-mcp"], "enabled": true } }# Codex ~/.codex/config.toml
[mcp_servers.papervpn]
command = "papervpn-mcp"
tool_timeout_sec = 900Give the client a generous tool timeout: a ScienceDirect fetch drives a real
browser and can take ~60–120 s. Downloads default to the current working
directory; if you pass output_dir, keep it inside the project (not a
sandbox-private /tmp).
Configuration
State lives in ~/.papervpn/ (cookies, token cache, download history) and is
independent of the working directory, so the CLI and MCP server share one
session.
Env var | Meaning |
| State directory (default |
| Default download directory (default: current working directory) |
| WebVPN host (default |
| Override the host-token codec key/IV |
| Source tree used by the skill launchers |
Troubleshooting
Symptom | Fix |
|
|
| Throttled — wait a few minutes and retry one paper with |
| Open the host once through the WebVPN, then |
IEEE | Re-login and retry a single paper |
ScienceDirect returns HTML, not PDF | Chrome must be available; run with a display ( |
Legal & acceptable use
Use only your own institutional entitlement, for personal reading. Download one paper at a time; do not mass-download, redistribute, or attempt to access content the university does not subscribe to. This tool automates a logged-in session that the user controls; it never asks for or stores an account password, one-time code, or CAPTCHA answer.
License
Available Tools
3 toolscheck_sessionA
Report whether the current WebVPN session is authenticated.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. 'Report whether' reasonably implies a read-only, side-effect-free query, which is useful context. However, it does not disclose the return shape (boolean vs. status text) or whether the check can refresh/alter the session, leaving real behavioral questions open.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the subject and condition are stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a trivial no-arg check this is nearly sufficient, but with no output schema the description should state what 'report' actually returns (e.g., a boolean or status string) so the agent knows how to interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and the schema coverage is 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Report') and resource ('current WebVPN session is authenticated'), making the purpose immediately clear. It is naturally distinct from the sibling download_* tools, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. The usage is only implied – an agent would infer it should be called before attempting downloads that need an authenticated session, but the description never says so.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_paperA
Download a single paywalled paper PDF through the Hainan University WebVPN. Accepts a WebVPN-proxied URL (https://webvpn.hainanu.edu.cn/https/...), a publisher URL, or a DOI. Downloads slowly and one at a time.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | WebVPN URL, publisher URL, or DOI | |
| filename | No | optional output filename | |
| output_dir | No | destination directory |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses operational traits beyond the schema: downloads are slow and serialized ('one at a time'), and the WebVPN requirement. It omits failure/error behavior, whether files are overwritten, and any dependency on an authenticated session, even though the sibling check_session implies auth is a factor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the core action and input forms; the example URL is high-value. Slight redundancy between 'slowly' and 'one at a time' keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a download tool with no output schema, the description should indicate what is returned (e.g., saved file path or confirmation) and whether a valid WebVPN session is a precondition. It covers inputs and performance but leaves those two agent-relevant points unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url ('WebVPN URL, publisher URL, or DOI'), filename, and output_dir. The description's phrasing of the accepted url forms duplicates the schema rather than extending it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Download), resource (a single paywalled paper PDF), and the mechanism (Hainan University WebVPN). It also implicitly separates itself from the sibling download_papers via 'single' and 'one at a time'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when it applies (single paywalled paper, via WebVPN, accepting WebVPN URL / publisher URL / DOI), which tells the agent what inputs qualify. It does not explicitly name download_papers as the alternative for batches, so the routing guidance stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_papersA
Download several papers sequentially with polite delays. Refuses large batches.
| Name | Required | Description | Default |
|---|---|---|---|
| max | No | max downloads this run (default 10) | |
| urls | Yes | ||
| output_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real traits: requests are sequential rather than parallel, delays are deliberately polite, and large batches are rejected. However it omits what counts as 'large', whether auth/session is required, and any rate-limit or failure semantics, so the disclosure is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and the key constraint (sequentially, polite delays) and no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter batch tool with no annotations and no output schema, the description leaves meaningful gaps: the threshold for refusal, auth/session expectations (relevant given the check_session sibling), and behavior on partial failure. One required parameter is named but its handling is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: 'max' is described in the schema, while 'urls' and 'output_dir' have no descriptions. The text does not fill those gaps, adding no format, ordering, or destination detail. Baseline is above 3 only if the description compensates, and it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (download) and resource (several papers), and 'sequentially' plus 'refuses large batches' marks it apart from the singular sibling download_paper. It is not fully explicit that this is the batch variant versus the single-paper sibling, but the plural and sequential framing get an agent most of the way there.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies when to use it (multi-paper, moderate batch) and signals a refusal case, but never names the alternative download_paper or states a concrete threshold or condition for choosing either. Usage is inferable rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
check_session - First observed
download_paper - First observed
download_papers
TDQS
Scored across 3 tools
download_paper and download_papers are distinct by singular vs plural, and check_session is clearly different. Minor risk of confusion between single and batch download, but descriptions clarify the difference.
All tools use snake_case with a verb_noun pattern (download_paper, download_papers, check_session). The convention is consistent and predictable.
Three tools cover the core functions: single download, batch download, and session check. The count is well-scoped for the narrow purpose of downloading paywalled papers via WebVPN.
The set lacks a tool to initiate or refresh authentication, leaving a dead end if check_session reports unauthenticated. Core download operations are present, but the auth lifecycle is incomplete.
Maintenance
Related MCP Connectors
MCP tools for AI agents: render URLs to image/PDF, check link health, convert HTML/CSV/JSON.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Read-only MCP tools for AI agent discovery, structured resources, and NIULAI information.
A paid remote MCP for AI agent browser DevTools MCP, built to return verdicts, receipts, usage logs,
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables LLM interaction with the Macau University of Science and Technology (M.U.S.T.) campus system, including automated login to Wemust and Moodle, retrieving class schedules, checking assignments and deadlines, downloading course materials, and managing course content.73MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to download, index, and semantically search PDF research papers using 8 MCP tools.2GPL 3.0
- FlicenseNot gradedqualityCmaintenanceLocal MCP server that enables downloading papers from arXiv, OpenReview, INFORMS, and Wiley by reusing your Chrome login state.1-
- FlicenseAqualityCmaintenanceEnables AI agents to search, read, and cite Chinese academic papers from CNKI using HUST single sign-on, with support for full-text retrieval, BibTeX export, and PDF download.72-