Skip to main content
Glama
spx86
by spx86

papervpn

Fetch papers you are institutionally entitled to read — IEEE Xplore, Elsevier ScienceDirect, and other publishers Hainan University subscribes to — through the university's WebVPN (webvpn.hainanu.edu.cn).

It works as a CLI, as an MCP server for AI agents, and as a bundled skill for opencode / Codex / Claude Code / the DeepSeek harness.

It downloads one paper at a time, with polite delays, and never bypasses a paywall the institution does not have — it only automates access the user already has.

How it works

The Hainan WebVPN is a 网瑞达 / Wengine gateway. A proxied URL looks like

https://webvpn.hainanu.edu.cn/https/<TOKEN>/<original-path>

where TOKEN = IV(16 bytes) + AES-CFB(hostname) and is deterministic per host. papervpn:

  1. reuses your browser's authenticated wengine_vpn_ticket session;

  2. builds/reuses the per-host token (and can regenerate tokens for new hosts after reading wrdvpnKey / wrdvpnIV from the portal's /user/info);

  3. rewrites the publisher page into its real PDF endpoint — IEEE stamp/stamp.jsp → stampPDF/getPDF.jsp, ScienceDirect "pdfDownload" → pdfft (clearing the JS challenge with a real Chrome window);

  4. saves the PDF, de-duplicates, and throttles.

Related MCP server: PDF Indexer MCP Server

Requirements

  • Python ≥ 3.10, requests, cryptography.

  • Optional: playwright (ScienceDirect), secretstorage (renew the session from the browser keyring).

  • An existing, logged-in WebVPN session in a local browser.

Install

pip install --user --break-system-packages -e .
# or just run the launchers in skill/papervpn/scripts/

Install the agent skill. By default it goes to ~/.agents/skills, the shared directory that opencode, Codex, Claude Code and dsh all read — one copy covers them all:

./install_skill.sh                 # -> ~/.agents/skills (default)
./install_skill.sh -a codex,dsh    # or pick specific agents (multi-select)
./install_skill.sh --all           # install to every known agent
./install_skill.sh --list          # show install status
./install_skill.sh -r --all        # remove from everywhere

Quick start

papervpn check                                     # authenticated: True ?
papervpn login --from-chrome                       # import the live ticket from your browser
papervpn download "10.1016/j.jare.2020.03.005"     # saves into the current directory
papervpn download "https://ieeexplore.ieee.org/document/11376648" -o some/dir

Accepted inputs: WebVPN URLs, publisher URLs, IEEE /document/<arnumber> URLs, DOIs, and ScienceDirect PIIs.

CLI

Command

Purpose

papervpn check

Validate the session; prints the token/key status

papervpn login

Authenticate: --from-chrome, --cookie "<header|curl|file>", or open a browser

papervpn download <urls…>

Download PDFs (-b/--batch, -o/--output, --max, --no-browser, --browser-headed)

papervpn key-discover

Read wrdvpnKey/wrdvpnIV from the portal to encode new hosts

papervpn token list|add

Inspect / extend the per-host token cache

Download politeness: serial, 3–8 s random delay, de-duplicated history, --max 10 per run, hard cap 50 items.

When the session expires

wengine_vpn_ticket is short-lived. When check reports authenticated: False, renew it — no password needed if the browser is still logged in:

papervpn login --from-chrome

This reads the live ticket from Chrome/Chromium (keyring via secretstorage) or Firefox (plaintext) and saves it to ~/.papervpn/cookies.json. Fallbacks:

papervpn login --cookie "<browser 'Copy as cURL'>"
papervpn login --cookie "wengine_vpn_ticket=...; route=..."
papervpn login            # open a browser and log in by hand

MCP server

papervpn-mcp speaks stdio MCP and exposes download_paper, download_papers, and check_session.

// opencode ~/.config/opencode/opencode.jsonc
"mcp": { "papervpn": { "type": "local", "command": ["papervpn-mcp"], "enabled": true } }
# Codex ~/.codex/config.toml
[mcp_servers.papervpn]
command = "papervpn-mcp"
tool_timeout_sec = 900

Give the client a generous tool timeout: a ScienceDirect fetch drives a real browser and can take ~60–120 s. Downloads default to the current working directory; if you pass output_dir, keep it inside the project (not a sandbox-private /tmp).

Configuration

State lives in ~/.papervpn/ (cookies, token cache, download history) and is independent of the working directory, so the CLI and MCP server share one session.

Env var

Meaning

PAPERVPN_STATE

State directory (default ~/.papervpn)

PAPERVPN_OUTPUT

Default download directory (default: current working directory)

PAPERVPN_HOST

WebVPN host (default webvpn.hainanu.edu.cn)

PAPERVPN_AES_KEY / PAPERVPN_AES_IV

Override the host-token codec key/IV

PAPERVPN_ROOT

Source tree used by the skill launchers

Troubleshooting

Symptom

Fix

authenticated: False

papervpn login --from-chrome

ScienceDirect anti-bot block (CPE00001)

Throttled — wait a few minutes and retry one paper with --browser-headed

must be accessed through the WebVPN but no token is cached

Open the host once through the WebVPN, then papervpn token add <url> or papervpn key-discover

IEEE HTTP 418 / 403

Re-login and retry a single paper

ScienceDirect returns HTML, not PDF

Chrome must be available; run with a display (--browser-headed)

Use only your own institutional entitlement, for personal reading. Download one paper at a time; do not mass-download, redistribute, or attempt to access content the university does not subscribe to. This tool automates a logged-in session that the user controls; it never asks for or stores an account password, one-time code, or CAPTCHA answer.

License

MIT

Available Tools

3 tools
check_sessionA

Report whether the current WebVPN session is authenticated.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'Report whether' reasonably implies a read-only, side-effect-free query, which is useful context. However, it does not disclose the return shape (boolean vs. status text) or whether the check can refresh/alter the session, leaving real behavioral questions open.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the subject and condition are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial no-arg check this is nearly sufficient, but with no output schema the description should state what 'report' actually returns (e.g., a boolean or status string) so the agent knows how to interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and the schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource ('current WebVPN session is authenticated'), making the purpose immediately clear. It is naturally distinct from the sibling download_* tools, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. The usage is only implied – an agent would infer it should be called before attempting downloads that need an authenticated session, but the description never says so.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_paperA

Download a single paywalled paper PDF through the Hainan University WebVPN. Accepts a WebVPN-proxied URL (https://webvpn.hainanu.edu.cn/https/...), a publisher URL, or a DOI. Downloads slowly and one at a time.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesWebVPN URL, publisher URL, or DOI
filenameNooptional output filename
output_dirNodestination directory

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses operational traits beyond the schema: downloads are slow and serialized ('one at a time'), and the WebVPN requirement. It omits failure/error behavior, whether files are overwritten, and any dependency on an authenticated session, even though the sibling check_session implies auth is a factor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences, front-loaded with the core action and input forms; the example URL is high-value. Slight redundancy between 'slowly' and 'one at a time' keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a download tool with no output schema, the description should indicate what is returned (e.g., saved file path or confirmation) and whether a valid WebVPN session is a precondition. It covers inputs and performance but leaves those two agent-relevant points unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents url ('WebVPN URL, publisher URL, or DOI'), filename, and output_dir. The description's phrasing of the accepted url forms duplicates the schema rather than extending it, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Download), resource (a single paywalled paper PDF), and the mechanism (Hainan University WebVPN). It also implicitly separates itself from the sibling download_papers via 'single' and 'one at a time'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when it applies (single paywalled paper, via WebVPN, accepting WebVPN URL / publisher URL / DOI), which tells the agent what inputs qualify. It does not explicitly name download_papers as the alternative for batches, so the routing guidance stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_papersA

Download several papers sequentially with polite delays. Refuses large batches.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxNomax downloads this run (default 10)
urlsYes
output_dirNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add real traits: requests are sequential rather than parallel, delays are deliberately polite, and large batches are rejected. However it omits what counts as 'large', whether auth/session is required, and any rate-limit or failure semantics, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and the key constraint (sequentially, polite delays) and no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter batch tool with no annotations and no output schema, the description leaves meaningful gaps: the threshold for refusal, auth/session expectations (relevant given the check_session sibling), and behavior on partial failure. One required parameter is named but its handling is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: 'max' is described in the schema, while 'urls' and 'output_dir' have no descriptions. The text does not fill those gaps, adding no format, ordering, or destination detail. Baseline is above 3 only if the description compensates, and it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (download) and resource (several papers), and 'sequentially' plus 'refuses large batches' marks it apart from the singular sibling download_paper. It is not fully explicit that this is the batch variant versus the single-paper sibling, but the plural and sequential framing get an agent most of the way there.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies when to use it (multi-paper, moderate batch) and signals a refusal case, but never names the alternative download_paper or states a concrete threshold or condition for choosing either. Usage is inferable rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedcheck_session
    • First observeddownload_paper
    • First observeddownload_papers

TDQS

A3.8/5.0

Scored across 3 tools

Disambiguation4/5

download_paper and download_papers are distinct by singular vs plural, and check_session is clearly different. Minor risk of confusion between single and batch download, but descriptions clarify the difference.

Naming Consistency5/5

All tools use snake_case with a verb_noun pattern (download_paper, download_papers, check_session). The convention is consistent and predictable.

Tool Count5/5

Three tools cover the core functions: single download, batch download, and session check. The count is well-scoped for the narrow purpose of downloading paywalled papers via WebVPN.

Completeness3/5

The set lacks a tool to initiate or refresh authentication, leaving a dead end if check_session reports unauthenticated. Core download operations are present, but the auth lifecycle is incomplete.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers