Skip to main content
Glama

LitLib for JLU

面向吉林大学学生的 Agent-ready 文献获取与本地知识库工作流。LitLib 优先从合法 OA 来源获取论文;非开放文献仅在用户具有权限时,通过校园网、CARSI/机构登录或吉大 WebVPN 小批量访问。Zotero 是唯一文献主库。

当前状态:0.3.0 alpha。元数据、OA、PDF 校验、任务状态、RIS 导入、LitLib MCP 已有自动化测试;出版社与 CNKI 路线依赖实时页面,只做有监督 smoke test。

最终框架

本项目固定为 一个 skill + 两个 MCP:

组件

责任

是否写入

litlib-literature-workflow skill

判断检索/下载/读取路线,执行合规边界、人工 checkpoint 和故障决策

不直接写

LitLib MCP

精确读取 Zotero/任务库元数据、PDF 路径、分页全文、任务状态

严格只读

ZotSeek MCP

对已索引 Zotero 文献做 semantic/hybrid/keyword 检索

MCP 只读

litlib CLI

建任务、下载、验证、生成 RIS、确认 Zotero 导入

会写

详细边界见 架构文档。可移植 skill 源码位于 skills/litlib-literature-workflow。分发 skill 时必须 连同 references/ 一起分发。

Related MCP server: zotero-mcp

能做与不能做

LitLib 能处理已有 DOI/PMID/PMCID/arXiv/准确标题,完成元数据解析、OA/机构获取、 严格 PDF 校验、任务追踪和 Zotero 入库。主题级“找论文”仍应先调用 PubMed、Crossref、 OpenAlex、Semantic Scholar 等学术数据库,再把标识符交给 LitLib;项目不把普通网页搜索 伪装成学术检索。

项目不提供 Sci-Hub/LibGen、paywall/CAPTCHA 绕过、代理轮换、无人值守机构批量抓取、 整卷下载,也不直接修改 zotero.sqlite。

安装

要求:Windows 10/11、Python 3.11-3.13、uv、Chrome/Edge、 Zotero 8/9。机构通道面向 JLU,其他学校需要新增 institution adapter。

git clone https://github.com/ganpingzhu904-dev/jlu-literature-agent-mcp.git
Set-Location jlu-literature-agent-mcp
Copy-Item .env.example .env
# 编辑 .env,至少填写真实联系邮箱 LITLIB_EMAIL
uv sync --locked --extra dev
uv run litlib doctor
uv run pytest

.env 不能保存吉大密码、Cookie、VPN token 或 Zotero 密码。吉大统一认证凭证仅通过 litlib inst set-cred 写入 Windows Credential Manager。

没有 D 盘的用户设置 LITLIB_REQUIRE_D_DRIVE=0;有大容量 D 盘的用户可设置为 1, 并把 LITLIB_RUNTIME_ROOT 指向 D 盘。所有设置见 .env.example。

快速工作流

单篇 DOI

uv run litlib download 10.xxxx/example
uv run litlib verify
uv run litlib status --verbose

该入口会自动建任务、解析元数据、下载、校验并登记。目标文件存在时默认拒绝覆盖;只有 用户明确要求时使用 --overwrite。

批量 OA

uv run litlib queue add --file examples/input.example.csv
uv run litlib run --stage fetch-metadata
uv run litlib run --stage oa
uv run litlib status --verbose

OA 候选顺序为 arXiv、可识别的 MDPI public static、Unpaywall、Europe PMC、OpenAlex, 最后才使用具有明确 OA license metadata 的 Crossref link。每个候选独立下载并验证。

JLU 机构通道

只处理 REQUIRES_INST,每批最多 10 篇、并发 1、间隔 8-15 秒:

uv run litlib inst set-cred
uv run litlib inst check-cred
uv run litlib inst open
uv run litlib run --stage inst --access-mode campus
# 校外使用 offcampus;网络未知使用 auto
uv run litlib inst close

遇到 Turnstile、滑块、OTP 或 CAPTCHA 时,程序必须暂停并由用户在可见专用浏览器完成。 出版社明确显示购买/租赁/HTML-only 时停止,不把登录成功等同于拥有 PDF 权限。

CNKI

uv run litlib inst open
uv run litlib cnki open
uv run litlib cnki search "检索词" --limit 10
uv run litlib cnki download "<详情页 URL>" --output "<目标.pdf>"
uv run litlib inst close

bar.cnki.net 滑块必须人工完成。CAJ 不是 PDF,当前不纳入 PDF 验证与全文管线。

Zotero 入库

uv run litlib proposal --doi-file <batch.csv>
# 用户在 Zotero 中导入 RIS 并检查条目/附件
uv run litlib review --doi-file <batch.csv>
uv run litlib import --lookup --batch <name> --doi-file <batch.csv>

只有精确匹配到一个 Zotero parent item、非空 item key 和可读取 PDF 附件时,状态才会变为 IMPORTED。Zotero 未运行或匹配歧义时返回非零且不虚假确认。

明确购买/租赁且无 PDF entitlement 的任务进入 terminal PAYWALLED,不会与 CAPTCHA 的 HUMAN_REQUIRED 或 429 的 RATE_LIMITED 混淆。

Agent 接入

完整配置见 Agent 接入 和 skill 内的 CLIENT_SETUP。OpenCode、Codex 或其他 MCP client 均需注册:

  • LitLib:stdio,<repo>\.venv\Scripts\python.exe -m litlib.cli mcp

  • ZotSeek:HTTP,http://127.0.0.1:23119/zotseek/mcp

ZotSeek 是外部 Zotero 插件,本仓库不捆绑 XPI。安装后先检查 active model 的 index coverage,再解释语义搜索未命中。中文查询英文文献效果弱时,优先改用英文同义查询,而 不是直接认定库中没有相关论文。

站点经验

完整、带日期和证据标签的实测档案位于 SITE_RECIPES.md,包括 Wiley signed pdfdirect、MDPI /pdf?version=、ScienceDirect pdfft、T&F CARSI/HTML-only、 OUP、Springer protocol、CNKI slider、ACS unsupported SP 和 RSC 429。

遇到新站点先按 DECISION_TREE.md 区分 authentication、anti-bot、partial response、HTML-only、paywall 和 wrong-PDF, 不要直接增加 DOI-prefix 硬编码。

个人经验自进化(litlib learn)

仓库内的站点档案是只读基线;每个用户自己的成功经验会自动保存在本地 <LITLIB_RUNTIME_ROOT>\experience\experiences.json(默认 D:\LitLibRuntime\experience\), 不进 Git、不上传、不随项目分发:

  • 机构通道或 CNKI 成功下载(通过严格 PDF 校验)后自动记录站点+路线。

  • 也可手动补充:litlib learn add --domain <域名> --route <路线> --note <说明>。

  • 查看:litlib learn list --domain <站点> 或 litlib learn export(Markdown)。

  • 删除:litlib learn remove <id>。

  • 付费墙/验证码/限流场景永远不会被当作成功经验学习;所有记录自动脱敏。

这样每个人"用一次、记一次",越用越顺,而项目本身不需要持续维护。

纤维素酶数据扩展

LitLib 还提供面向 UniProt accession 的数据准备基础能力:

uv run litlib uniprot P12345 --output output\uniprot_P12345.json
uv run litlib evidence scan staging\downloads\paper.pdf
uv run litlib supplement discover "https://publisher.example/article"
uv run litlib supplement download "https://publisher.example/supp.xlsx" --doi 10.xxxx/example
uv run litlib cellulase validate data\measurements.jsonl
uv run litlib cellulase maxima data\measurements.jsonl --output output\maxima.jsonl

该扩展允许缺失字段,保留原始单位、底物、实验条件和 DOI/页码/表图证据位置。最大值只在 同一构建体、底物、指标族、单位和 assay method 内比较;相对活性、图表估读和待补充材料 不会被静默当作精确绝对活性。图片处理队列和外部多模态模型属于后续可选层,不影响基础 文献获取工作流。

开发与发布

uv run pytest
uv run ruff check .

CI 在 Windows 上测试 Python 3.11 和 3.13。贡献新站点 adapter 前阅读 CONTRIBUTING.md,并用 fixture 测试;真实机构会话不能进入 CI。

运行数据、PDF/CAJ、全文 cache、Zotero/SQLite 数据库、XPI、浏览器 profile、日志、 Cookie、WebVPN token 和 .env 都被排除在 Git 之外。发布前检查见 RELEASE_CHECKLIST.md。

License

LitLib 自有代码与文档使用 MIT License。第三方组件状态见 THIRD_PARTY.md。出版社内容及下载论文不因本项目许可证而获得再分发 权限。

Available Tools

5 tools
library_get_fulltextA

返回匹配文献的分页全文(默认 10000 字符,上限 20000,只读)

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
offsetNo
limit_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses read-only, pagination, default 10000 characters, and maximum 20000 characters, which are useful behavioral traits beyond the schema. However, it does not explain offset semantics, query syntax, error behavior, or how pagination proceeds when results exceed the limit. With no annotations, the description carries the full burden and provides only partial transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, compact sentence with no filler. Every clause adds useful information: the action, the pagination, the default/max limits, and the read-only guarantee. It is front-loaded and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a 3-parameter tool with 0% schema coverage, the description omits essential details about query syntax and offset mechanics, which are necessary for correct invocation. It is adequate for a simple retrieval call but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning only for limit_chars (default and max), but leaves query vaguely defined as 'matching literature' without format or syntax, and offset is merely implied by 'paginated' without clarifying its unit (characters vs. pages) or behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific action (returns paginated full text) and resource (matching literature), and clearly distinguishes itself from siblings like library_search_metadata (metadata) and library_get_pdf_path (PDF paths). The verb '返回' and object '全文' make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or alternatives are mentioned, but the read-only full-text behavior implies it is for retrieving content rather than metadata or file paths. The usage context is inferable from the name and description, yet there is no guidance on exclusions or relationships with siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_get_pdf_pathB

返回匹配文献的 PDF 本地路径(只读)

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Although no annotations are provided, the description explicitly states '只读' (read-only), which conveys behavioral intent (no side effects). This is a key behavioral disclosure for a tool that retrieves a local path, though it doesn't mention potential errors or access restrictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that is front-loaded with the core purpose. It avoids unnecessary details and is appropriately concise for a simple tool, though it could have added a brief parameter hint without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and the tool is simple (1 param), the description covers the main purpose and read-only nature. However, the lack of parameter semantics and usage context leaves gaps for an agent deciding between this and sibling tools, especially 'get_fulltext'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate for the parameter 'query', but it provides no details about how to formulate the query (e.g., title, DOI, full-text keywords). The single parameter is ambiguous, and the description adds no meaning beyond 'query'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('get') and resource ('PDF path') and indicates it returns the local path for matching documents. It is clearly distinct from siblings like 'get_fulltext' or 'search_metadata', though it doesn't explicitly name the alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies using this tool when you need a PDF path, and 'read-only' hints at safety, but it doesn't specify when to use it instead of 'library_get_fulltext' or 'library_search_metadata'. No explicit exclusions or alternatives are mentioned, but the purpose is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_list_collectionsA

列出 Zotero 分类(只读)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does disclose the key read-only nature of the operation, but it adds no further context such as pagination, scope of collections returned, or any side-effect caveats. This is minimal but not misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. Every element — the list operation, the resource, and the read-only qualifier — contributes essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only list tool with an output schema present, this description is nearly complete. It tells the agent what the tool does and that it is safe, while the output schema covers return details. It falls short only by not pointing to sibling tools for alternative operations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema description coverage, so the baseline is 4. The description does not need to add parameter-level meaning because there are no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('列出' / list) and a specific resource ('Zotero 分类' / Zotero collections), and explicitly marks the operation as read-only. This makes it clearly distinct from siblings like library_get_pdf_path or library_get_fulltext.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to choose this tool over sibling tools such as library_search_metadata or library_get_fulltext. Context signals list the alternatives, but the description itself does not state any selection conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_search_metadataA

按标题/DOI 在 Zotero 与本项目库中搜索文献元数据(只读)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral transparency burden. It explicitly states the tool is read-only and covers both Zotero and the project library, which is valuable context. However, it does not disclose potential error behaviors, result limits, or how the two sources are combined, though these are relatively minor for a search operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the search action and criteria. It contains no filler and every word contributes to conveying the tool's purpose and scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential purpose, input semantics, and the dual-source scope, while the output schema exists to define the return structure. It does not mention pagination or edge cases, but for a read-only search operation with a defined output schema, the provided information is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains that 'query' accepts a title or DOI, adding meaningful semantics beyond the generic parameter name. However, it does not explain the 'limit' parameter at all, though its purpose is relatively self-evident from its name and default value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (search), resource (literature metadata), and identifies two concrete sources (Zotero and the project library), while specifying the search criteria (by title/DOI). This clearly differentiates it from siblings like library_get_pdf_path and library_get_fulltext, which handle different retrieval actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when metadata is needed and clarifies it is read-only, but does not explicitly contrast with sibling tools or state when not to use it. Given the distinct sibling functions, the intended context is fairly clear without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_task_statusA

查询下载/导入任务状态(只读),state 如 READY / REQUIRES_INST

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and explicitly marks the operation as read-only ('只读'), which is the key behavioral trait. It also gives example state values, adding useful context about the domain. It does not cover rate limits or auth, but for a simple status query this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence with no filler. The core purpose and key behavioral note are front-loaded, and every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-optional-parameter tool with an output schema, the description is close to adequate. However, it relies on inference for whether 'state' filters results, and it gives only example values rather than full enumeration. A sentence clarifying the state parameter's effect would make it fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning by giving example state values (READY / REQUIRES_INST), which the schema alone does not provide. However, it never explicitly states that the state parameter is an optional filter or what omitting it returns, leaving some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it queries download/import task status. The read-only note and the state examples (READY / REQUIRES_INST) further clarify the domain. It is clearly distinct from sibling tools like metadata search, PDF path, fulltext, and collections.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this when checking download/import task status. It does not explicitly name alternatives or exclusion conditions, but the sibling tool names are sufficiently different in purpose that an agent can route correctly without further guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.3.0
    • First observedlibrary_get_fulltext
    • First observedlibrary_get_pdf_path
    • First observedlibrary_list_collections
    • First observedlibrary_search_metadata
    • First observedlibrary_task_status

TDQS

A3.9/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct concern: metadata search, PDF path retrieval, fulltext access, collection listing, and task status. There is no meaningful overlap among the operations.

Naming Consistency4/5

Most tools follow a clear library_verb_noun pattern, but library_task_status uses a noun phrase instead of a verb phrase like library_get_task_status. The prefix and overall style are consistent enough to be predictable.

Tool Count5/5

Five tools is a well-scoped size for a read-only literature library helper. Each tool serves a clear purpose without redundant surface area.

Completeness4/5

The core read-only workflows are covered: search metadata, retrieve PDF path, access fulltext, and inspect collections. Minor gaps exist, such as no way to fetch collection contents or retrieve detailed item records, but these are not severe for the apparent purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Unified academic search MCP server that searches open literature (arXiv, bioRxiv, medRxiv, PMC), CNKI, and Web of Science, with browser-backed authentication, local paper library, and export to multiple formats.
    21
    2
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    MCP server that grants AI tools read-only access to a Zotero library via search, citekey lookup, and on-demand fulltext retrieval, with low token usage and support for Claude Code, Claude Desktop, and Codex.
    5
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Read-only MCP server for your local Zotero library. Browse collections, inspect paper metadata, and extract full-text from PDFs via FastMCP tools.
    4
    MIT