chinese-char-counter-mcp
You can count and extract Chinese characters from text via three read-only MCP tools.
Count Chinese characters in a single text, excluding punctuation, spaces, letters, digits, emoji, etc.
Get total characters, Chinese ratio, and a breakdown of Chinese, letters, digits, punctuation, spaces, and other.
Batch-count up to 500 texts with per-item results and a total Chinese count.
Extract only the Chinese characters from mixed text.
Use all tools safely since they are read-only and idempotent.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@chinese-char-counter-mcp数一下这段文案有几个中文字:新品上市,全场五折!"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
chinese-char-counter-mcp
一个用 Python 实现的 MCP 服务(Model Context Protocol Server,STDIO 传输), 用于统计一段文本里的中文字数——标点、空格、英文字母、数字、emoji 等一律不计入。
传输方式:STDIO(本地进程,无网络、无 API Key、无环境变量)
运行时依赖:Python >= 3.10、
mcp>=1.0.0,<2已通过 ModelScope MCP 广场托管部署所需的配置形态(
command为uvx,包发布到 PyPI)
这个服务解决什么问题
大模型写中文文案时,"字数"常常对不上:把标点、空格、英文单词都算进去,或者把 emoji 也算成字。本服务给出一个确定的答案:只数中文字符,并把标点/字母/数字/空白 等分项一并返回,方便直接用于文案校验、作业字数检查、标题长度限制等场景。
Related MCP server: Twitter MCP Server
客户端配置
把下面这段配置加入任意支持 MCP 的客户端(Claude Desktop、Cursor、Cherry Studio、通义灵码、 ModelScope MCP 实验场等):
{
"mcpServers": {
"chinese-char-counter": {
"command": "npx",
"args": ["-y", "chinese-char-counter-mcp@latest"]
}
}
}说明:
该配置用 npm 上的零依赖 Node 实现(
npx会自动下载并运行,无需手动安装)。本服务不需要任何环境变量,因此配置里没有
env字段。想用 Python 实现:
pip install chinese-char-counter-mcp安装后,把配置里的command改成uvx、args改成该包名(即uvx chinese-char-counter-mcp@latest)即可;两种实现的工具名、参数、返回结构完全一致。command用npx或uvx是 ModelScope(魔搭)MCP 广场托管部署检测的硬要求,其他写法不会被托管。
安装
# 方式一:pip 安装后直接启动(控制台命令)
pip install chinese-char-counter-mcp
chinese-char-counter-mcp
# 方式二:模块方式启动
python -m chinese_char_counter_mcp
# 方式三:不安装,临时运行(需要 uv)
uvx chinese-char-counter-mcp@latestSTDIO 服务启动后不打印任何内容、等待客户端的 JSON-RPC 请求,这是正常现象。
工具
1. count_chinese_characters
统计单条文本的中文字数。
参数 | 类型 | 必填 | 说明 |
| string | 是 | 待统计文本,长度上限 200000 个字符 |
返回字段:
字段 | 类型 | 说明 |
| integer | 中文字数(唯一需要关心的结果) |
| integer | 文本总字符数(含标点、空格等全部字符) |
| number | 中文占全部字符的比例,保留 4 位小数 |
| object | 分项计数: |
| string | 失败原因;成功时为空字符串 |
调用 count_chinese_characters(text="你好,World 2026!") 的返回:
{
"chinese_count": 2,
"total_characters": 14,
"chinese_ratio": 0.1429,
"breakdown": {
"chinese": 2,
"letters": 5,
"digits": 4,
"punctuation": 2,
"spaces": 1,
"other": 0
},
"error": ""
}2. count_chinese_characters_batch
批量统计多条文本,并给出合计,适合一次校验多段文案。
参数 | 类型 | 必填 | 说明 |
| array | 是 | 文本列表,最多 500 条,每条上限 200000 个字符 |
返回字段:results(与输入顺序一致,每项含 index / total_characters /
chinese_count / chinese_ratio)、text_count、total_chinese_count、error。
3. extract_chinese_text
抽取文本中的中文字符,剔除标点、空格、英文、数字等其它内容。
参数 | 类型 | 必填 | 说明 |
| string | 是 | 待处理文本,长度上限 200000 个字符 |
调用 extract_chinese_text(text="Hello 世界 2026") 的返回:
{
"chinese_text": "世界",
"chinese_count": 2,
"error": ""
}三个工具都是只读、幂等操作(MCP readOnlyHint / idempotentHint 已标记为 true),
客户端可以安全地自动调用。
计数规则
计入中文:
CJK 统一表意文字基本区(U+4E00–U+9FFF)与扩展 A–I 区
兼容表意文字(U+F900–U+FAFF、U+2F800–U+2FA1F)
表意数字零"〇"(U+3007)
不计入:
中文标点与英文标点:
,。、;:!?""''()《》、,.;:!?()<>等空白:空格、制表符、换行
英文字母与其它拉丁字母、阿拉伯数字
emoji、各类符号
其它文字系统:日文假名、韩文谚文、西里尔字母等(归入
other/letters)
已知边界(刻意为之,不视为缺陷):
日文汉字与中文汉字同属 CJK 表意文字区,Unicode 层面无法区分,一律计入中文。
日文迭字符"々"(U+3005)不是表意文字本体,不计入。
全角数字"123"属于数字,不计入;汉字数字"一二三"计入。
项目结构
chinese-char-counter-mcp/
├── src/chinese_char_counter_mcp/
│ ├── counter.py # 计数核心:纯函数,无副作用,可单独复用
│ ├── server.py # MCP 工具定义与显式注册、控制台入口
│ └── __main__.py # python -m 入口
├── npm/ # 零依赖 Node 实现(npm 包,供 npx 使用)
│ ├── bin/server.js # MCP 服务入口(STDIO)
│ ├── lib/counter.js # 计数核心(与 Python 版逐字段对齐)
│ ├── lib/protocol.js# 极简 MCP JSON-RPC 服务端
│ └── test/ # node --test:计数规则 + 协议链路
├── tests/ # pytest 单元测试(计数规则 + 工具信封契约)
├── scripts/
│ └── e2e_stdio_client.py # STDIO 端到端验证:initialize -> list_tools -> call_tool
├── docs/publish-guide.md # 发布到 GitHub / PyPI / ModelScope MCP 广场的步骤
└── pyproject.toml开发
# Python 实现
pip install -e ".[dev]"
pytest -q # 单元测试
python scripts/e2e_stdio_client.py # STDIO 端到端验证(需已安装 mcp)
# Node 实现
cd npm && npm test # node --test:计数规则 + STDIO 协议全链路设计约定:
counter.py只做纯计算,不 importmcp,方便复用与测试。server.py的工具永远返回 dict 信封、永不向 MCP 层抛裸异常:成功时error为空字符串, 失败时error给出可直接阅读的原因,其余字段类型保持稳定。服务端不向 stdout 打印任何内容(STDIO 传输里 stdout 是协议通道),日志走 stderr。
在 ModelScope(魔搭)MCP 广场上架
本仓库已按魔搭"从 GitHub 仓库快速创建"的要求准备:根目录 README 正文中包含可解析的
STDIO 服务配置(即上面的 mcpServers JSON 块),command 为 uvx,包已发布到 PyPI。
完整步骤与部署检测自查清单见 docs/publish-guide.md。
License
Available Tools
3 toolscount_chinese_charactersARead-onlyIdempotent
统计一段文本的中文字数(不算标点、空格、英文字母、数字等)。
Args: text: 待统计的文本,长度上限 200000 个字符。
Returns: chinese_count 为中文字数(CJK 表意文字,含扩展区与“〇”); total_characters 为文本总字符数(含标点、空格等全部字符); chinese_ratio 为中文占全部字符的比例(保留 4 位小数); breakdown 为分项计数:chinese/letters/digits/punctuation/spaces/other; error 为失败原因,成功时为空字符串。
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| breakdown | Yes | |
| chinese_count | Yes | |
| chinese_ratio | Yes | |
| total_characters | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: a 200,000-character input limit and the fact that failures surface via an 'error' field rather than throwing. That is useful but not rich; no mention of performance implications at the upper length bound.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then separates Args and Returns cleanly. The Returns block is somewhat long given an output schema already exists, which is mild redundancy, but every element remains readable and no sentence is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter counting tool with annotations and an output schema, the definition covers input limits, scope of the count, and error signaling. The main gap is sibling disambiguation, which the description leaves entirely to the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter alone, and it does: it defines 'text' as the text to be counted and gives a concrete 200,000-character upper bound. That is the key semantic an agent needs before calling, though nothing is said about encoding or multibyte handling.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource (counting Chinese characters in a text) and precisely scopes what counts (CJK ideographs including extensions and '〇', excluding punctuation, spaces, letters, digits). It does not, however, distinguish itself from its siblings count_chinese_characters_batch or extract_chinese_text, so an agent must infer the division of labor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the narrowly defined scope of what gets counted, letting an agent infer this is the single-text counter. But there is no explicit when-to-use statement, no mention that a batch sibling exists for multiple texts, and no guidance on selection versus extract_chinese_text.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
count_chinese_characters_batchARead-onlyIdempotent
批量统计多条文本的中文字数,并给出合计。
Args: texts: 文本列表,最多 500 条,每条长度上限 200000 个字符。
Returns: results 与输入顺序一致,每项含 index/total_characters/chinese_count/chinese_ratio; text_count 为实际统计条数;total_chinese_count 为所有文本的中文字数合计; error 为失败原因,成功时为空字符串。
| Name | Required | Description | Default |
|---|---|---|---|
| texts | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| results | Yes | |
| text_count | Yes | |
| total_chinese_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, closed world), so the bar is lower. Beyond them the description adds concrete operational limits (max 500 texts, 200000 chars each) and failure semantics via the 'error' field, which are genuinely useful and not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose followed by tightly organized Args and Returns sections. Every sentence carries information (limits, ordering, aggregate field, error field) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though an output schema exists, the description usefully enumerates the result fields and the error convention, and supplies the input limits. The only minor gap is no explicit statement of when to prefer this over the single-text sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (the schema only says 'array of strings'), so the description must compensate — and it does, documenting that texts is a list capped at 500 items with a 200000-character per-item ceiling. That constraint meaningfully exceeds the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a specific verb and resource ('批量统计多条文本的中文字数') plus the aggregation behavior ('给出合计'), so the operation is unambiguous. It does not explicitly name the singular sibling count_chinese_characters, but '批量' (batch) clearly separates it from that sibling by inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via the word '批量' — an agent can infer this is for multiple texts versus the single-text sibling, but there is no explicit when-to-use, when-not-to-use, or alternative routing. No mention of extract_chinese_text as a related option.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_chinese_textARead-onlyIdempotent
抽取文本中的中文字符,剔除标点、空格、英文、数字等其它字符。
Args: text: 待处理的文本,长度上限 200000 个字符。
Returns: chinese_text 为仅由中文字符按原顺序组成的文本; chinese_count 为抽取出的中文字数; error 为失败原因,成功时为空字符串。
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| chinese_text | Yes | |
| chinese_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a safe, read-only, idempotent, closed-world operation, and the description does not contradict them. The description adds genuinely useful context beyond the annotations: a hard input length cap of 200000 characters and the semantics of the error field (empty string on success). This is meaningful behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first line, followed by cleanly separated Args and Returns sections. The Returns section slightly duplicates the existing output schema, but the structure is tight and each line earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, non-destructive utility with an output schema, the definition covers everything needed: purpose, the input constraint, and the meaning of the returned fields including the error indicator. Since an output schema exists, the Returns prose is redundant rather than necessary, which keeps this just short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the input schema alone carries no meaning for the sole parameter. The description compensates by defining 'text' as the text to be processed and stating its maximum length (200000 characters), which is exactly the information an agent needs before invoking.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (抽取/extract) and resource (中文字符/Chinese characters), and it enumerates precisely what is excluded: punctuation, spaces, English, numbers and other characters. An agent understands the transformation unambiguously. It does not, however, explicitly distinguish itself from the sibling tools count_chinese_characters / count_chinese_characters_batch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance: the description never states the conditions under which this extraction is preferred over the sibling counting tools, nor any prerequisites or exclusions. The purpose is clear, but routing guidance is entirely absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
count_chinese_characters - First observed
count_chinese_characters_batch - First observed
extract_chinese_text
TDQS
Scored across 3 tools
The three tools have related but fairly distinguishable purposes: single counting, batch counting, and Chinese-only extraction. There is some overlap between count_chinese_characters and count_chinese_characters_batch, but the batch version is clearly scoped for multiple texts.
All tool names follow a consistent verb_noun-ish snake_case pattern: count_chinese_characters, count_chinese_characters_batch, extract_chinese_text. The relationship between the count tools is also clearly signposted by the _batch suffix.
Three tools is a reasonable, focused set for a Chinese character counter server. It is not thin for the core purpose, though batch counting could arguably be folded into the main count tool.
The server covers counting, batch counting, and extracting Chinese characters, which is solid for the domain. There are minor gaps around broader text-processing needs such as normalization or validation, but the stated purpose is well covered.
Maintenance
Related MCP Connectors
Count occurrences of any character in your text instantly. Specify the character and get precise c…
Exact character/word counting, reversal, palindrome checks, indexing, sorting; Unicode-safe.
Character count, text discarded
Chinese web novel MCP: 36 tools (outline, prose, review, coach, KD export). BYOK, no API key.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables counting characters or bytes in text with options to include or exclude whitespace. Provides a simple tool for text analysis and length measurement.1MIT
- FlicenseAqualityDmaintenanceProvides tools for accurate Twitter/X post character counting, validation, and optimization using official counting methods. It enables users to extract entities like URLs and hashtags to ensure content fits within platform constraints.41-
- AlicenseNot gradedqualityDmaintenanceProvides tools for AI models to count characters and words in text, supporting English and other space-delimited languages.1MIT
- AlicenseNot gradedqualityDmaintenanceMCP server that counts the number of Chinese characters (excluding punctuation, spaces, English, and numbers) in a given text, compatible with Streamable HTTP and JSON-RPC.MIT