bitHuman
Server Details
Animated AI characters in your chat: short clips of a character saying your words, or a live talk.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: listing characters, generating a character video, starting a live voice conversation, or reporting widget capabilities. Although make_character_video and talk_to_character both use a character id and produce speech, their modalities and descriptions make the boundary unambiguous.
All names use snake_case and begin with a verb, which is consistent overall. The pattern is mostly verb_noun, with talk_to_character adding a preposition, so it qualifies as a minor deviation rather than a fully uniform convention.
Four tools is well-scoped for a narrow character media service. Each tool earns its place by covering discovery, video generation, live voice interaction, and widget telemetry.
The surface covers character discovery and the two main interaction modes, plus the internal widget reporting need. Minor gaps exist, such as no explicit stop/end control for the live conversation or separate video status retrieval, but agents can still accomplish the core workflows.
Available Tools
4 toolslist_charactersList charactersARead-onlyIdempotentInspect
List the bitHuman house characters that can speak in a video: id, name, a short description and a picture.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| characters | Yes | |
| make_your_own_url | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds only the scope qualifier ('house characters that can speak in a video'), which is modest added context; no auth, rate-limit, or freshness/caching behavior is mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action and the scope before listing the returned fields. No filler, no repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and a full annotation set, the description need not explain return values, and it is sufficient to call the tool. Minor gap: it does not state what 'house' characters means relative to any user-created characters the sibling video tool might accept.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate beyond the schema, and it correctly describes a no-argument call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb plus resource: 'List the bitHuman house characters that can speak in a video', with the returned fields enumerated. It implicitly separates itself from make_character_video and talk_to_character, but never names a sibling or states the contrast explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'that can speak in a video' implies the use case (discover valid character ids before calling make_character_video), but there is no explicit when-to-use statement, no mention of alternatives, and no note on whether this is the required entry point for video creation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_character_videoMake character videoAIdempotentInspect
Make a short video of one bitHuman house character saying the given words, in its own voice. Use a character id from list_characters. The text is said word for word: 3 to 200 characters, about 15 seconds of speech. Returns a video URL that plays inline, labelled AI character.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The exact words the character says. | |
| character | Yes | The character id from list_characters. |
Output Schema
| Name | Required | Description |
|---|---|---|
| talk | No | |
| label | Yes | |
| character | Yes | |
| share_url | No | |
| video_url | Yes | |
| poster_url | No | |
| share_text | No | |
| duration_seconds | Yes | |
| make_your_own_url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnly=false, idempotent=true, destructive=false, so the safety profile is already covered. The description still adds real behavioral value: the output is a video URL that plays inline and is labelled AI character, and speech is rendered verbatim at roughly 15 seconds for 3–200 characters. It does not mention latency or generation cost, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action before the parameter hints. The '3 to 200 characters' restates the schema's minLength/maxLength, a minor redundancy, but nothing else is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, yet it still notes the video URL and inline playback. Combined with the character-id source and length/duration guidance, an agent has enough to call it correctly; only sibling routing guidance is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum is fully enumerated, so baseline would be 3. The description goes modestly beyond the schema by clarifying that 'text' is spoken word for word and that its 3–200 character range corresponds to about 15 seconds of speech, which is genuinely new information an agent can use when composing input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Make a short video of one bitHuman house character saying the given words, in its own voice.' An agent immediately knows this produces a rendered video rather than a live exchange. It does not explicitly contrast with talk_to_character, so the sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage pointer — 'Use a character id from list_characters' — which tells the agent where valid ids come from. However, it offers no when-to-use vs. when-not guidance and never names talk_to_character as the alternative for interactive speech, so selection between siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_widget_capabilitiesReport widget capabilitiesBIdempotentInspect
Used by the bitHuman widget only: records which display modes and browser features the host allows (booleans, no user data).
| Name | Required | Description | Default |
|---|---|---|---|
| pip | No | ||
| webgpu | No | ||
| wss_ok | No | ||
| mic_api | No | ||
| rtc_api | No | ||
| storage | No | ||
| fetch_ok | No | ||
| direct_ok | No | ||
| rtc_srflx | No | ||
| fullscreen | No | ||
| opened_tab | No | ||
| mic_granted | No | ||
| pip_granted | No | ||
| wss_blocked | No | ||
| frame_loaded | No | ||
| render_local | No | ||
| frame_blocked | No | ||
| ended_on_close | No | ||
| fullscreen_granted | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds only that the payload is booleans and carries no user data; it says nothing about where the data goes, whether repeated reports merge or overwrite, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the caller restriction and the verb, with the payload constraint trailing. Nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 19-parameter telemetry tool with no output schema and no enum or description coverage, the description answers who calls it but not what the flags mean or what a call produces. That leaves the agent unable to populate the payload meaningfully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are 19 flags, many highly cryptic (rtc_srflx, direct_ok, render_local, ended_on_close, wss_blocked). The description supplies no per-parameter meaning at all, and 'booleans' merely restates the schema's type declarations, so an agent cannot tell what any individual flag asserts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'records which display modes and browser features the host allows.' The parenthetical '(booleans, no user data)' narrows the payload, and the 'bitHuman widget only' clause separates it from the sibling character/video tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Used by the bitHuman widget only' is an explicit caller restriction, effectively telling the agent when not to use it. It does not name a specific alternative, but the exclusivity clause makes the routing decision unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
talk_to_characterTalk to characterAInspect
Open a live voice conversation with one bitHuman house character in the chat. The character listens through the microphone and answers out loud for up to a minute. Use a character id from list_characters.
| Name | Required | Description | Default |
|---|---|---|---|
| character | Yes | The character id from list_characters. |
Output Schema
| Name | Required | Description |
|---|---|---|
| label | Yes | |
| character | Yes | |
| agent_code | No | |
| max_seconds | Yes | |
| session_url | Yes | |
| make_your_own_url | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a non-read-only, non-idempotent action, and the description usefully adds that the character listens via microphone, speaks aloud, and the session lasts up to a minute. These are behavioral facts absent from the structured fields. It stops short of stating microphone-permission requirements or what happens when the minute expires.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the action, the runtime behavior, and the argument source. No filler and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the interaction model (voice in, voice out, one-minute cap). It is nearly complete for a one-parameter tool, with only permission/blocking nuances left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema coverage is 100% with a full enum plus its own description. The text only points back to list_characters as the source, which the schema description already says verbatim. Baseline 3 applies since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Open a live voice conversation with one bitHuman house character.' This is clearly distinguishable from list_characters (enumeration) and make_character_video (video generation) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent where to source the required argument: 'Use a character id from list_characters.' That is a real routing instruction to a sibling. It does not, however, say when NOT to use this tool or contrast it with make_character_video, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
list_characters - First observed
make_character_video - First observed
report_widget_capabilities - First observed
talk_to_character
Related MCP Connectors
Generate AI talking-head videos with custom characters and voices.
- PrimetaOAuthai.primeta
Give your AI a face, a voice, and a personality. 3D avatars with custom personas.
AI support employee for any website: learns the site, answers visitors by chat and voice.
Create and edit AI videos from chat: plan shots, generate scenes, and export stories and ads.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceLet AI agents create interactive visualizations that render live inside your chat — no code required.1BSD 3-Clause
- AlicenseNot gradedqualityCmaintenanceGenerate and refine AI images/audio/video through natural conversation.408Apache 2.0
- AlicenseNot gradedqualityCmaintenanceTurn words, images, and audio into an animated video with MP4 export.MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI to control 3D VRM models via natural language, supporting expressions, animations, and bone manipulation in real-time through a web browser.-
Glama MCP Gateway
Add one secure layer between your agents and this server.