Skip to main content
Glama
IzzyFuller

intentional-masking

by IzzyFuller

intentional-masking

MCP server for rendering Ready Player Me avatars with lip-sync using Remotion and React Three Fiber.

Features

  • render_frame - Render a single still image of an avatar with expression, pose, camera, and lighting options

  • render_speaking_video - Render a video of an avatar speaking with real phoneme-based lip-sync from audio

Related MCP server: Remotion Video Generator

Requirements

  • Node.js 18+

  • macOS, Linux, or Windows

Installation

cd intentional-masking
npm install
npm run build

Usage

As MCP Server

Add to your Claude Code configuration (~/.claude.json or project .claude/settings.local.json):

{
  "mcpServers": {
    "intentional-masking": {
      "command": "node",
      "args": ["/path/to/intentional-masking/dist/server/index.js"],
      "env": {
        "INTENTIONAL_MASKING_ROOT": "/path/to/intentional-masking"
      }
    }
  }
}

MCP Tools

render_frame

Render a single still frame of an avatar.

{
  "avatar_path": "/path/to/avatar.glb",
  "expression": "happy",
  "pose": "greeting",
  "camera_preset": "closeup",
  "lighting_preset": "soft",
  "background": "#1a1a2e"
}

Options:

  • expression: neutral | happy | thinking | surprised

  • pose: default | greeting | listening

  • camera_preset: closeup | medium | full

  • lighting_preset: soft | dramatic | natural

  • background: Hex color (default: #1a1a2e)

Returns: { "success": true, "image_path": "/path/to/output.png" }

render_speaking_video

Render an avatar speaking with lip-sync from audio.

{
  "avatar_path": "/path/to/avatar.glb",
  "audio_path": "/path/to/speech.wav",
  "camera_preset": "closeup",
  "lighting_preset": "soft",
  "background": "#1a1a2e"
}

Audio requirements:

  • 16kHz 16-bit mono PCM WAV (as produced by info-dump)

Returns: { "success": true, "video_path": "/path/to/output.mp4", "duration_seconds": 5.2 }

Architecture

src/
├── server/
│   ├── index.ts              # MCP server entry point
│   ├── services/
│   │   └── lip-sync.ts       # Rhubarb lip-sync integration
│   └── tools/
│       ├── render-frame.ts   # Still image rendering
│       └── render-speaking-video.ts  # Video rendering
├── config/
│   └── viseme-map.ts         # Rhubarb → RPM morph target mapping
└── remotion/
    ├── index.ts              # Remotion entry point
    ├── Root.tsx              # Composition registration
    ├── AvatarFrame.tsx       # Still frame composition
    ├── AvatarSpeaking.tsx    # Speaking video composition
    └── components/
        ├── Avatar.tsx        # GLB model loader
        ├── Scene.tsx         # Three.js scene setup
        └── LipSyncController.tsx  # Morph target application

Lip-Sync Pipeline

  1. Audio analysis: rhubarb-lip-sync-wasm analyzes 16kHz audio

  2. Phoneme mapping: Rhubarb shapes (A-H, X) → Ready Player Me viseme morph targets

  3. Frame interpolation: Smooth blending between viseme shapes

  4. Video rendering: Remotion captures Three.js scene frame-by-frame

Rhubarb Shape Mapping

Shape

Phonemes

RPM Morph Targets

A

P, B, M (closed)

viseme_PP

B

K, S, T (teeth)

viseme_kk, viseme_nn

C

EH, AE (vowels)

viseme_I, viseme_E

D

AA (wide open)

viseme_aa

E

AO, ER (rounded)

viseme_O, viseme_aa

F

UW, OW, W (puckered)

viseme_U

G

F, V (teeth-on-lip)

viseme_FF

H

L sound

viseme_TH, viseme_nn

X

Silence

viseme_sil

Integration with info-dump

This server pairs with info-dump for complete TTS → avatar rendering:

info-dump generate_audio("Hello!", voice, output_path)
    ↓
intentional-masking render_speaking_video(avatar_path, audio_path)
    ↓
MP4 video with lip-synced avatar

Avatar Requirements

Avatars must be Ready Player Me GLB files with standard viseme morph targets:

  • viseme_aa, viseme_E, viseme_I, viseme_O, viseme_U

  • viseme_PP, viseme_FF, viseme_TH, viseme_DD, viseme_kk, viseme_nn, viseme_sil

Create avatars at readyplayer.me

Development

# Run tests
npm test

# Watch mode
npm run dev

# Preview Remotion compositions
npm run remotion:preview

Environment Variables

  • INTENTIONAL_MASKING_ROOT - Project root directory (default: cwd)

  • INTENTIONAL_MASKING_OUTPUT - Output directory for rendered files (default: {root}/output)

License

MIT

Available Tools

1 tool
render_videoA

Render an avatar speaking with lip sync from audio, optionally with body animations

ParametersJSON Schema
NameRequiredDescriptionDefault
animationsNoOptional body animation timeline
audio_pathYesPath to audio file (from TTS like info-dump)
backgroundNoBackground color (hex)#1a1a2e
avatar_pathYesPath to the avatar .glb file
output_pathNoOptional output path (default: auto-generated in output/)
camera_presetNoCamera angle presetmedium
lighting_presetNoLighting stylesoft

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for disclosing behavioral traits. It only states what the tool does, not side effects (e.g., output file creation), processing requirements, or constraints (e.g., avatar format, animation requirements). This is a significant gap for a rendering tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It immediately states the action, subject, and key parameter relationships, making it highly scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 7-parameter schema with full descriptions, the description doesn't need to repeat parameter details. However, it omits the tool's output behavior (e.g., generates a video file) and any mention of the rendering process's cost or complexity. It's adequate for a simple tool, but not comprehensive for a rendering operation with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by clarifying relationships: audio drives lip sync, animations are optional, and the avatar is the subject. This contextual framing helps understand the key parameters beyond their individual schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Render') and clearly identifies the resource ('an avatar speaking with lip sync from audio'). It also mentions the optional body animations, which distinguishes it from a simple audio-to-video tool. Even without siblings, the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you need to generate an avatar video from audio, but it does not state prerequisites, alternatives, or explicit when-to-use/when-not-to-use guidance. The schema hints at audio source (TTS) but the description itself provides no direct usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedrender_video

TDQS

A3.6/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The single tool has a clearly defined purpose of rendering an avatar with lip sync.

Naming Consistency5/5

The tool name 'render_video' follows a clear verb_noun pattern, and with only one tool, there are no inconsistencies to evaluate.

Tool Count2/5

The server has only one tool, which feels too few given the implied scope from the server name 'intentional-masking' and the tool's rendering functionality. A single tool is unlikely to cover the expected domain.

Completeness1/5

The server name suggests a broader purpose (e.g., masking), but the only tool is a video renderer. There are no supporting tools for managing, previewing, or configuring renders, leaving the surface severely incomplete for the stated purpose.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers