Motion Capture
Recording real-actor motion
Also known as: Mocap · Performance capture
Definition
Motion capture records the movement of a real actor's body and face, then applies it to a 3D character's rig. Modern systems range from full optical studios (Vicon, OptiTrack) costing six figures, to inertial suits (Rokoko, Xsens) for indie budgets, to AI-from-video tools (Move.ai, Rokoko Vision) that need only a single camera. Mocap delivers animation volume that keyframe alone cannot match, especially for cinematic sequences.
In production
Raw mocap data is rarely usable directly — it always needs cleanup for jitter, foot sliding, and retargeting to the game rig's proportions. Cascadeur and MotionBuilder are the standard tools for that cleanup phase, with MotionBuilder's Story tool commonly used to retarget a capture session onto rigs of different proportions. Facial mocap is captured separately, usually via a head-mounted camera (Faceware) or an iPhone's TrueDepth sensor through Unreal's Live Link Face app, then solved to blend shapes or a FACS rig. Full-body optical setups like Vicon require reflective markers and a calibrated multi-camera volume, delivering sub-millimeter accuracy but needing hours of studio setup per session, while inertial suits like Xsens drift over time and need periodic re-calibration against a T-pose. A classic pitfall is "foot sliding," where the captured root motion doesn't match the character's actual stride length, forcing animators to manually key foot-lock constraints frame by frame.