Ray 3.2: Intro & Core Concepts
Ray 3.2 Modify Video is built for one clear job: transforming footage you already have.
You can access and use Ray 3.2 a few different ways:
- Through the agent chat window by selecting your video and restyled keyframes, and giving the agent natural language instructions.
- Via right-click on a video asset > Modify Video > Ray 3.2 – this is the most common and recommended approach, as it pulls up the full Ray 3.2 UI modal.
- Through the toolbar via Generate Video > Ray 3.2 which allows you to create a keyframe video in a simple video generation UI modal.
Ray 3.2 is not a text-to-video model. It is not an image-to-video animator. It is not an extender. Ray 3.2 starts with a source video and returns a re-imagined version of that same clip: same duration, new look, new material, new environment, new styling, or new visual direction.
That makes Ray 3.2 especially useful for production work. Instead of starting from a blank prompt, you begin with real motion, real framing, real edits, and real timing. Ray 3.2 uses the existing clip as the foundation, then lets you control how far the result should depart from the original.
Use it when you want to restyle footage, rescue a plate, explore look development, localize a campaign, create many variants from one master edit, or turn rough source material into a more finished visual direction.
The simplest way to think about Ray 3.2 is this:
Ray 3.2 is the production-grade V2V workhorse for transforming existing video while preserving the structure of the source.
What Ray 3.2 is best at
Ray 3.2 is strongest when you already have a video that contains something worth preserving: a camera move, a performance, a product angle, a scene layout, a timed edit, or a specific motion path.
From there, you can use it to create a new version of the clip without rebuilding the shot from scratch.
Common uses include:
- Turning stock footage into pitch-ready campaign visuals
- Restyling a product shot into a new material, colorway, or finish
- Converting live-action plates into anime, claymation, painterly, or graphic styles
- Updating wardrobe, signage, logos, props, skies, weather, or environments
- Creating seasonal, regional, or demographic variants from one master spot
- Exploring multiple VFX looks before committing to a final direction
- Creating polished temp comps while the final VFX pipeline continues separately
The key is that Ray 3.2 does not invent the whole shot from nothing. It works from the source clip. That is what gives it production value.
How Ray 3.2 differs from Ray 3.14 Modify
Ray 3.2 adds several major production controls compared with Ray 3.14 Modify.
The biggest difference is keyframes. Ray 3.14 supports a start-and-end style workflow. Ray 3.2 supports up to 64 keyframes at arbitrary source-frame indexes. This means you can anchor the look at specific moments in the source timeline instead of only guiding the beginning and end.
Ray 3.2 also gives you controls for preserving or transforming the source: Adherence, Characters, prompt, keyframes, Enhance, resolution, HDR, and inference mode.
Ray 3.2 also supports HDR video-to-video, EXR export for color grading and finishing, and Speed vs Quality inference modes. It also preserves the exact source duration, rather than trimming or looping the result to a fixed 5-second or 10-second output.
Use Ray 3.14 Modify only when you specifically need to change duration or create a loop. Use Ray 3.2 for standard V2V transformation, multi-keyframe guidance, HDR, EXR export, or deeper production control.
The three core inputs
Every Ray 3.2 job is driven by three primary inputs:
- Source video (Required)
- Keyframes (Optional with prompt)
- Prompt (Optional with keyframes)
The source video is required. Ray 3.2 always transforms an existing clip, and the output duration always matches the source duration, up to a maximum of 20 seconds. There is no duration parameter. A 7.3-second source produces a 7.3-second output.
The prompt describes the target end state. It should not describe a sequence of changes or a story over time. Ray 3.2 already has the timing and motion from the source video. The prompt tells it what the final frames should look like.
Keyframes are optional guide images attached to exact source-frame indexes. They are one of the most powerful controls in Ray 3.2 because they let you art-direct specific moments in the timeline.
You can use a prompt alone, keyframes alone, or both together.
A prompt alone creates a pure prompt-driven restyle. Keyframes alone create a visual-reference-driven transformation. Prompt plus keyframes gives the strongest control because the prompt defines the creative intent while the keyframes lock specific moments.
You need at least one of those two: a prompt or keyframes. A source video with neither will error.
Why keyframes matter
Keyframes are the biggest new control surface in Ray 3.2.
A keyframe is a still image that tells Ray 3.2 what the output should look like at one exact frame in the source video. You can provide Ray3.2 with up to 64 keyframes.
This makes keyframes extremely useful for production. You can export a few frames from the source, edit them with 3rd party software like Photoshop or Nuke, or edit them directly in Luma with any image editing model like Uni-1, Nano Banana, GPT Image, then feed them back into Ray 3.2 at the same original source-frame indexes. Ray 3.2 uses those anchors to interpolate the look across the rest of the shot.
That workflow is especially useful for art-directed restyles, product close-ups, music-video beats, brand-safe openings and endings, and VFX look development.
Adherence
Adherence protects the original. Use it for targeted edits, subtle changes, and cases where the source needs to remain highly recognizable.
Adherence is useful when you want to preserve the original shot’s layout, camera movement, body motion, product position, or general scene structure while changing the look, style, material, environment, or surface details.
Adherence includes separate controls for Motion and Structure.
Both Motion and Structure controls use a numeric scale from 1 to 9, where 9 represents the strongest adherence/preservation.
Motion controls how strongly Ray 3.2 preserves the movement of the source clip. Higher motion adherence keeps the timing and motion behavior closer to the original. Lower motion adherence gives the model more freedom to reinterpret how movement feels through the transformation.
Structure controls how strongly Ray 3.2 preserves the spatial arrangement of the scene. Higher structure adherence protects layout, geometry, camera framing, and object placement. Lower structure adherence gives the model more freedom to reshape the scene.
When uncertain, start near the balanced middle and adjust based on what is drifting. If the motion feels too loose, increase Motion. If the scene layout is changing too much, increase Structure.
Characters
The Character controls decide how Ray 3.2 holds a person or animal together through a transformation.
These controls act like independent locks. Each one you enable tells the model that a specific attribute of the subject must survive the restyle. Toggle on what matters. Toggle off what you want freed up.
The Character controls stack together. Faces, Bodies, and the skeletal tracking mode can work together to preserve identity, physical build, and pose behavior.
Faces
Faces locks facial identity and expression. It helps preserve the geometry that makes someone recognizably themselves, plus the expression they are making in the source: a smile, a squint, a reaction, a snarl, or a subtle emotional beat.
Turn Faces on when the performance lives in the face, such as dialogue, reactions, close-ups, or emotional scenes. It is also useful when you are changing the look or style, but the same person must remain the same person.
Turn Faces off when there are no faces in frame, when the faces are tiny or distant, or when you are replacing the character entirely. For example, if you are changing a person into a robot, creature, or different character, keeping Faces on can fight the swap and pull the old face back into the result.
Faces is one of the most important locks to check. If identity is drifting, check whether Faces should be on. If a character swap will not take, check whether Faces should be off.
Bodies
Bodies locks the overall body: silhouette, build, proportions, and broad pose or posture in space.
Turn Bodies on when the subject’s physical presence matters. This is useful for wardrobe restyles, athletic performances, dance, fashion shots, or any transformation where the same body shape should carry through the look change.
Turn Bodies off when you want to change the body type itself. For example, turn it off when changing adult proportions to child proportions, a human into a robot chassis, or a creature into a much slimmer or bulkier form.
Faces and Bodies are separate on purpose. You can lock Faces while freeing Bodies if you want to preserve someone’s identity but reshape their physique. You can lock Bodies while freeing Faces if you want to preserve a build or performance while changing who the character is.
Poses and Blocking
Poses and Blocking are skeletal tracking modes. They both help Ray 3.2 follow the subject’s pose, but they track motion at different levels of detail.
Blocking is the sparser, more flexible mode. It shows the model the key joint positions without the connecting bone lines. It is the lighter signal, allowing more room to diverge from the source form while still retaining a loose sense of pose.
Use Blocking when you want the motion to be loosely guided by the source, but not tightly locked to it. This is useful when mapping a human performance onto something less anatomically humanoid, like a creature, robot, abstract character, or energy form.
Putting Characters controls together
For a dialogue close-up with a style change, use Faces on, Bodies on, and Blocking. This preserves identity, expression, and broad posture.
For a pianist’s hands while restyling the scene, Faces may be optional, Bodies should usually stay on, and Poses is the better skeletal mode because it is the heavier, more adherent mode, giving the model the most detailed skeletal tracking to preserve the performance.
For turning a dancer into a glowing energy being, Faces should usually be off if the identity is changing. Bodies can stay on or off depending on whether the performer’s build should remain. Blocking is usually the better choice for sweeping choreography.
For a crowd of distant extras, consider turning Faces off. If the faces are too small to matter, the lock may add unnecessary constraint without improving the result.
Ray 3.2: Prompting, Outputs & Controls
How to write Ray 3.2 prompts
Prompting Ray 3.2 is different from prompting a text-to-video model.
With text-to-video, the prompt often needs to describe the full scene, motion, timing, and cinematic structure. With Ray 3.2, the source video already provides the timing, motion, composition, and camera behavior.
Your prompt should describe the target end state.
Do not write commands. Do not describe the transformation process. Do not describe what happens over time. Describe the final image qualities that should exist in the transformed video.
Instead of:
Change the sky to purple.
Use:
A purple sky over the mountain landscape.
Instead of:
Remove the people from the street.
Use:
An empty cobblestone street at dawn.
Instead of:
The colors shift throughout the shot.
Use:
Saturated teal-and-orange color palette.
Ray 3.2 also responds better to positive phrasing. Avoid “no,” “not,” “without,” “devoid of,” or “missing.” Negation can bias the model toward the thing you are trying to exclude.
For removal, describe the empty result.
For style transfer, describe the finished style.
For material swaps, describe the final material.
For wardrobe changes, describe what the subject is wearing.
For environment swaps, describe the new environment directly.
A strong Ray 3.2 prompt often ends with a preservation cue:
Preserve all other elements including subject identity, pose, and lighting.
Drop “and lighting” when the lighting is the thing you are changing.
Enhance
Ray 3.2 now includes an Enhance toggle for prompts.
Enhance is on by default. When enabled, Ray 3.2 can rewrite your prompt with V2V-style guidance before the run. This helps turn a rough instruction into a stronger transformation prompt that better fits how Ray 3.2 interprets source footage.
Use Enhance when you have a simple idea, a rough instruction, or a prompt that has not yet been optimized for V2V. It can help clarify the target look and make the request more model-friendly.
Turn Enhance off when you have already crafted a careful prompt and want it sent to the model exactly as written. This is especially useful for production workflows where phrasing, constraints, brand language, or preservation cues have already been tested and approved.
A simple rule:
- Rough prompt or fast exploration: Enhance on.
- Carefully written production prompt: Enhance off.
Prompt templates
Material swap
A [subject] made of [target material], with [surface qualities]. Preserve all other elements including subject identity and pose.
Example:
A sports car made of brushed black titanium, with subtle satin reflections and crisp panel lines. Preserve all other elements including subject identity and pose.
Environment swap
[Subject] in [new environment], [time of day], [weather]. Preserve subject identity, wardrobe, and pose.
Example:
A runner on a quiet Tokyo street at night, wet pavement, soft neon reflections, light rain. Preserve subject identity, wardrobe, and pose.
Style transfer
[Scene description] rendered as [target style]. Preserve composition and motion.
Example:
A city street scene rendered as hand-painted anime background art, clean linework, soft cel shading, graphic color blocks. Preserve composition and motion.
Wardrobe or product change
[Subject] wearing [new wardrobe / holding new product]. Preserve identity, environment, pose, and lighting.
Example:
A presenter wearing a tailored cream linen suit and holding a matte black smartphone. Preserve identity, environment, pose, and lighting.
What to avoid in prompts
Avoid imperative commands like “make,” “turn,” “change,” or “transform.” Describe the result instead.
Avoid temporal language like “throughout,” “as it moves,” “over time,” “when,” “then,” or “gradually.” Ray 3.2 is using restyled source frames, not writing a new story timeline.
Avoid negation. Use positive descriptions.
Avoid long mood essays. Ray 3.2 rewards specific, declarative visual direction.
Avoid generic adjectives that do not add concrete visual information. Words like “whimsical,” “dreamy,” “vibrant,” or “hyper-realistic” can reduce precision if they are not grounded in specific details.
Output controls
Ray 3.2 supports four resolution options:
- 360p draft
- 540p
- 720p
- 1080p
You may not need to generate in 1080p too early unless the job is ready for final review or delivery.
A practical ladder is:
- Use 360p or 540p for internal exploration
- Use 720p for polished review
- Use 1080p for final delivery, hero shots, broadcast, or social finals
Ray 3.2 also supports HDR and EXR export.
HDR is off by default. Turn it on for premium brand work, automotive, fashion, broadcast, streaming, modern OLED delivery, or anything that needs expanded dynamic range.
EXR export requires HDR. When enabled, Ray 3.2 attaches an EXR ZIP variant to the output. This gives downstream finishing tools such as Nuke, Resolve, Flame, or Baselight a frame sequence with full linear floating-point color data. Use this when the output needs professional color grading or VFX finishing.
Inference mode gives you another production lever:
- Quality mode is the default and should be used for finals.
- Speed mode is faster for quicker turnaround, useful for exploration, quick-turn review rounds, and batch testing.
A reliable production pattern is to explore in Speed mode at 540p, then render the selected direction in Quality mode at the final resolution.
Cost-aware iteration
Ray 3.2 pricing scales by resolution and output duration.
Inference Mode (Speed vs. Quality) saves time, but not cost.
Because the output duration always matches the source, the length of your source clip directly affects cost. A 10-second clip costs twice as much as a 5-second clip at the same resolution.
Resolution also matters dramatically. A 10-second clip at 1080p costs much more than the same clip at 540p. That means you should avoid using final settings for early exploration.
For most workflows:
- Start at 720p
- Use Speed mode while exploring
- Test several directions cheaply
- Choose the strongest result
- Escalate only the winning direction to 1080p, HDR, EXR, or Quality mode
This makes Ray 3.2 practical for creative exploration without burning final-render costs on options that will not ship.
Auto controls
Adherence controls are set to Auto ON by default. Most Ray 3.2 users can leave Auto on until there is a reason to change settings.
Adherence controls
Adherence controls how strongly Ray 3.2 preserves the source video while transforming it.
Both Motion and Structure controls use a numeric scale from 1 to 9, where 9 represents the strongest adherence/preservation.
The updated UI includes Motion and Structure controls.
Motion controls how closely the result follows the movement in the source clip. Increase Motion when the timing, gesture, camera movement, or action needs to stay close to the original. Lower it when the transformation needs more freedom.
Structure controls how closely the result follows the spatial layout of the source clip. Increase Structure when composition, geometry, object placement, or scene layout needs to remain stable. Lower it when you want Ray 3.2 to reinterpret the scene more freely.
Use Motion when movement is the thing that must survive.
Use Structure when layout is the thing that must survive.
Use both when the source clip needs to remain highly recognizable.
Character controls
The Characters group is where Ray 3.2 decides how it holds a person or animal together through a transformation.
The big idea: Characters controls are independent locks on different aspects of a subject. Each one you enable tells the model that attribute must survive the restyle. Toggle on what matters and toggle off what you want freed up.
They stack together. Faces, Bodies, and a skeletal mode work as a combined preservation system.
Faces
Faces locks facial identity and expression.
It preserves the geometry that makes someone recognizably themselves, plus the moment-to-moment expression: a smile, a squint, a snarl, a reaction, or an emotional beat.
Turn Faces on when:
- The performance lives in the face
- You are working with dialogue, reactions, close-ups, or emotional beats
- You are changing the style or look, but the same person must stay the same person
Turn Faces off when:
- There are no faces in frame
- You are replacing the character
- Faces are tiny or distant and not the point of the shot
Faces is the single most common control to check when identity results feel wrong. If identity drifts, try turning Faces on. If a character swap will not take, try turning Faces off.
Bodies
Bodies locks the overall body: silhouette, build, proportions, and broad pose or posture in space.
Turn Bodies on when:
- The subject’s physical presence matters
- You are doing wardrobe restyles
- You want to keep an athlete’s build or a recognizable stance
- You want the same body shape carried through a look change
Turn Bodies off when:
- You are changing the body type itself
- You are changing adult proportions into child proportions
- You are turning a human into a robot chassis
- You are slimming, bulking, or reshaping a creature
- No coherent body is present
Faces and Bodies are separable on purpose.
Lock Faces and free Bodies when you want to keep someone’s face while changing their physique.
Lock Bodies and free Faces when you want to keep a build, pose, or performance but change who the character is.
Poses vs Blocking
Poses and Blocking are skeletal tracking modes. They both help Ray 3.2 understand the subject’s pose, but they give the model different amounts of skeletal information. Joints is the default and more forgiving option.
The important difference is signal strength.
Poses is the heavier, more adherent mode. It includes the joints plus the bones connecting them, so the model sees a fuller skeletal structure: not just where the key points are, but how the limbs connect across the body. Because Poses gives Ray 3.2 more visible pose information, it creates a louder control signal and holds the body closer to the source.
Use Poses when the original pose, body mechanics, or choreography needs to stay strongly intact. It is useful for full-body performances, dance, walking, running, fight choreography, sports movement, broad gestures, and shots where the body’s structure should remain recognizable through the transformation.
Poses is the better choice when preservation matters. It gives the model more strict guidance, so it is less likely to drift away from the source body form.
Blocking is the sparser, more flexible mode. It shows the model the key joint positions without the connecting bone lines. Because the model sees fewer skeletal features, it has more room to diverge from the source form while still retaining a loose sense of pose.
Use Blocking when you want the result to follow the general pose but not be tightly bound to the original body structure. It is useful for stylized character transformations, body reshaping, creature or robot conversions, exaggerated proportions, or any case where the source pose should guide the result without over-constraining it.
Blocking is not the “more detailed” mode. It is the lighter signal. It can be useful precisely because it gives Ray 3.2 less skeletal information, allowing the transformation to move farther from the original body shape.
The quick rule:
- Stronger pose/body adherence: use Poses.
- More freedom to diverge from the source body form: use Blocking.
- Unsure: start with Blocking, and switch to Poses if the pose drifts too much.
Practical Characters recipes
Dialogue close-up, style change
Faces: on
Bodies: on
Poses: Blocking
This locks identity and expression while preserving broad posture.
Pianist’s hands, restyle the scene
Faces: optional
Bodies: on
Poses: Poses
Poses is used here because its strong adherence is needed to preserve the subtle and complex choreography of the hands, which is the point of the shot.
Turn a dancer into a glowing energy being
Faces: off if identity is being replaced
Bodies: on/off depending on whether the dancer’s build should remain
Poses: Blocking for the sweeping choreography.
Crowd of distant extras
Faces: off
If faces are too small to benefit, the lock may add overhead or fight the transformation without improving the result.
A practical Ray 3.2 workflow
Start by choosing the right source clip. Keep it as short as the job allows, because cost scales with duration.
Next, decide what needs to change and what must remain protected. If the shot mostly needs a subtle fix, increase preservation through Adherence and Characters. If it needs a meaningful but recognizable transformation, keep Motion and Structure balanced. If it needs a stronger creative reinterpretation, give Ray 3.2 more freedom through lower adherence, prompt direction, and keyframes.
Write a prompt that describes the end state. Avoid commands, temporal language, and negation. Add a preservation cue at the end when identity, pose, lighting, or layout should remain.
Use Enhance when you want Ray 3.2 to improve a rough prompt for V2V. Turn Enhance off when you have already written a careful prompt and want it used exactly as written.
If the look needs art direction at specific moments, export frames from the source, paint or generate the desired stills, and reattach them as keyframes with exact source-frame indexes.
Start with 720p for balanced iteration. Drop to 540p when you want cheaper and faster exploration. Use Speed mode when turnaround matters. Once the team chooses the strongest direction, move to Quality mode and the final resolution.
Turn on HDR only when the deliverable needs expanded dynamic range. Turn on EXR only when the output needs a professional finishing or color-grading handoff.
Use Character controls intentionally when people or animals matter. Lock Faces for identity and expression. Lock Bodies for silhouette, proportions, and broad posture. Use Blocking for stable full-body motion and Poses for detailed hand or joint articulation.
When not to use Ray 3.2
Do not use Ray 3.2 when you need to generate video from text alone. Use a text-to-video model instead.
Do not use Ray 3.2 when you need to animate a still image from scratch. Use another model like Ray 3.14 instead.
Do not use Ray 3.2 when you need to extend a clip. Use the dedicated Extend feature (available on Ray 3.14) instead.
Do not use Ray 3.2 when you need to change duration or create a loop. Use Ray 3.14 Modify for that.
Ray 3.2 is at its best when the source footage matters and the transformation needs to stay grounded in that footage.
Key takeaway
Ray 3.2 is Luma’s production-grade video-to-video transformation model. It keeps the duration, motion, and structure of your source video while giving you powerful controls for restyling, keyframing, output quality, HDR, EXR, and advanced conditioning.
Start with the source. Describe the end state. Choose the right Adherence settings. Use keyframes when art direction matters. Keep Auto On until you have a reason to change settings – that is the core Ray 3.2 workflow.
Ray3.2 Workflows Deep Dive
What each control does, how to read its slider, and how to combine them.
How to think about the controls
Modify Video gives you a set of controls that tell the model how closely to follow different parts of your source footage. They fall into three families: Motion (how movement carries over), Structure (how tightly shapes and forms are held), and Characters (how a performance is captured). Each works independently, so you choose which parts of your footage matter and let go of the rest.
Ask yourself: how much do I care about the movement here? How much about the shapes, structure or edges? How much about the performance? Turn each control up to match, Start simple — run the lightest setup first, see what's missing, and add a control only to fix that specific gap.
One more thing worth knowing up front: with this system — especially when you're using multiple keyframes — your settings often matter more than your prompt. The model is good at inventing a believable look on its own to bridge your keyframes; the controls are what hold it to your footage. It's often good to make improvements with prompts, but you may find that saving the keyframes that you created from text and then making more detailed changes on top of them end up being even more valuable.
Motion
Range: Off, or 1–9.
Motion decides how much of your source footage's movement is carried into the output, and how densely. It's a separate system from Structure: it pays attention to movement across the whole scene rather than to shapes or edges. The simplest way to read the slider is as motion density.
Reading the slider
- Low (around 1) — broad strokes. The model picks up only the largest, general movement. In practice it mostly locks onto your camera move, because with few moving things there's little else to hold. Ideal when you want the camera nailed but want freedom in how things move within the frame.
- High (around 9) — dense capture. The model grabs as much movement across the frame as it can in bursts, so small or fleeting movement (an individual limb, a butterfly passing through) is far more likely to survive If it happens to be in a frame where points are propagated.
Good to know
- Overlapping movement drops out. When two moving things cross or pass in front of each other, that overlapping movement tends to be lost. Fast, crossing action won't all carry through — you keep the broad back-and-forth, not every detail.
- Closeness preserves detail. A subject that fills more of the frame and moves less holds its movement far better than something small and fast. A creature filling the frame can carry surprisingly specific motion — effectively acting like motion capture for non-human subjects — while a small object flying past may carry only a trace of its movement by the time it exits.
- It handles motion blur well. Where traditional motion tracking tends to break down on blur, Motion stays reliable.
- It loves gradients. Motion responds to gradients of color and light rather than flat shapes. If you drop a placeholder object into a scene (a stand-in for a creature or prop), give it shading or a gradient — a flat, single-color blob gives weaker motion control than a shaded one. 3D objects work naturally here, because their lighting creates gradients on its own (even rough lighting helps).
Reach for Motion when you care about reproducing movement not shape — a camera move, an object's path, the motion of a performance — more or less independently of the exact shapes involved. Even at the highest settings, small things in the frame can feel like they are very generally following the overall motion of them instead of being super specific.
Structure
Range: Off, or 1–9.
Heads up on direction: higher means more adherence. This direction was flipped recently, so if you're used to the older behavior, note that bigger numbers now mean a tighter hold on your footage.
Structure controls how tightly your output sticks to the shapes and forms of your original footage. Low lets the model reinvent shapes freely; high locks the output to what you shot. A useful mental image: at high settings it's as if the whole scene were shrink-wrapped, so everything stays glued in place.
Reading the slider
- Low (around 1–3) — very soft and loose. The model works from soft blobs: a rough sense of where things sit in the frame relative to one another, with little read on the camera. It mostly just keeps things roughly where they belong. Use it when you want to transform shapes heavily but keep approximate placement. Around 1 is the loosest, broadest version.
- Middle (around 5) — balanced. Real shape detail comes through, but it's still loose enough to transform the look.
- High (around 8) — tight on the subject. A firm hold on your subject's shapes with a quick fall-off behind them: roughly, anything more than a few feet back drops away to changing based on the motion of the foreground, not really adhering to the shapes of the background. This probably is the closest to a “Green screen” setting — the subject is held firmly while the background is freed up — and it's often the sweet spot for close-ups. A side note: The green screen effect is an analogy. The system does not notice color at all, so you are better off having a bookshelf and a lamp ten feet behind you than you are having a bright green wall four feet behind you.
- Top (around 9) — whole-frame hold. The model holds onto every shape, angle, and edge across the entire frame, regardless of distance: the most aggressive “give me exactly what I shot” setting. This is a genuinely different mode than 8, not just more of it — it stops caring about distance and locks the whole frame.
Good to know
- Override it with keyframes. Even at the top setting you can fight the hold in specific areas with keyframes. Keep most of the scene locked while you change one character or one background element — add enough keyframes and you'll win against the structural hold in those regions.
- A nuance for faces. At the very top setting, fine facial detail can actually come through slightly less than at 8 in some respects. For close-up faces, try 8 first (tight subject, With less concern for background) before jumping to the top.
Reach for Structure when shapes and forms matter — keeping a location, a silhouette, a face, or a product looking like the real thing.
Characters
Face
Face captures facial performance. With it on, the model reads your expression as a set of proportions — how far the mouth corners lift, how much the eyes squint, and so on — and transfers that to the output character. It works well even when the output character looks nothing like you.
Bodies & Poses
Pose captures the performance of people (or stand-ins) in your footage. Bodies and Poses together capture a person's body and limb movement, and you turn them on independently of Motion and Structure.
Poses-only is powerful. With just Bodies + Poses on (Motion and Structure off), the model focuses purely on the person you're driving and captures their motion tightly. As a bonus, it's quite good at inferring the camera move from how a body moves. That makes pose-only especially handy in two situations:
- Cramped spaces. You don't drag in walls, furniture, or other clutter as noise — only the body.
- Keeping the rest of the scene free. A character's full motion carries through while the rest of the scene, including other characters, is left to do its own thing.
Poses gives you the full performance. It carries every actual movement, and the model is robust enough to smooth over small jitters and assume a natural continuation of human motion.
Blocking is for looser motion blocking. Blocking is the broader pose mode (think block-out or pre-vis level) that captures only the general back-and-forth of a body rather than every limb. It's useful when your performance is a rough stand-in and you'd rather the model generate natural motion than follow your rough puppeteering exactly. For most performance work, stick with Poses.
Blind spots worth knowing
Face pays less attention to the center of the forehead and the upper cheeks, so very subtle cues there (a faint nose wrinkle, fine brow movement) may not fully carry. Broad expressions transfer best — though on a tight close-up a big smile can read as someone else's smile. To preserve a real performance on a real face at the highest fidelity, lean on high Structure (8, sometimes the top) with Face on.
Good to know
- Face-only is an option. Run Face by itself to puppeteer just a character's expression — drive a wild performance with keyframes and nothing else.
- Turning yourself into a creature. For a close-up where you become a different character or creature, use little or no Structure plus Face, so the performance carries without locking you to your own shapes.
Combining controls
Combining is the hardest part, and the guiding rule is simple: turn off anything you don't need, Turn settings on to help you tune towards the things you're wanting the model's attention on. If you want a tight facial performance and accurate motion on a chimpanzee, that may be nothing more than Face + Motion — adding Structure or extra toggles may only pull the model’s focus.
- Balance camera against subject. Decide how much you care about structure versus motion and dial them independently. “I need the camera nailed but the character can be loose” suggests low Motion (around 1) for a strong camera lock plus low Structure (1–2) to keep the character free — a solid camera move with a loose subject.
- Preserve a real face. Go high Structure (try 8 before the top) with Face on.
- Recolor or tattoo the same face. Keep Structure high — you're keeping the same shapes, just changing the surface.
The big advantage of separating Motion from Structure is that you can hold a movement exactly while freely changing shapes. Push Motion to keep the motion and keep Structure low, and you can — for example — keep a camera move and a performance while giving a character an entirely new silhouette. The older shape-based approach couldn't do that without losing either the motion or the camera.
Coming from Adhere / Flex / Reimagine
If you're used to the previous modes, these landmarks help you translate:
- Structure is the old Adherence scale, reversed. Higher now means more adherence, not less.
- Reimagine 2 ≈ Motion off, Structure off, Bodies + Poses, Face on. Match it by turning Structure off, not by setting Structure to 2 — the old Reimagine 2 had no structural component at all.
- Rough endpoints. Bodies + Poses with Structure at the top is close to the old Adhere (maximum adherence); Structure around 1 lands closest to Reimagine 3.
- Motion is new. It has no equivalent in the old scale, so don't try to map it onto Adhere, Flex, or Reimagine. In some ways, it's the most advanced tool in the toolkit because you can keep some very clear control of broad motions without also having to adhere to really tight edges from the original footage. The downside is that it operates in bursts, so you can leave your points behind by panning away without keeping anything in the frame. In this case, structure can actually end up saving you.
Quick reference
A control-by-control summary. Treat the numbers as starting values and test from there.
Control & range
Motion
Off / 1–9
At the low end: Broad strokes; mostly the camera move and some large object motions
At the high end: Dense capture; lots of motions are picked up across the frame
Use it when: You care about reproducing movement — a camera move, an object's path, a performance's motion
Structure
Off / 1–9
At the low end: Soft blobs, free to reshape; little camera read
At the high end: Tight, whole-frame hold — “exactly what I shot”
Use it when: The shapes and forms matter — a location, silhouette, face, or product
Bodies + Poses
On / Off
Off: The body is nothing more than another motion or shape in the frame.
On: full body & limb performance is held tightly; also infers the camera; superb on its own in tight spaces
Use it when: A person's motion is what matters and you.
Faces
On / Off
Off: No special notice of performance outside of motion or shape
On: transfers expression as amounts of different facial motions, even onto a very different character
Use it when: You want the facial performance carried over, on your character or a new one
Rule of thumb: Try the simplest version first, then add one control to fix what's missing — and save the settings you used alongside any clip you'll want to reproduce.

