FLUX 3 Complete Guide

Davicho Barona
Davicho Barona

How to prompt FLUX 3: Directing motion, camera, and timing

FLUX 3 is Black Forest Labs’ first video model, built as part of a multimodal system trained across image, video, and audio.

It can generate video from text, animate images, work with keyframes, continue existing video, and generate synchronized audio.

But the most important thing to understand is how to talk to it.

With FLUX 3, think less like you are describing a picture and more like you are directing a shot.

What is happening?

How does the subject move?

What is the camera doing?

How should the scene progress over time?

A clear answer to those questions usually gives FLUX 3 more useful direction than simply adding more visual adjectives.

Black Forest Labs recommends defining explicit motion, intentional shot language, and a clear narrative rather than describing a loose collection of visual details.

Start with the event

A strong FLUX 3 prompt usually has one clear event at its center.

Instead of:

Cinematic woman in a rainy city, neon lights, beautiful reflections, moody, 35mm

Give the scene something to do:

A woman walks quickly through a rain-soaked night market as vendors pull tarps over their stalls. The camera tracks beside her at shoulder height while neon signs reflect across the wet pavement.

The second prompt establishes a subject, an action, a camera behavior, and environmental motion.

That gives FLUX 3 a sequence to resolve over time.

A useful principle from BFL’s prompting guidance is to favor concrete nouns and verbs that the camera could actually see.

Short prompts give the model freedom

You do not need a large prompt every time.

FLUX 3 can take something as simple as:

A red fox leaping through fresh snow, telephoto.

With a short prompt, the model makes more decisions about framing, movement, pacing, and atmosphere.

That can be useful while exploring.

Longer prompts become useful once you know what you want to control.

For example:

A red fox bounds through deep fresh snow at sunrise. Telephoto shot with strong background compression. The camera pans smoothly with the fox as powder explodes beneath each landing. Pale golden sunlight catches the airborne snow while the distant forest remains softly out of focus.

The principle is simple:

Start short enough to discover. Add detail when you need more control.

Longer is not automatically better. Packing too many instructions into the same shot can make the result harder for the model to resolve coherently.

A useful FLUX 3 prompt structure

When you need more control, think about the prompt in six layers:

  • Core idea: What happens during the clip?
  • Scene: Where are we? What are the lighting and atmosphere?
  • Subject: What needs to remain recognizable or consistent?
  • Motion + camera: What does the subject do, and how does the camera respond?
  • Audio: What speech, ambience, effects, or music belongs in the scene?
  • Style: What visual language, palette, texture, or level of realism holds everything together?

You do not need to label every section in the final prompt.

For a simple shot, those ideas can live naturally in one description:

Low tracking shot of a cyclist racing through a narrow European street at sunrise. The camera stays beside the bicycle as morning light flickers between buildings and the rider leans hard into a corner. Tires hum against the pavement, with a quick metallic rattle from the bicycle chain.

The important thing is that every detail has a job.

Give the camera one clear job

FLUX 3 understands traditional cinematography language, including:

shot size,

camera angle,

camera movement,

focus,

lenses,

lighting,

POV,

transitions,

and specialty camera setups.

That does not mean you should use all of them at once.

A useful rule is:

One framing choice + one camera movement + one subject action.

For example:

Close-up, slow push-in, the chef looks up from the cutting board as someone enters the restaurant.

Or:

Wide low-angle tracking shot following a horse running through shallow water.

Compare that with:

Low aerial handheld orbit tracking push-in around the horse.

Every phrase describes a real camera behavior, but together they compete with each other.

Use camera language to clarify the shot, not decorate the prompt.

Use timing when an action needs to land

When the progression of the shot matters, FLUX 3 also understands timestep prompting.

Instead of describing the entire clip in one paragraph, you can give it a simple timeline:

0–2s — Locked wide shot of an empty harbor before sunrise.

2–4s — Gulls lift from the water as the camera begins a slow push forward.

4–6s — The sun breaks over the horizon and warm light spreads across the boats.

You do not need to choreograph something new every second.

Two or three meaningful beats are often enough.

Timing is especially useful for:

reveals,

transformations,

product actions,

character beats,

and shots where something needs to happen at a specific point.

Multi-shot prompts can describe an edit

FLUX 3 can also create multiple scenes or camera angles inside a generation.

If you want cuts, describe them deliberately rather than leaving the edit entirely to the model.

For example:

SHOT ONE: Wide aerial of a desert highway at dawn. A red sports car races through the empty landscape.

HARD CUT.

SHOT TWO: Interior close-up of the driver's hands tapping against the steering wheel.

HARD CUT.

SHOT THREE: Static roadside telephoto shot as the car disappears into heat haze.

A single restrained electronic music bed continues across all three shots.

At this point, prompting starts to feel less like describing an asset and more like writing a miniature shooting plan.

A helpful reminder

More prompting does not automatically mean more control.

FLUX 3 already understands a lot about cinematography, movement, scene structure, and physical interactions.

A useful question while editing your prompt is:

What decision am I trying to stop the model from making for me?

If you care about the camera movement, specify it.

If a reveal needs to happen at a particular moment, use timing.

If you do not care about a decision, you may not need another sentence in the prompt.

Key takeaway

Treat FLUX 3 like a director treats a shot.

Start with a clear event.

Give the subject something to do.

Give the camera one clear job.

Use timing when specific beats need to land.

And only add more detail when there is a decision you actually want to control.

FLUX 3 audio and iteration: Directing dialogue, sound, and better generations

FLUX 3 can generate synchronized audio as part of the video generation itself.

That includes dialogue, ambience, sound effects, music, multilingual speech, and lip-synced performance.

This changes how you should think about prompting.

Sound is not something that necessarily needs to be added after the visual is finished.

It can be part of the scene from the beginning.

The same principle that helps with visual prompting also applies here:

describe what should actually happen, rather than relying on vague adjectives.

Describe sound sources

Instead of:

Moody ambience.

Try:

Rain ticking against the metal awning, distant traffic passing through wet streets, and the soft hum of the restaurant refrigerator.

The second prompt gives FLUX 3 identifiable sound sources.

For an action, you might write:

The ceramic cup clicks against the saucer as she puts it down.

Now the model has both a visible event and the sound associated with it.

Sound works especially well when it feels causally connected to the scene.

You do not need every audio layer

A generation does not need speech, music, ambience, and sound effects simply because FLUX 3 supports all of them.

Choose the layers the scene actually needs.

A quiet conversation might need:

dialogue,

subtle room tone,

and perhaps one or two environmental sounds.

A product film might use:

mechanical effects,

environmental sound,

and music,

without any speech at all.

The goal is not to fill every possible audio channel.

It is to make the scene sound intentional.

Prompt dialogue like performance direction

For spoken dialogue, put the words in quotation marks and make it clear who is speaking.

For example:

The mechanic looks toward the driver and says, "Try it now." His voice is low and matter-of-fact, spoken from several feet away inside the garage.

This gives FLUX 3 more than a line of text.

It gives the model:

a speaker,

the words,

a performance,

and an acoustic context.

Those details can influence how the dialogue feels.

Make off-screen speech explicit

If the speaker is not visible, identify the speech as narration or voiceover.

For example:

A calm female voiceover says, "Some roads are meant to be taken slowly."

Making this explicit helps separate spoken language from text that might otherwise be interpreted as something visible in the scene.

Direct voices with audible characteristics

Vague directions such as:

Professional, warm and engaging voice.

leave a lot of interpretation to the model.

Instead, describe characteristics that could actually be heard:

A woman in her early thirties speaking softly to a close friend, relaxed pace, close dry recording, lightly amused delivery.

Now the direction begins to resemble casting and performance notes.

Useful details can include:

age range,

pace,

energy,

distance from the microphone,

emotional delivery,

volume,

accent or language when relevant,

and whether the recording should feel intimate, distant, polished, or environmental.

Give dialogue enough time to exist

Speech takes time.

If the clip is short, the dialogue needs to fit realistically inside it.

Trying to combine a long script with several speakers, complex camera choreography, multiple actions, and a transformation gives FLUX 3 many things to resolve simultaneously.

Keep the amount of dialogue proportional to the duration.

When a line matters, give it room.

Think about continuity across shots

If you are generating a multi-shot sequence, audio can also help make the edit feel continuous.

For example:

SHOT ONE: Wide aerial of a desert highway at dawn. A red sports car races through the empty landscape.

HARD CUT.

SHOT TWO: Interior close-up of the driver's hands tapping against the steering wheel.

HARD CUT.

SHOT THREE: Static roadside telephoto shot as the car disappears into heat haze.

A single restrained electronic music bed continues across all three shots.

The visual perspective changes, but the shared music gives the sequence an audio throughline.

The same technique can work with ambience, narration, or recurring environmental sound.

Iterate by changing one decision at a time

Your first generation does not need to solve everything.

If the concept works but the camera does not, change the camera.

If the composition works but the movement feels weak, strengthen the action.

If the voice sounds too polished, change the performance direction.

If a keyframe transition feels unnatural, reduce the visual distance between the keyframes.

Small, targeted changes make it easier to understand what actually improved the result.

If you rewrite the entire prompt after every generation, you may get a different result without learning which change mattered.

Explore first, then lock decisions

A useful FLUX 3 workflow is:

Explore → identify what works → add control → refine.

Begin with a relatively simple prompt and see what the model gives you.

Then ask:

What should stay?

What needs to change?

What decision did the model make that I would rather make myself?

Add direction specifically for those things.

This is usually more productive than trying to predict every possible failure before generating anything.

Draft before committing to the final

At the model and API level, FLUX 3 also supports a draft-to-enhance workflow.

The idea is to explore using lower-cost draft generations, choose the version that works, and then enhance that selected generation at higher quality.

Importantly, enhancement is intended to preserve the chosen draft rather than reinterpret the original prompt as an entirely new take.

That creates a useful production mindset:

Do not spend final-quality resources while you are still deciding what the shot should be.

Explore first.

Choose the direction.

Then finish it.

Even when you are using FLUX 3 through a different interface, that broader workflow remains useful.

A helpful reminder

When something is wrong with a generation, try to diagnose the specific problem before adding more words.

If the movement is wrong, adjust the movement.

If the sound is wrong, change the sound direction.

If the performance is wrong, change the performance direction.

If the composition is already working, leave it alone.

The goal of iteration is not to make the prompt longer.

It is to make your intent clearer.

Treat sound as part of the scene, not an afterthought.

Describe real sound sources.

Direct dialogue like a performance.

Give speech enough time to fit naturally inside the clip.

Then iterate by changing one meaningful decision at a time.

FLUX 3 gives you a lot of control, but the most effective workflow is still simple:

explore first, lock the decisions that matter, and finish last.