How I Use Gemini Omni to Plan AI Video Assets Before Final Production
Summary: Hirofumi Onde shares their experience using AI video tools, emphasizing that planning and decision-making are crucial parts of the process. They describe a structured workflow using Gemini Omni, which involves breaking ideas into sequences, distinguishing visual direction from motion, and generating key shots first. They note that even failed generations provide valuable insights. The author suggests that managing the creative workflow is becoming increasingly important in AI video production and seeks opinions on whether others use a similar structured approach or rely on experimentation.

AI video tools are improving quickly, but one problem I keep running into is that generating the final video is often not the hardest part.
The harder part is deciding what should actually be generated.
Before starting production, there are usually dozens of small decisions to make: the visual direction, shot order, camera movement, pacing, subject consistency, transitions, and which scenes are worth spending more generation time on. If these decisions are made too late, it is easy to burn through multiple generations without getting closer to a usable result.
Recently, I have been experimenting with a more structured workflow using Gemini Omni as part of the planning and generation process.
Instead of treating AI video generation as a single prompt-to-video step, I have started treating it more like pre-production.
Start with the sequence, not the prompt
One mistake I made early on was trying to write the “perfect prompt” immediately.
That sounds efficient, but in practice it often creates a beautiful individual clip that does not fit well with the rest of the project.
Now I first break the idea into a simple sequence.
For example, if I am creating a short product launch video, I might divide it into:
an opening establishing shot
a closer product-focused shot
one or two motion-heavy scenes
a transition scene
a final hero shot
At this stage, I am not worrying too much about exact wording. I am mainly trying to understand what each scene needs to accomplish.
This makes the generation process much more intentional.
Separate visual direction from motion
Another useful change has been separating two questions:
What should the frame look like?
and
What should happen inside the frame?
These are easy to mix together in a long prompt.
For the visual side, I think about composition, lighting, environment, lens feel, subject placement, and overall style.
For motion, I think about camera movement, subject movement, speed, direction, and how the shot should end.
Thinking about these separately usually produces clearer prompts and makes it easier to diagnose why a result did not work.
If the composition is wrong, I adjust the visual description.
If the scene looks good but the movement feels unnatural, I change the motion instructions instead of rewriting everything.
Generate the important shots first
I also stopped generating scenes in chronological order.
Instead, I generate the most important or most difficult shots first.
Usually there are one or two scenes that define whether the whole concept works. They may require a specific camera move, consistent character appearance, complex object interaction, or a very particular visual style.
If those scenes fail repeatedly, it is better to discover that early.
Once the difficult shots are working, the simpler transition and supporting scenes become much easier to build around them.
This approach has saved me quite a bit of unnecessary iteration.
Use failed generations as planning information
One of the underrated parts of AI generation is that bad outputs are still useful.
A failed result often tells you something about the idea itself.
Maybe there is too much action in a five-second shot.
Maybe the camera instruction conflicts with the subject movement.
Maybe the scene requires several important objects to remain consistent at once.
Instead of simply generating the same idea again, I try to identify which part of the request is creating uncertainty.
Sometimes simplifying the scene produces a much better result than adding more prompt details.
AI video is becoming a workflow problem
The more I work with these tools, the more I think AI video creation is becoming less about individual generations and more about managing a creative workflow.
The model obviously matters, but so does everything around it:
idea → shot planning → reference selection → generation → evaluation → revision → editing.
For founders, marketers, and solo creators, this may be especially useful because we often do not have a traditional production team separating these responsibilities.
AI tools can compress the production process, but they do not completely remove the need for planning.
In some ways, better generation models actually make planning more important because there are now far more creative directions available.
My current approach is therefore to spend slightly more time deciding what I want before generating anything.
The result is fewer random iterations and a much clearer path from an initial idea to something that can actually be published.
I am curious how other people working with AI video are approaching this.
Do you treat generation mostly as experimentation, or have you started developing a more structured pre-production workflow?