Has anyone switched from traditional audio editing to prompt-first audio generation?
Summary: sarah wilson discusses their experience with audio editing for demo videos, noting that generating audio elements separately took longer than expected. They experimented with using AI to describe entire scenes and found it reduced editing time, although it required refining prompts rather than detailed sound effects. They used Seed Audio 1.0 for generating multiple elements, but are more interested in prompt engineering's impact on audio editing. They seek input from others about whether prompt-first approaches are replacing traditional methods in startup projects.
I'm curious whether anyone else has experienced this.
While preparing short demo videos for a startup project, I noticed something unexpected: editing the audio consistently took longer than creating the video itself.
My usual workflow looked something like this:
Write a short script
Generate narration
Search for background music
Add ambient sounds
Adjust volume levels
Export everything
None of these steps were difficult on their own, but together they added a surprising amount of time to every iteration.
Recently, I tried approaching the problem differently.
Instead of generating each audio element separately, I experimented with describing the entire scene in a single prompt and letting the AI generate dialogue, ambience, music, and sound design together.
The interesting part wasn't that the output was perfect—it wasn't.
The biggest improvement was that I spent less time editing and more time rewriting prompts.
That completely changed my workflow.
One thing I also learned was that adding more details didn't necessarily improve the results. My shortest prompts often produced cleaner audio than prompts packed with dozens of sound effects.
For these experiments, I happened to use Seed Audio 1.0, since it supports generating multiple audio elements within one workflow, but I'm much more interested in the broader question than in any specific tool.
For founders building demos, landing pages, or MVP presentations:
Have you found that prompt engineering is gradually replacing parts of traditional editing?
Or do you still prefer generating each audio component separately and assembling everything manually?
I'd love to hear how other startup teams are approaching this.