EP 040 · 75 MIN VEO 3 MAKES YESTERDAYS BEST LOOK LIKE BETA
Veo 3 Makes Yesterday's Best Look Like a Beta Test +Runway & Midjourney Check-In
The short version
Google's new Flow workspace bundles Veo 3, Gemini 2.5, and Imagen 4, and Rory Flynn calls its prompt coherence and physics a clear step up from Veo 2, though Extend still falls back to Veo 2 and drops sound and quality. Rory and Drew Brucker also dig into Runway's References feature for compositional control, Figma's new AI website builder, and Midjourney's office-hours roadmap: faster V7, an improved --sref, and a video model still roughly two months out.
Key takeaways
- 01
Prototype cheaply on Veo 2 Quality before spending Veo 3 credits: Rory burned through ~12,500 monthly credits in two days at roughly 150 credits per generation.
- 02
Extending a Veo 3 clip silently falls back to Veo 2: you lose sound and sharpness, and the footage skews noticeably bluer, Rory found.
- 03
Runway's References feature turns rough wireframe sketches (boxes labeled 'car here, person here') into accurately composed shots, complete with matched lighting and color.
- 04
Midjourney's video model won't fully launch for roughly two months, per office-hours notes; both hosts want text-to-video first, then image-to-video, to avoid a Sora-style feature dump.
- 05
Turn a reference image into a reusable JSON prompt template, then swap only the variable fields (breed, color, material) to batch-generate a consistent icon set fast.
- 06
Midjourney's near-term roadmap, per office hours: a faster V7 next week, an improved --sref (without SREF Random at first), then multi-subject referencing.
- 07
Rory still prefers Midjourney's unpredictability to fully coherent models: he wants it to 'spit out 30% of what I didn't prompt for,' not total control.
Receipts · said on the record
“I think we went through a real OG era of analog-like AI, with the level of context and things that we had to really apply there, especially after that first wave.”
“You need internal adopters that are just curious by nature in the company. And you need multiple people at the moment, because most companies are not going to spend at the level that they would need to to upskill their own employees.”
“I think people just start to sleep on Runway. Candidly, the quality is not on a Veo 3 level or a Kling 2.0 level. But sometimes you don't need that. Sometimes you just need something fun like this, and Runway can do that.”
“I can't get anything that's, you know, Mid Journey imagination quality on any other tool. There is no other tool that has that magic.”
“And sometimes I want new ideas. Like I want Mid Journey to spit out 30% of what I didn't prompt for. Yeah, so I can take it.”
Chapters
- 0:00:00 When did this madness begin?
- 0:02:19 AI video is finally getting spicy
- 0:03:29 Google’s Flow Suite: Veo 3, sound, and coherence
- 0:05:02 Google's confusing product soup: Flow, Gemini, Imagen, Whisk
- 0:10:45 Pricing pain: Is Veo 3 worth the $125?
- 0:13:09 Veo 2 vs Veo 3: Best value tips and tradeoffs
- 0:15:08 Prompt accuracy and physics: Is Google really listening?
- 0:17:53 Why less prompt effort = better results now
- 0:19:40 Veo 3 vs Kling vs Midjourney: Prompting philosophies
- 0:20:52 Scene builder: Longer takes and smart extension workflows
- 0:22:34 The catch: extending drops quality and loses sound
- 0:24:17 New image-to-video support + third-party images
- 0:25:41 Ingredients-based generation and persistent characters
- 0:27:10 Frame extraction: finally, a feature we all needed
- 0:28:08 Timeline editing, upscaling, and staying inside the tool
- 0:29:48 Sora vs Veo 3 vs Runway: usability and consistency
- 0:31:43 Canva, Figma, Framer: Tools are becoming monsters
- 0:35:33 Figma’s new AI website builder is wild
- 0:36:40 Prompting sneaker ads and JSON-based design
- 0:37:09 Why training teams on AI is almost impossible
- 0:38:07 Hedra who? Veo 3 makes fast pivots a must
- 0:39:55 Midjourney's next move: what video could look like
- 0:41:11 Runway’s underrated features and clever reference hacks
- 0:44:26 Scene sketching and layout prompting: mind blown
- 0:47:25 Interior design from mood board to layout to render
- 0:49:45 Lighting direction via floorplans = next-gen hack
- 0:52:53 Try-on tech and Chrome extensions
- 0:54:22 Style consistency with JSON + ChatGPT
- 0:58:23 Mass-generating stylized icons and dogs with jobs
- 1:02:36 Midjourney updates: V7.1, personalization, and video
- 1:05:01 What Midjourney must get right with video
- 1:07:18 The one-shot window to impress
- 1:09:23 Bring back the Midjourney magic
- 1:11:14 Wrap-up: chaotic times, coherent thoughts, caffeinated takes
Questions this episode answers
How does Veo 3 compare to Veo 2, and what does it cost?
Rory says Veo 3's output quality and prompt coherence beat Veo 2, especially on physics and detail, though Veo 2 was stronger on extreme motion. Pricing starts around $125/month (roughly 12,500 credits, about 150 per generation) and jumps to about $250 after an introductory period, so Rory recommends prototyping cheaply on Veo 2 Quality before spending Veo 3 credits.
What is Runway References, and what can it do?
Runway References lets creators feed in a rough wireframe sketch, labeled with subjects and positions, and get an accurately composed shot back, complete with matched lighting, color, and style. Rory and Drew show it handling multi-subject scenes, a virtual try-on Chrome extension, and an interior-design mockup built from a floor-plan sketch plus a mood-board reference image.
When is Midjourney's video model launching?
Per office-hours notes read on the show, internal testing by guides and moderators was set to begin in the final days of May, with text-to-video and image-to-video both in development across multiple pricing tiers. A full public launch depends on server infrastructure upgrades expected in about 45 days, so Rory and Drew estimate real availability is roughly two months out.
What does Midjourney need to get right with video, according to the hosts?
Rory argues Midjourney needs a real differentiator beyond raw video quality, specifically bringing its style-reference, personalization, and 'ingredients'-style tools into video rather than just rules-based camera controls. Both hosts also want a simple, staged rollout (text-to-video first, then image-to-video, then extras) instead of a Sora-style feature dump, since a bad first impression will be hard to recover from.
What's Drew's JSON-prompt technique for consistent AI characters?
Drew reverse-engineers a reference image into a JSON prompt template in ChatGPT, capturing variables like pose, material, and finish, then swaps out only specific fields (like dog breed or color) to batch-generate a large, visually consistent set of stylized characters or icons without rewriting the full prompt each time.
Show notes
Fast Hours Podcast Episode 40—Veo 3 Makes Yesterday's Best Look Like a Beta Test +Runway & Midjourney Check-In
After a short hiatus (blame conferences and caffeine dependency), the Rory Flynn and Drew Brucker break down Google’s shiny new Flow suite — with its Veo 3 video model, sound + dialogue generation, and confusing-as-hell product naming. They talk strategy, cost, coherence, and why it still feels like Midjourney has that “magic dust” no one else can replicate.
Along the way: Runway love, layering hacks, JSON secrets, interior design with arrows, and 3D dogs with job titles. It’s fun. It’s weird. It’s chaotic. But you’ll probably walk away with 3 ideas you want to try right away.
Also, someone paid $125 just to tell you whether it's worth it. (You're welcome.)
