Protein Puffs

AI Video Production for RIVET Protein Puffs

A CPG Commercial Without a Film Crew

Scroll Down

Services

Marketing Goal

Brand awareness commercial

Target Audience

Protein snack consumers

Project Timeline

5 weeks

AI Video Production for RIVET Protein Puffs

Project Overview

The Client: RIVET — a US protein snack brand. Their flagship product, Protein Puffs, is baked (not fried), delivers 20g of protein and 3g net carbs per serving, in flavors like Loaded Sour Cream & Onion.

The Task: Produce a brand awareness commercial without a traditional film shoot. Two deliverables: a 10-second hero cut and a 5-second compressed version, both upscaled and resized for large-screen and CTV placements.

The Concept: An active lifestyle without compromise. RIVET Protein Puffs fits any moment of the day — outdoors, waiting on the grill, in good company. "Baked, not fried" isn't just a product spec here; it's the positioning the entire commercial is built around.

Project Execution

The Script Logic: Show the Craving, Not the Chewing

The first decision that shaped the entire script was technical, not creative.

Showing a close-up of someone eating a snack in an AI-generated video means losing quality exactly where it matters most. Neural networks still struggle badly with mouth movement and chewing animation — facial expressions glitch, the product's texture breaks apart mid-bite, and fingers holding food lose their shape on contact. The result isn't an appetizing snack; it's a visible AI artifact working against the brand.

The fix: cut the eating moment entirely. The focus shifts to the moment before — the product in the hero's hand as part of the mood, not a how-to-eat instruction. This isn't a workaround. It's a stronger directorial choice: the viewer fills in the moment themselves.

Ten-second commercial structure:

  • Scene 1 (2 sec): Wide establishing shot — the outdoor setting
  • Scene 2 (3 sec): Close-up of the hero holding a bag of RIVET
  • Scene 3 (2 sec): Medium shot — hero within the scene
  • Scene 4 (3 sec): Packshot — product, logo, tagline

The Core Pipeline Decision: AI for the Scene, 3D for the Product

This wasn't an aesthetic choice — it was a technical necessity.

AI models generate convincing environments and characters. But ask one to accurately reproduce a specific package with a real logo, exact typography, and precise brand colors, and you get artifacts, distortion, and text hallucination. For a brand with a codified visual identity like RIVET's, that's not an option.

Our solution: AI owns everything that builds atmosphere — characters, locations, lighting, motion. 3D owns the product. The RIVET Protein Puffs bag is composited into finished AI-generated scenes using 3D graphics, with no neural network involved in rendering the package itself. That's what guarantees 100% brand accuracy in every frame.

For a deeper look at how this hybrid pipeline works in practice, see our AI Video Production & 3D Motion Graphics case study.

Stage 1: Moodboarding and Casting

Before any generation began, we locked the visual direction using traditional moodboards built from reference photography, color palettes, and location examples — zero AI involved at this stage.

The moodboard's job is alignment before a single prompt gets written. If the client and the creative team have two different pictures in mind for "active outdoor lifestyle," that surfaces in an hour on a moodboard — not a week before deadline in a final edit review.

Once the visual direction was locked, we moved to character casting. We prepared several AI-generated avatar options targeted at RIVET's core audience — looks that read as authentically "their" customer, not generic stock talent.

Stage 2: Generation, Animation, and the Real Problems of an AI Pipeline

Once static frames were approved for every scene, we moved into animation. AI video isn't a "generate" button. It's an iterative process where every intermediate result needs review, judgment calls, and — frequently — a full regeneration from scratch.

The commercial features three characters: a father, a mother, and their son. A family scene outdoors, waiting on the grill. Early in animation, it became clear that three characters in one frame doesn't add complexity linearly — it triples every typical AI failure mode at once.

Directorial Call: Where the Characters Look

By the third iteration of the scene, we tested two versions of the final beat:

  • Version 1: the family looking at each other
  • Version 2: their gaze drifting away from each other — toward the grill

Version 2 won. The look toward the grill lands harder: it closes the loop on the anticipation the commercial builds from frame one. The viewer gets the story without a word of dialogue — the meal is close, and RIVET is what fills the wait.

Hands and Fingers

Neural networks handle hands poorly — a well-known failure point in every AI production. Here the challenge compounded: three characters, and every set of hands needed to move naturally. Glitches on hands showed up in the majority of first-pass generations.

A specific case: the son's hands in the first frame. Regeneration wasn't producing a stable result — one take would nail the hand position but glitch elsewhere, the next would fix that and break somewhere new. The fix: mask the one clean phase of hand motion and freeze it in place, while letting the body's animation continue underneath. Technically more demanding than simply regenerating again, but it delivers a predictable result exactly where the model keeps failing.

The Father's Eye and a Floating Object

The first frame carried two unrelated artifacts that had to be fixed separately.

First: the father's eye was animated incorrectly — a micro-shift subtle enough to miss on a casual watch, but registered by the brain as "something's off with this face." Classic uncanny valley, at the level of micro-expression.

Second: a yellow object on the table near the mother was visibly "floating" in the animation — a clear generation artifact with no relation to the scene. Same fix pattern: freeze the object on a neutral, static phase and remove it from the motion entirely.

The Grill: Three Unrelated Problems in One Object

The background grill turned out to be the source of three separate issues.

First — spatial inconsistency across shots. AI models don't reliably preserve object position relative to other elements from scene to scene. In one frame, the grill sat at one angle to the porch; in the next, it had rotated roughly ninety degrees. In the edit, the grill visibly "jumped" — the viewer is looking at the same backyard, but the grill keeps relocating within it. The sense of a single, coherent space falls apart.

Second — the AI added an extra grill. In several generations, a second, unrequested grill appeared in the background, smoking away on its own. The model has no concept that the scene's internal logic only allows for one. Both artifacts had to be manually removed.

Third — texture degradation. Around the tenth generation pass, the grill's metal surface started to "melt": the structure lost readability entirely. Fixed manually — a corrected texture layer composited over the affected frame.

The Slipping Mask and Package Color Mismatch

In the third frame, the mask on the product slipped — a compositing error where the 3D product layer loses its tracking lock on the character's hand motion. The RIVET bag starts to visually exist independently of the person holding it.

A separate issue turned up comparing frames three and four: the package color didn't match between the two scenes. Each AI-generated scene carries its own lighting profile, and the 3D render of the package was built independently for each one. The fix went back to the client's brand book — a recalculation and re-render of both packages against a single, locked colorimetric reference.

Stage 3: Sound Design

A voiceover script timed to 10 and 5 seconds isn't just a read — it's a compressed copywriting brief. Tone, pacing, and delivery all need to land in sync with the picture, with zero room to spare.

Beyond the voiceover and music bed, full sound design was layered in: ambient outdoor atmosphere, subtle SFX. Without this layer, even a well-edited AI commercial feels flat. Sound carries half the experience.

The final audio pass balanced voiceover, music, and SFX, with loudness leveled across every placement format.

Results

Final delivery in 5 weeks:

  • A 10-second hero commercial — two edit versions
  • A 5-second compressed cutdown
  • Upscaling and resizing for large-screen and every required placement

Produced with no film crew, no location rental, no live-action casting — while maintaining full creative control over the visual aesthetic and pixel-accurate brand representation in every frame.

Need an AI Commercial for Your Brand?

Lava Media is an AI commercial production company built around exactly this kind of hybrid pipeline — AI-driven commercial production paired with human-led 3D compositing, so the final frame is both fast to produce and pixel-accurate to your brand. If you need AI video advertising with precise product integration — let's talk.

Related work: