The Skill Was Named After a Model It Didn't Use: An Amazon SBV Video Build Log
← Back to the journal

The Skill Was Named After a Model It Didn't Use: An Amazon SBV Video Build Log

John Aspinall · · 13 min read

The skill lives at ~/.codex/skills/amazon-sbv-veo-video. The name has "Veo" in it. The SKILL.md is 2,089 words, dated May 22, and it is built around Google Veo: three 5-second clips per concept, each continuing from the last clean frame of the one before, editor-added text overlays, a 15-second default and an AAC stereo audio track.

The only real run of it on my machine used none of that.

On August 26 I ran it on B0FXY55MY9, Old Spice Aqua Reef deodorant. The run used a different model. It produced seven independent clips, not a chain of three. It cut them into a 10-second video with ten hard cuts, no overlays and no sound. And it wrote four of the eight files the skill says every run has to produce.

The output is good. The file describing how it was made is wrong. And the QA checklist passed a frame with garbled label text in it. This post covers all three, because this was the first build where the model changed underneath the skill and the skill never found out.

One disclosure first. Old Spice is not a client. The run sits in my amazon-research folder, not a client folder. I picked a big-brand ASIN with a busy label on purpose, because a label is the hardest thing to keep intact in a generated video, and I wanted to break the skill on packaging I don't have to defend. So there's no campaign data in this post. No CTR, no ACOS. What I have is the file trail.

The problem: SBV is the ad format brands skip

Sponsored Brands Video is the format we see missing from Q4 plans more than any other on accounts doing $100K-$500K a month. The reason is never strategy. It's production. Somebody has to shoot or animate the product, cut it to Amazon's specs and get it through moderation. Quotes we've been sent for a short product video have usually been four figures with a two-to-three-week turnaround. So the video gets pushed to "after peak", and after peak it gets pushed again.

The format itself is unforgiving in ways that suit a machine better than you'd expect:

  • Autoplay is muted. The ad has to make sense with no sound.
  • The first second decides it. The video plays in a results page next to everything else on it, so the product has to fill the frame straight away.
  • The specs are fixed and checkable. 16:9 only. 1280x720, 1920x1080 or 3840x2160. 6-45 seconds, with 20 or less recommended. 23.976-30 fps. H.264 or H.265, Main or Baseline profile. Progressive scan. No black frames at the start or end. No letterbox or pillarbox bars.
  • The content rules are a list, not a vibe. No star ratings, no prices, no deal language, no flashing or spinning, no Amazon UI, no unsupported claims.

A format with fixed specs and a written list of prohibitions is exactly the kind of thing a skill file handles well. That was the bet in May.

The build: what the skill actually says

The May file has three ideas worth copying even though the run ignored two of them.

1. A product fidelity gate before any prompt gets written. Verbatim:

Do not write prompts that invent exact label text, badges, certifications,
ingredients, specs, warnings, or package claims.
If the product must appear clearly and the references are not strong enough,
ask for better images or an approved render.

2. Text belongs in the edit, not in the generation. The skill says to keep exact text overlays, logos and CTA copy "as an editor overlay plan whenever possible, because video generators are unreliable with precise small text." Hold on to that sentence. It's the most accurate line in the file, and the run proved it in a way I didn't catch at the time.

3. Continuation chaining. Every clip after the first has to start from the previous clip's last frame:

Clip 02 - Continuation/detail ([target seconds])
Use the final clean frame from Clip 01 as the reference image/frame.
Continue from that exact product position, setting, lighting, and camera style...
Do not reset the scene, introduce a new product, change the variant,
or animate the product as a character.

This was the Veo design. Continuous camera, continuous scene, one story told across 15 seconds. It's also the part that tied the skill to one model's reference-frame behaviour, and that's where things went wrong.

The run: what actually happened on August 26

The folder timestamps tell the story:

Time What landed
18:51:28 11 source assets: official pack shot, closed and cap-off identity shots, three phone photos of a physical stick (front, cap off, back), five reference comps
18:53:28-38 7 keyframe stills, 16:9, generated with Grok image_edit from the pack references
18:55-18:57 7 motion clips, Grok image_to_video, 1280x720, 24 fps, 6.04 seconds each
18:59:25 Master export plus QA frames at 1s, 5s and 9s
19:00:02 Edit notes and QA checklist

That's about nine minutes from source assets to documented master. The session started earlier, because the browsing and reference-gathering came before 18:51. I don't keep session logs, so I won't pretend I know how long the whole thing took.

The shot list is where the design changed. Instead of three chained clips, the run generated seven separate clips, each from its own still:

# Clip Motion Fidelity lock
01 hero Aggressive push-in Closed cap + 24/7 sticker, exact label
02 label Horizontal whip across the mark Glyphs sharp at mid-whip
03 open Cap flies off, no hands Cap is the short clear shell; gel is cyan
04 truth Micro-push on the gel dome Gloss, grain, dimple, not white solid
05 spill Water burst + orange peel tumble Product stays the Aqua Reef stick
06 overhead Slow spin Scatter readable, pack identifiable
07 settle Slow orbit into still end frame Loop-friendly, label readable

Then it cut them to a 10-cut timing table. No cut is longer than 1.75 seconds, and most are under one:

Cut Time Source In-point
1 0.00-0.85 hero 0.00
2 0.85-1.65 label 0.40
3 1.65-2.55 open 1.00
4 2.55-3.45 truth 0.50
5 3.45-4.75 spill 0.80
6 4.75-5.55 overhead 0.50
7 5.55-6.35 truth 2.50
8 6.35-7.55 spill 2.50
9 7.55-8.25 label 0.50
10 8.25-10.00 settle 3.50

Master: 1920x1080, 30 fps, H.264 Main, about 8 Mbps, silent, exactly 10.00 seconds and 300 frames, full-bleed. It passes every technical spec on Amazon's list.

Cost. Seven clips at 6.04 seconds is 42.3 seconds of generated video. At xAI's published 720p rate for Grok Imagine video, $0.07 a second, or $0.08 on the 1.5 model, that's roughly $3.00-3.40 in motion. Add seven image edits and input charges and the whole run is under $5 in generation. That's from the published price list, not an invoice line. It's still the cheapest thing in this post by two orders of magnitude.

The number that matters more: 10 of 42 generated seconds shipped, about 24%. The rest wasn't bad luck. It was planned for, and the next section explains why.

What broke, part one: the label has a clock on it

Open clips/review/01-hero-t1.jpg, a frame from one second into the hero clip. "Old Spice" and "AQUA REEF" are clean. The small sticker on the cap is not. "LEGENDARY FRESHNESS" renders as "LEGENBARY". "ODOR FIGHTERS" renders as "COOR FIGHTERS".

Two seconds later, at 01-hero-t3.jpg, the splash has moved up the stick and "AQUA REEF" has become "HQUA REEF". The edit notes also record that the push-in "slightly warps the stick feet into the splash after ~1s."

That's the real reason for the 10-cut structure. I'd love to call it a creative choice, but it wasn't really. Cuts average about a second because that's roughly how long a generated label stays true. Cut 1 uses 0.00-0.85 of the hero clip and throws the rest away. Cut 9's in-point was moved from 2.00 to 0.50 "so the flashback stays on the sharp AQUA REEF mark." The fast-cut style is what you get when you edit around decay.

This is the practical version of the line in the May file about generators being unreliable with small text. The file was right. It just assumed the fix was keeping text out of the generation. On packaging you can't do that, because the label is the product. The generator has to render the label, and the label starts decaying as soon as the camera moves.

What broke, part two: QA passed a frame with gibberish on it

The QA checklist in the run folder has this line ticked:

- [x] Packaging matches lock (red Red Collection stick, gold hex, citrus-wave band)

Now open exports/qa-frame-09s.jpg, the end frame. The one the loop lands on and the one that sits on screen longest. "Old Spice" is clean. "AQUA REEF" is clean. "ALUMINUM FREE" is clean. The cap sticker is unreadable: the two lines that should say "LEGENDARY FRESHNESS" and "with ODOR FIGHTERS" are a jumble of letter-shaped marks. "SCENT INTENSITY" is garbled. And the net weight line, the one piece of regulated copy on the front of the pack, is garbled too.

The check was honest about what it checked. The lock listed the colour, the gold hex, the citrus band and the gel colour, and every one of those is right. Nobody put the small print in the lock, so the QA confirmed everything except the part most likely to be wrong. It sampled 3 frames out of 300, or 1%, and the frame it passed was the one with the problem.

Will a shopper scrolling a muted results page on a phone read a net weight line on a 10-second loop? Almost certainly not. But that's the wrong test. The right one is whether you'd put a frame with invented text on the front of a product into an ad under your brand name. On Old Spice it's a curiosity. On a supplement, where the small print is a serving size or a "no artificial" claim, it's a moderation rejection at best and a claim you never made at worst.

What broke, part three: the file describes a pipeline that doesn't exist

This is the part I'd have missed if I hadn't sat down to write this post.

The skill says every run produces eight files. The run produced four: shot-list.md, edit-assembly-notes.md, qa-checklist.md and production-specs.md. The four missing ones are creative-brief.md, sbv-video-strategy.md, editor-overlay-plan.md and veo-generation-prompts.md.

That last one is the problem. The prompts that made these seven clips aren't in the folder. I can show you the timing table, the trim points, the codec and the bitrate. I can't show you the sentence that produced the hero clip, because it only existed in a session that's gone. The run can be verified, but it can't be repeated.

Then there's the name. Anyone who opens amazon-sbv-veo-video in six weeks, including me, will reasonably assume the pipeline is Veo-based, chained, 15 seconds long, with overlays and audio. Every one of those assumptions is wrong. The only place the real generator is written down is production-specs.md, which to its credit records Grok image_to_video · resolution_name: 720p · 6s.

This is the same class of failure as the folder split in last week's orchestrator log and the screenshot frame in the rollout log. The model did something reasonable, the file allowed it, and nothing wrote the difference down. The difference this time is that it wasn't a naming rule that drifted. The whole method drifted, and the documentation stayed behind with the old model.

The accidental good news: the drift made it portable

In the week of September 14 I wrote about two deadlines landing six days apart: the Sora API shutting off on September 24 and the Gemini Omni Flash preview being deprecated on September 30. The point was that a chained, model-specific video pipeline is a vendor dependency you don't see until the end-of-life email arrives.

The May design, with continuation clips each built on the last frame of the one before, is exactly that kind of dependency. It relies on one model's reference-frame behaviour, and swapping the model means rewriting every continuation prompt and re-checking every seam.

The August run, which I didn't design deliberately, has no seams to re-check. Every clip is an independent image-to-video generation from a still, joined with hard cuts. Swap the generator and you regenerate seven 6-second clips from the same seven keyframes, then reuse the same trim table. The keyframes are the asset. The motion model is replaceable.

I'd like to say I planned that. I didn't. It came out of cutting around label decay. But it's now the rule I'm writing into the file.

The rewrite, this week

Five changes, all to the file, none to the workflow that produced the video:

  1. Rename it amazon-sbv-video. The generator becomes a required field in the run manifest, not a word in the folder name.
  2. Make fast-cut from independent keyframes the default structure. Continuation chaining becomes an option, flagged as model-specific.
  3. Prompts are a required artifact, written before generation. A run folder without the prompts counts as incomplete, however good the export looks.
  4. Add small print to the fidelity lock, with a binary rule: every piece of on-pack text is either legible or deliberately out of focus. Invented glyphs fail. No third option.
  5. Sample QA frames every half-second, not three times. That's 20 frames on a 10-second master, which takes about 90 seconds to look through, and is still cheaper than one moderation rejection.

What an operator could copy

  • Treat the label as having a shelf life measured in seconds. Plan cuts around how long text holds up, not around how long the model can run. Shorter cuts aren't a style choice here. They're how you keep the product honest.
  • Keep keyframes and motion as separate steps. Make the still, approve the still, then animate it. The approved still is the asset that survives a model shutdown.
  • Write the lock before you look at the output, and put the small print in it. A QA that checks the colour and the logo will pass a frame with a made-up net weight on it.
  • Save the prompts in the run folder or accept that you can't repeat the run. A beautiful master with no prompt trail is a one-off, not a pipeline.
  • Record the generator in the artifact, not in the name. Models change more often than folder names.
  • Put the video on SBV before the detail page. A muted, skimmed ad in a results grid is the most forgiving place to learn whether generated product video works for your catalogue. It gets a CTR read in two to three weeks. The PDP is where a garbled label costs you conversions on your warmest traffic.

The workflow was the cheap part again. Nine minutes of timestamps, under $5 in generation, a spec-perfect 10-second master. The expensive part was the file, and this time it wasn't a missing constraint sentence. It was a file describing a pipeline I'd already stopped running.

The most useful thing in the run folder wasn't the video. It was the frame from one second into the hero clip, the one where "ODOR FIGHTERS" became "COOR FIGHTERS." That frame explains every edit decision in the timing table. The QA should have looked at it first.

Related reading: the listing pipeline orchestrator build log, the image stack rollout build log, and the Sora API shutdown and your video pipeline.

Enlarged image preview