Lipstick vs moisturizer AI video benchmark
Three controlled Lamina runs show a clear render-speed gap, while brand fit, product accuracy, and edit-time results still need a scored repeat test.

Lamina Team
Product Team @ Lamina

What you have is a throughput gap: Lamina returned the lipstick run in about 42 seconds, while the moisturizer and normalized runs each came back in about 61 seconds. Use it to size generation capacity. It does not pick a winning beauty category, because the supplied runs recorded no brand-fit or product-accuracy scores, first-draft pass rates, or hands-on edit time.
Run the next test as a controlled repeat: hold the creative brief, 9:16 format, duration, brand kit, reference-pack completeness, review process, and routing configuration steady, then change only the product category. That gets at the question a beauty ecommerce team actually needs answered—whether lipstick’s shade and application constraints create more rework than moisturizer’s texture and skin-finish constraints.
| Metric | Value | Source |
|---|---|---|
| Lipstick render time | 42 seconds | uselamina.aias of 2026-08-15 |
| Face moisturizer render time | 61 seconds | uselamina.aias of 2026-08-15 |
| Controlled cross-category normalization render time | 61 seconds | uselamina.aias of 2026-08-15 |
| Generation cost per measured asset across all three variants | $0.040 | uselamina.aias of 2026-08-15 |
| Median time to generate an asset | 230s | Lamina platform telemetryas of 2026-08-15 |
What does the lipstick-versus-moisturizer test actually prove?
It shows that the measured lipstick asset rendered materially faster than the other two controlled variants, while all three runs shared a per-generation cost of four cents. The moisturizer needed roughly 19 seconds more than lipstick; normalized landed less than a second from moisturizer. That matters in a large batch, though human review, revisions, approvals, and paid-media costs were outside the render measurement.
The broader Lamina telemetry figure is a guardrail, not a conflict. Its 230-second median spans assets across platform activity; these three beauty measurements are individual experimental runs under a specified setup. Read the 42- to 61-second range as observed return time from one test, never as a production promise across every model route, reference pack, or asset type.
No quality result came with this experiment. There is no observed evidence that moisturizer had better brand fit, more accurate packaging retention, a higher review-pass rate, or fewer edit minutes than lipstick. Those are sensible hypotheses, especially with cosmetic packaging, color, labels, and application anatomy leaving so little room for drift. They are not benchmark findings.
What one-line brief should both beauty categories get?
Give both categories the same production brief: “Create a 15-second 9:16 premium beauty social ad: hook in the first second, show product proof and one application moment, end on a clean packshot and CTA; preserve the supplied product exactly.” Keep it fixed. Otherwise a richer moisturizer instruction or an easier lipstick script has decided the outcome before generation starts.
Lamina’s published brief guidance calls for the channel, aspect ratio, first-second hook, product proof, and a brand-kit non-negotiable. Put all of that in the common instruction, then provide an equally complete category reference pack: front, side, and back package views; a close product-detail crop; approved brand colors, fonts, and copy; plus a hand-for-scale reference. Products change. The evidence standard stays put.
Before the full ad, run a three-distance preflight. Close checks label and material preservation; mid checks usable product scale; wide checks whether the pack still belongs in the visual world. Guidance on locked product references calls for checking silhouette, color, material, logo, and plausible scale at close, mid, and wide distances before scaling video. A bad packshot reference contaminates every scene after it.
What storyboard should the agent return before it renders?
Before generating video, require a five-scene storyboard: a 0–1.5-second hook and reveal, 1.5–4 seconds of macro product proof, 4–8 seconds of application or use, 8–12 seconds of finish or payoff, then a 12–15-second locked packshot with approved CTA. Now the first draft is auditable. Reviewers no longer have to guess which instruction led to a bad label, a missing application beat, or a product reveal that arrived late.
Keep the storyboard verbatim with every scene prompt, selected model, seed or settings, generation count, render duration, chosen take, and rejection reason. Lamina describes a workflow that takes a brief and brand, coordinates image, video, and try-on models, scores against a brand kit, and delivers selected assets. Log that orchestration in any benchmark: Lamina routes across more than 15 models, and a routing or model-version change can shift a later batch.
The storyboard is a request and audit record, not an observed Lamina output from the supplied runs. That matters. Its scene list lets the team compare like with like even if the video agent shifts framing, camera movement, or generation approach from one product to another.
How to run the controlled beauty-video benchmark
Lock the production conditions
Use the same 15-second vertical brief, brand kit, CTA, product-proof requirement, model-routing configuration, and review rubric in both cells. Change only the lipstick or moisturizer product and its equally complete reference assets.

Check the product at three distances
Generate or inspect close, mid, and wide product views before building the full sequence. Reject the reference setup if pack geometry, label, cap, color, material, or scale wobbles. A multi-scene ad only multiplies that error.

Request and save the scene plan
Get the five-scene storyboard before render. Store its scene prompts and production settings with the final output, so a second reviewer can reproduce the job instead of rebuilding it from memory.

Generate matched first-draft batches
Run at least 10 first drafts per category, using the same reference-pack standard. Randomize the order reviewers receive them. The category name should not tip the score before anyone sees the asset.

Score blind; log every intervention
Have two independent reviewers score brand fit, product accuracy, and edit efficiency without seeing the category label where practical. If scores differ by more than one point, send it to a third reviewer for adjudication. The editor then logs each correction by type and minutes.

Report category results apart from render speed
Publish the mean, median, standard deviation, hard-fail rate, usable-first-draft rate, median hands-on edit time, retry count, and edit-category counts. Keep render latency and nominal generation cost apart from the cost of an approved asset.

How should lipstick and moisturizer videos differ creatively?
For lipstick, put the detail into shade proof, the bullet or applicator, the lip-contact moment, and the final packshot. Establish the product quickly, then give the swatch and finish room to read. A non-Lamina lip-tint example uses macro swatch, model application, beauty imagery, and a final hero product shot—a useful pattern to test, not evidence of Lamina performance.
Face moisturizer needs inspection time for cream texture, fingertip pickup, spread, absorption, and the visible finish around the product story. Let transitions run slower here; viewers need to read texture and skin-state cues, not just register a shade. Runway’s beauty guidance likewise names slow reveals, cream macros, campaign-grade beauty shots, and product references that retain packaging detail as common building blocks.
The application scene is where both cells get stressed, for different reasons. Lipstick needs credible lip anatomy, contact, shade placement, and stable final color. Moisturizer needs a believable product amount, hand-to-skin movement, texture behavior, and a finish that does not imply an unsupported product claim. Beauty UGC guidance recommends testing shopper-facing hooks at scale with AI while treating exact swatches and true application as claims needing especially careful validation. Keep the generated visual; scrutinize its product proof.
Which edits should reviewers log?
Use fixed edit categories: packaging correction; product-scale or hand correction; application anatomy or physics correction; skin or texture realism correction; brand-style correction; claim or legal correction; pacing or shot-order correction; text or CTA overlay; and discard or regenerate. That log shows whether lipstick burns time on color and application fixes while moisturizer burns it on cream and skin-finish fixes, instead of collapsing all rework into an editor’s vague impression.
Add exact legal copy, shade names, and CTA text in post-production overlays. A repeatable AI-video evaluation framework flags label and logo changes as common failure modes and recommends adding exact typography during editing. This is not surface polish. A wrong product name or altered package copy can hard-fail an otherwise attractive draft.
Give published ads a separate platform-rules review. Meta says advertising must clearly represent the company, product, or brand; stay relevant to the advertised offering; and match the landing page. Its cosmetics guidance requires an adult audience and bars prohibited claims. Brand fit does not rescue a misleading product representation.
What scoring rubric makes this comparison repeatable?
Use a weighted 0–5 rubric: product accuracy at 45%, brand fit at 35%, and edit efficiency at 20%, with a non-negotiable product-accuracy hard gate. That weighting matches beauty creative in practice. A polished video showing the wrong lipstick shade, cap, label, package shape, or application result cannot run as an ad.
Judge brand fit against the approved palette, font treatment, tone, composition, luxury level, and category-appropriate pace. A 5 is publishable with no aesthetic correction; a 3 is clearly on brand but needs one targeted pass; a 1 looks generic or visibly off brand. Lamina’s score-gated approval use case supports formal brand review rather than automatic publication after generation.
Score product accuracy across pack geometry, cap, logo and label, shade or color, material or cream texture, scale, and credible application. Hard-fail any materially wrong packaging, shade, product text, or application moment. Subject fidelity deserves the heaviest weight: the published evaluation framework also prioritizes it, while separately testing motion obedience, temporal stability, camera control, and editability.
Score edit efficiency by hands-on minutes and retries: 5 for zero to five minutes with no regeneration, 4 for six to 15 minutes, 3 for 16 to 30 minutes, 2 for 31 to 60 minutes, and 1 for more than an hour or a discarded draft. Report usable-first-draft rate separately—the drafts passing the accuracy hard gate divided by all drafts. Otherwise an appealing average can hide a stack of unusable output.
We’re an indie brand at heart. We always have more ideas than time to execute. With Adobe Firefly, I can visualize the mood and world of a campaign in minutes. Once that vision is clear, editors and designers know exactly how to build it out.
What is the practical decision rule for beauty teams?
Pick a category only after repeated batches show a persistent gap in accuracy-gated first-draft rate and median hands-on edit time. One faster render proves little. Run 10 or more matched drafts per cell, compare means and medians, include standard deviation, and repeat the batch if one category appears ahead. One strong video is a creative sample, not a production benchmark.
For now, budget this test’s generation layer at four cents per asset: approximately 42 seconds for lipstick, versus 61 seconds for the moisturizer and normalized runs. That excludes review, regeneration, compositing, legal clearance, and publishing work. It is not cost per approved ad; the edit log and accuracy gate turn raw generations into an operating metric you can decide from.
The working hypothesis still deserves a test: moisturizer may allow more visual latitude through texture and hydration imagery, while lipstick demands exact shade fidelity and highly scrutinized application. The available data do not validate it. Hold the system steady, retain the evidence trail, keep human art direction and approval as the final gate, and let the scored batch—not a category stereotype—show where Lamina’s agent needs the most prompt specificity.
Methodology
Original Lamina experiment run 2026-08-15. Hypothesis: Given matched one-line briefs, Lamina’s AI product-video agent will produce stronger first-draft brand fit and require less edit time for face moisturizer than lipstick, because moisturizer can rely on broad skin-texture and hydration imagery, while lipstick demands tighter shade fidelity, lip anatomy, and application accuracy. A controlled, repeated benchmark will quantify the gap and identify the category-specific prompts and edit interventions that close it.. Measured 3 variant(s) for cost and latency on the Lamina image engine; numbers cited here are our own measurements.