Product PhotographyData reportAug 16, 2026·Data as of Aug 15, 2026

AI product photography model showdown: what 12 prompts prove

Variant A was about 2.7× faster than Variant B at the same reported cost, but the supplied 12-prompt plan has no quality scores to establish a fidelity winner.

Lamina Team

Lamina Team

Product Team @ Lamina

Two ecommerce sunscreen bottle hero images displayed side by side with a product-reference sheet, brand color swatches, and packaging-label quality checks.

In the supplied Lamina test data, Variant A is the operational pick: it generated an asset in about 22 seconds, against about 58 seconds for Variant B, at the same stated $0.04 cost. The data provides no image-quality scores for product fidelity, brand consistency, packaging text, editability, or ad readiness. Treat the 12-prompt plan as a controlled acceptance test, not grounds for publishing a model ranking.

Speed buys iteration room before a campaign deadline. That matters. A fast image with a changed SPF claim, altered cap shape, or unreadable packaging still cannot serve as a hero asset. Start with the faster candidate for throughput, then require SKU-level approval against the product reference before anything reaches a PDP or paid placement.

What the supplied evidence measures—and what it does not
MetricValueSource
Variant A reported generation cost$0.040/assetuselamina.aias of 2026-08-15
Variant A reported generation time21551msuselamina.aias of 2026-08-15
Variant B reported generation cost$0.040/assetuselamina.aias of 2026-08-15
Variant B reported generation time58341msuselamina.aias of 2026-08-15
Best base-model complete product-fidelity result in Photoroom’s benchmark29.0%photoroom.comas of 2026-07-31
Top result with Photoroom’s Fidelity Layer in that benchmark38.2%photoroom.comas of 2026-07-31
BrandFusion human-evaluation preference rate for its framework66.11%openaccess.thecvf.com

What does this 12-prompt model showdown actually establish?

It establishes a real throughput gap, not a creative-quality winner. Once blind scores are recorded, the stated design can support a defensible comparison: the same original product-reference pack, 12 locked ecommerce briefs, three fixed seeds per brief, and the same native defaults for both Lamina candidate models.

The proposed subject gives the models nowhere to hide: a fictional 50 mL matte coral sunscreen bottle with a fixed wrap reading NORTHLINE, SOLAR VEIL, SPF 50, BROAD SPECTRUM, and 50 mL. The pack also requires front, 45-degree, rear, cap-off, and label-close-up source photographs, plus an SVG wordmark and approved coral, cream, and charcoal swatches. Those inputs separate retaining established facts from simply inventing a good-looking scene.

The planned 72 base generations give both candidates equal coverage. At the stated per-asset price, the base run costs $2.88 total, or $1.44 per variant. Useful trial-budget math, nothing more: it excludes local edits, rerolls, creative review, compliance review, and media spend behind the final ad.

Why start this test with Variant A?

Variant A is the practical first run because it was about 37 seconds faster per generation at the same reported $0.04 asset cost as Variant B. In plain terms, A took roughly 22 seconds and B roughly 58 seconds. The supplied measurements make A about 2.7 times faster.

That difference piles up in a live creative queue. Candidate hero images, alternate crops, and revised backgrounds clear generation sooner with Variant A, leaving calendar for the work that protects the SKU: matching approved source photography, checking visible claims with legal, and getting art-direction signoff. This is one controlled latency result, not a delivery promise; upload, review, revision, and export time sit outside it.

Do not mistake faster for more faithful. The brief assumes the candidate with the higher composite score will retain the reference product and brand, keep text readable, and survive standardized edits. No composite or component scores were supplied. That question remains open.

Can this evidence name the best model for ecommerce hero shots?

No. The supplied Lamina comparison has no scored quality outcomes, so no model earns the title of best ecommerce hero-shot model from it. What exists is a planned head-to-head with operational measurements, not completed evidence on fidelity, text accuracy, repeatability, edit success, or final ad usability.

There is one relevant external, model-level signal. AMALYTIX ranked GPT Image 2 first in its updated comparison of typical ecommerce scenarios and Amazon-catalog products, citing gains in text rendering, infographic structure, prompt accuracy, and product realism. Put GPT Image 2 in a future hero-shot bake-off, especially for brief-led images where packaging or graphic copy must stay visible.

That still leaves plenty unresolved. AMALYTIX’s evaluation is not the specified 12-prompt Lamina protocol, and it does not establish universal superiority across every category, SKU geometry, or brand system. CNET separately calls Nano Banana Pro a strong general image generator for editing, realism, character consistency, and legible text, while warning that graphic information can be inaccurate. General image quality overlaps with ecommerce product truth; the tests are not interchangeable.

How reliable is AI product fidelity in a paid ecommerce ad?

AI product fidelity still needs a review gate, particularly for packaging, logos, color, materials, and regulated copy. Photoroom’s vendor-published Product Fidelity Benchmark found that even the strongest frontier model kept complete product accuracy only 29.0% of the time across 850 products. Its Fidelity Layer raised the reported top result to 38.2%.

Never approve a hero image because the scene looks convincing. A beauty bottle can keep its coral silhouette while the label hierarchy, finish, cap geometry, net contents, or claim wording drifts. One defect is enough to make it unfit for a PDP, retail marketplace, or conversion ad.

A fidelity layer or targeted repair may improve the odds; neither figure justifies automatic publishing. Photoroom’s guidance gets the workflow right: compare output against an accurate item image, identify the altered logo, color, material, text, or component, then correct that area only. Local repair preserves approved styling better than regenerating the full composition and risking fresh changes elsewhere.

How should teams score product text, labels, and claims?

Treat product text, labels, and claims as hard-fail facts, never aesthetic details. In the sunscreen test, compare every visible instance of the fictional brand, product name, SPF value, broad-spectrum claim, and 50 mL size at full size against the approved label source. One substituted character, invented claim, or missing word fails the output for publication.

That holds even for a model with a reputation for legible text. AMALYTIX gives GPT Image 2 a positive signal on text rendering, yet legibility is not factual accuracy. A crisply rendered wrong claim is worse than an obvious typo because a rushed creative review may let it through.

The supplied research already has the right product-truth fields to lock before scoring: identity, variant, dimensions, label text and logo placement, material, color, visible claims, and certifications. Give each fixed fact a binary result. Keep lighting, composition, shadow, and crop in a separate craft score; a polished image does not erase a product-truth failure.

How should teams run the 12-prompt acceptance test?

  1. Create an immutable reference pack

    Start with original source imagery showing the item front-on, at 45 degrees, from the rear, cap-off, and in a close label crop. Include the approved vector wordmark and color swatches. Before generation, record the SKU, exact variant, dimensions, material finish, required label copy, and claims in a product-truth sheet.

    Create an immutable reference pack
  2. Lock creative variables before you compare models

    Write 12 specific hero-shot briefs around placements you actually buy: primary PDP image, premium lifestyle scene, ingredient-led composition, close crop, square social ad, vertical story placement, and other defined formats. Keep the reference pack, prompt text, crop requirement, resolution, reference-strength setting, seed set, and run count identical for every candidate.

    Lock creative variables before you compare models
  3. Run three seeds for each brief and candidate

    Use the planned fixed seeds 1101, 1102, and 1103 for every locked brief. Capture the exact model ID and version, time, cost, output resolution, and generation settings. Three seeds show whether a result repeats or was merely lucky.

    Run three seeds for each brief and candidate
  4. Score product truth before creative appeal

    Blind-score every output against the approved reference. Mark identity, variant, shape, color, logo, label text, materials, components, claims, and certifications pass or fail. Any mismatch in product identity, packaging copy, price, claim, or certification is a hard failure, regardless of visual polish.

    Score product truth before creative appeal
  5. Test one local correction on viable images

    Request one standardized localized change, such as a background swap or a correction to a failed label area. Check that approved product pixels, brand colors, composition, and nearby details hold steady. Record whether the required edit lands without creating another product-truth defect.

    Test one local correction on viable images
  6. Approve ad-ready exports only

    Check intended placement dimensions, safe crop, legibility at serving size, absence of invented claims, and alignment with approved brand styling. Report approved-image rate alongside per-generation cost and latency. That rate turns generation economics into a real publishing call.

    Approve ad-ready exports only

What should the five scoring dimensions cover?

Use the five dimensions to isolate failure modes; combine them only after removing hard failures. Product fidelity asks whether the exact SKU survives. Brand consistency covers palette, lighting, photography genre, composition, and wordmark treatment within the approved system. Text accuracy checks that every visible word and number is legible and correct. Editability tests a targeted adjustment without collateral drift. Ad readiness covers delivered crop, message safety, and publishing suitability.

Brand consistency needs its own score. A prompt alone does not reliably carry a visual system across assets. BrandFusion research identifies lighting, photography genre, palette, and composition as difficult brand traits for text-to-image models; its framework received a 66.11% preference rate in human evaluation. That supports structured brand alignment, not a universal winner among commercial generators.

Keep the scorecard order strict. Reject factual defects first. Then test whether the model holds style across seed changes and required ad formats. Only then score creative quality among surviving assets. Otherwise, a reviewer can get pulled in by atmosphere and miss a changed product fact.

What limits the reported comparison?

The supplied comparison covers two unnamed Lamina candidate variants and two operational observations. It cannot support a broad quality ranking. Exact model IDs and versions are to be captured during execution, but they are absent here, along with the blind-scoring rubric, rater count, individual prompt outputs, OCR results, edit results, confidence intervals, and approved-image rates.

The external evidence has limits as well. Photoroom’s fidelity numbers come from vendor-published benchmarking, not an independent cross-vendor result. BrandFusion evaluates a research framework rather than broadly ranking production generators. AMALYTIX offers the clearest supplied model recommendation, though its seven-scenario evaluation cannot replace a brand’s own references, claims, layouts, and placement rules.

This does not mean sending every hero asset back to a traditional shoot. AI generation suits new concepts, complex styling, on-model presentation, and campaign variation at scale. Run it with discipline: art-direct the brief, guard fixed product facts, repair isolated faults, and approve each brand-critical asset before spend starts.

What should ecommerce teams do now?

Choose Variant A for faster initial iteration, put GPT Image 2 on any broader model shortlist, and wait for the completed 12-prompt scorecard before selecting a final model. That uses the actual operating edge in the Lamina measurements without claiming latency proves product fidelity.

Build each run around approved reference assets, not descriptive prompting alone. Preserve existing product information where possible, provide an explicit brand pack to the generator, and put the product-truth sheet in front of the reviewer. For text-heavy, regulated, or certification-bearing packaging, require a dedicated full-size label review before approval. It should beat a reshoot on time because reviewers are checking a known list of factual fields.

Publish the final report with model version, prompt set, seed count, run count, product category, hard-fail rate, average creative score among passing assets, local-edit pass rate, generation time, and fully loaded cost per approved image. Then the buyer can answer the question that matters: which system reliably turns this brand’s real products into deployable creative under its own rules?

Methodology

Original Lamina experiment run 2026-08-15. Hypothesis: With identical source assets, 12 locked ecommerce briefs, three fixed seeds per brief, and blind scoring, the Lamina model/version with the highest composite score will preserve the reference product and brand most reliably while still producing legible packaging, surviving standardized edits, and yielding usable ad placements. Create original ground truth first: photograph a real 50 mL matte coral sunscreen bottle on a turntable (front, 45-degree, rear, cap-off, and close label crop), using a custom fictional wrap reading exactly “NORTHLINE / SOLAR VEIL / SPF 50 / BROAD SPECTRUM / 50 mL”. Upload these five photos plus an SVG wordmark and approved coral/cream/charcoal swatches as the same reference pack for every run. In Lamina, run every one of the 12 briefs three times at fixed seeds 1101, 1102, and 1103; preserve native model defaults otherwise, record exact model ID/version, reference-strength setting, resolution, generation time, and cost. This creates 72 scored images for two models (12 briefs × 3 seeds × 2 models), plus a reproducible original reference pack.. Measured 2 variant(s) for cost and latency on the Lamina image engine; numbers cited here are our own measurements.