Product PhotographyAug 5, 2026·Data as of Aug 2, 2026

Data report: AI flat-lay-to-model photography vs. studio shoots for ecommerce — a repeatable benchmark using Lamina. Test the same apparel SKUs across catalog accuracy (silhouette, color, print, trims), realistic drape/fit, pose control, brand consistency, production time, and cost per approved asset. Publish failure examples and a decision matrix for when AI on-model imagery is fit for PDPs, paid social, and campaign concepts versus when a physical studio shoot remains necessary.

A reproducible protocol for comparing AI flat-lay-to-model images with studio references—without claiming a Lamina result before the SKU-matched test is run.

Lamina Team

Lamina Team

Product Team @ Lamina

Apparel flat lay beside AI-generated on-model variants, reviewed against a garment-accuracy scorecard on a merchandising desk

What does this AI flat-lay-to-model benchmark actually establish?

This protocol does not establish that Lamina matches or outperforms a studio shoot: the supplied evidence contains no completed Lamina-versus-studio apparel test. It gives an ecommerce team a repeatable way to find out—run the same apparel SKUs through AI on-model and studio-reference production, blind-score each against the physical garment, then publish the approval-adjusted result instead of a flattering gallery of selects.

Photorealism is not product truth. Garment-consistency research found that broad quantitative image measures miss fine-detail consistency, and the supplied virtual-try-on evidence points toward interpretable review of texture fidelity, shape plausibility, and realism rather than a generic similarity score. A convincing image is a candidate. It is not yet an approved catalog asset.

Run the comparison at SKU level. Keep the flat lay, front and back references, detail crops, size and color information, plus the generated prompt and settings for every garment; then make matched studio-reference and AI-on-model deliverables in the same intended aspect ratios. That trail lets you explain a rejection instead of calling it subjective.

Benchmark facts that determine the scoring design
MetricValueSource
Catalog-accuracy scoring range per attribute0–4link.springer.comas of 2026-06-11
Garment attributes operationalized by the consistency framework20link.springer.comas of 2026-06-11
Technical and human QA stages2dev.toas of 2026-08-02
Drape-and-fit checks specified for review8loomadesign.aias of 2026-04-25
Median time to generate an asset207sLamina platform telemetryas of 2026-08-04

How do you build a fair, SKU-matched apparel test?

Hold the garment, channel brief, output specification, and reviewer information constant. Change only the production route. Build a deliberately difficult, representative SKU set: simple tees; printed or logo garments; dark and light colorways; textured fabrics; metal hardware; sheer materials; layered pieces; and styles where fit is purchase-critical. Do not let a clean basic tee settle the program for an embellished jacket.

Build one evidence packet per SKU before generation begins. Include every available flat-lay view and close-up of buttons, zips, seams, labels, print, neckline, sleeve opening, hem, pockets, lining, and any included component. Flat lays do not show body tension, pocket depth, sleeve behavior, lining, or fabric movement directly, so reviewers need a clean record of what the source proves and what the model may be inferring.

Match the briefs for pose, crop, background, model direction, and channel. Save every prompt, reference image, generation setting, retry, edit, and timestamp. A benchmark built from hand-picked winners is just a creative review; one with a fixed prompt set and complete run log can be repeated next season.

How to run the benchmark

  1. Set the SKU panel and evidence standard

    Predeclare a SKU panel with both straightforward and failure-prone construction. Photograph or document the actual garment as the adjudication reference, then list each SKU’s required product-truth attributes: silhouette, color, print or logo, material cues, trims, neckline, sleeves, hem, and included components.

    Set the SKU panel and evidence standard
  2. Generate fixed AI on-model variants

    Use Lamina for the same planned poses and crops across every SKU. Keep the brief fixed within each test cell, log all inputs and settings, and set the retry cap before work starts. Preserve the entire output set, rejects included. You cannot calculate usable-image rate from approved examples alone.

    Generate fixed AI on-model variants
  3. Score catalog accuracy before visual appeal

    Give blinded reviewers the evidence packet and have them score each mandatory attribute from 0 to 4. Score silhouette and proportions; color; print, logo, and text; texture or material; seams, trims, buttons, zips, and hardware; neckline, sleeves, and hem; and included components on their own. Log the exact failure—missing button, altered stripe direction, whatever it is.

    Score catalog accuracy before visual appeal
  4. Treat drape and fit as a separate gate

    Assess garment length, shoulder and waist placement, ease or tension, sleeve and inseam behavior, fold plausibility, transparency, pattern alignment, and concealed areas. Flag every claim the flat lay cannot support. That keeps a realistic pose from passing as evidence of real-world fit.

    Treat drape and fit as a separate gate
  5. Run technical QA, then blind human review

    Automate resolution, aspect-ratio, background-compliance, missing-asset, and basic artifact checks across every output. Then let a merchandiser or product specialist judge product truth and a creative or brand reviewer judge commercial suitability. Keep the production route hidden until they submit their scores.

    Run technical QA, then blind human review
  6. Publish approval-adjusted time and cost

    Report elapsed production time from a complete evidence packet to approved delivery, total labor and tool spend, total outputs, approval count, and cost per approved asset. Lamina reports a median generation time of 207s per asset. That is not publishing time: human review, revisions, source preparation, and any media spend sit outside that figure.

    Publish approval-adjusted time and cost

What should a catalog-accuracy rubric measure?

Score each observable garment attribute separately, then fail any output that misses a mandatory attribute. Use 0–4: 0 is absent or materially wrong; 1 is clearly wrong; 2 is partly preserved with a visible discrepancy; 3 is substantially accurate with a minor discrepancy; 4 is verified against the evidence packet. This follows the supplied garment-consistency framework’s attribute-level approach.

Set mandatory fields by SKU. A blank tee may need silhouette, color, neckline, sleeve, and hem; a branded track jacket also needs logo geometry, stripe placement, zip type, pocket treatment, and contrast-panel boundaries. An average can hide a deal-breaking logo substitution. Publish the average and the mandatory-attribute pass rate.

Do not turn the rubric into one automatic verdict. The supplied evidence puts automated checks on technical compliance, while human reviewers assess product fidelity, material, scale, and brand suitability. That is the useful split: software can catch the wrong canvas size; it cannot reliably judge whether a pleat, cuff, or translucent layer makes a merchandising promise the garment cannot support.

Which failures belong in the report?

Publish representative failures by error class, alongside the source evidence and the reviewer’s rejection reason. Prioritize changed print or logo, omitted hardware, altered trim, incorrect neckline or hem, implausible transparency, misplaced pattern alignment, unsupported pocket or lining behavior, and poses that hide a decision-critical feature. Buyers can then see where the process needs stronger inputs, a tighter prompt, or a different asset treatment.

Do not bury retries. For every published pass, retain the number of variants generated, the first-pass result, every correction pass, and the final decision. A beautifully rendered image that took extensive human intervention belongs in approval-adjusted cost and time—not in a headline generation-speed claim.

Failure reporting keeps the benchmark useful. It shows a merchant whether a category is ready for repeatable on-model production, and it shows the creative team which evidence gaps, rather than vague AI worries, are driving avoidable rejects.

When does AI on-model imagery belong on PDPs, paid social, and campaign concepts?

Put AI on-model imagery on a PDP only after every mandatory product-truth attribute passes, no fit or drape claim exceeds the supplied evidence, the destination channel permits the asset, and blinded human reviewers approve it. Catalog imagery sits close to the purchase decision, so the supplied operational guidance demands the strictest preservation here. Approval is SKU- and output-specific. It does not clear every garment in the range.

Use AI more readily for paid-social variants where the product remains materially accurate and the asset is supplementary or short-lived. The same product-truth gate applies, though the format can carry more pose, crop, background, and audience variation than a primary PDP image. Keep the review record. Social speed does not excuse an altered logo, print, or garment construction.

Use AI most freely for campaign concepts, moodboards, and pre-production exploration, and label concepts as concepts rather than product proof. For complex, reflective, sheer, highly patterned, fit-critical, construction-critical, high-value, regulated, or marketplace-sensitive SKUs, raise the evidence threshold and send outputs through closer product-specialist review. AI generation remains the production premise. The control is better evidence and stricter approval, never lower standards.

PDP primary image: all mandatory product-truth attributes pass; drape and fit claims are supported; and human review approves. Use AI on-model output only after full SKU-level approval.

Paid social variant: the product remains materially accurate, and channel and brand review approve. Use controlled pose, crop, and background variants.

Campaign concept or moodboard: keep it clearly separate from product-proof imagery. Use freely for creative exploration and art direction.

Evidence-sensitive SKU: complex construction, transparency, patterning, reflectivity, fit claims, or channel restrictions require heightened review. Generate with fuller source evidence and require product-specialist sign-off.

How do you calculate cost per approved apparel asset?

Divide total production cost by the number of assets that pass the published approval gate. Total production cost must include source-evidence preparation, generation, retries, editing, human review, and revisions; if paid distribution is part of the deliverable, record media spend separately. A low per-generation amount can mislead badly when a SKU yields many candidates and few approved outputs.

Track two clocks. First, machine generation time: Lamina telemetry reports a median of 207s per asset. Second, end-to-end elapsed time from complete evidence packet to approved delivery. Procurement should use the second number because it includes the people and decisions that make an asset publishable.

Put test conditions beside every result: SKU count, garment types, variants per SKU, prompt set, reviewer roles, pass threshold, and test dates. One run is a benchmark observation. It is not a general performance guarantee.

What is the practical call for ecommerce teams?

Use the benchmark as an approval system, not a one-off contest between a generator and a studio reference. Start with a controlled SKU panel, define mandatory attributes before anyone sees an output, then expand AI on-model production only where approval-adjusted results meet your channel standard. You protect customer-facing product truth and keep the speed and variation generated imagery can offer.

Ali Alkan, founder of Fipera, puts the operating pressure plainly: consistency and cost must be solved together, not traded blindly against each other. The benchmark exposes both at the approved-asset level.

Fashion brands don’t have a creativity problem, they have a consistency and cost problem.
Ali Alkanfounder, Fipera