AI virtual try-on benchmark: protocol and results
A reproducible ecommerce virtual try-on benchmark protocol, with the measured Lamina cost and latency results separated from quality metrics that remain pending.

Lamina Team
Product Team @ Lamina

What does this AI virtual try-on benchmark actually establish?
Right now, this benchmark establishes unit cost and measured generation latency. It does not establish that any Lamina prompt variant produces better garment fidelity or fewer publish-ready failures. The frozen design runs Standard, Apparel-lock, and Detail-first across the same 720 garment-model-pose cells, using three fixed seeds per cell; quality columns stay pending until blinded scoring is finished.
That split matters. An output that alters a logo, drops a closure, or redraws a hem cannot go on a PDP, no matter how polished the person appears. So the benchmark tracks product fidelity, body and pose integrity, request-to-delivery time, and publish-ready failure separately rather than burying them inside one aesthetic score.
The evidence package is the real deliverable: a frozen benchmark pack, scoring guide, anonymized rater-level CSV, run manifest, and a contact sheet showing the same 24 randomly selected cells across variants. Buyers can inspect the breakage themselves instead of taking selected vendor examples on faith.
| Metric | Value | Source |
|---|---|---|
| Frozen garment set | 18 Lamina-generated catalog garments: 6 tops, 4 dresses, 4 outerwear pieces, and 4 bottoms | uselamina.aias of 2026-08-08 |
| Test matrix | 720 garment-model-pose cells from 18 garments, 8 synthetic adult model identities, and 5 poses | uselamina.aias of 2026-08-08 |
| Planned output volume | 6,480 outputs: 2,160 per prompt variant, with 3 fixed seeds per cell | uselamina.aias of 2026-08-08 |
| Measured cost per generated output | $0.040 for all three variants | uselamina.aias of 2026-08-08 |
| Standard prompt measured generation time | ~25 seconds | uselamina.aias of 2026-08-08 |
| Apparel-lock prompt measured generation time | ~24 seconds | uselamina.aias of 2026-08-08 |
| Detail-first prompt measured generation time | ~27 seconds | uselamina.aias of 2026-08-08 |
| Human-alignment comparison in OpenVTON-Bench | Kendall’s τ of 0.833 versus 0.611 for SSIM | arxiv.orgas of 2026-01-30 |
Lamina’s supplied frozen virtual try-on protocol measured cost and generation time for three prompt variants. Blinded quality scoring, publish-ready labels, and failure outcomes have not yet been published.
Generation cost per output
over Measured protocol; 2,160 planned outputs per variant
Measured generation time
over Measured protocol
Measured generation time
over Measured protocol
Why don’t FID and KID cover ecommerce virtual try-on?
FID and KID cannot tell an ecommerce team whether a single try-on image kept the sellable garment intact. Production try-on usually has no true reference image of that exact person wearing the target garment, and FID and KID measure distribution-level similarity rather than the perceptual quality of one output.
Review images one by one. Put the source person image and garment references in front of the reviewer, then ask whether the output retained color, print, logo or text, fabric character, construction details, and silhouette without damaging the person’s body or pose.
The research points the same way: OpenVTON-Bench evaluates texture, identity, background consistency, shape plausibility, and realism; VTON-QBench explicitly covers garment fidelity and person-specific detail preservation. One similarity score leaves ample room for an attractive image that misleads shoppers about the product.
What belongs in an ecommerce virtual try-on test set?
Lock the test set before model testing. Build it around garment and person cases most likely to expose product mistakes. Stratify the holdout by category, dominant color and contrast, print or logo density, material texture, silhouette, layering, body-size presentation, pose, occlusion, and background complexity.
Keep the hard cases in their own slice. Raised arms can warp sleeves and armholes; seated and walking poses put hem geometry and garment-body alignment under pressure; hands, hair, and accessories show whether the model can retain occlusion without inventing anatomy. VTBench names texture preservation, complex-background consistency, cross-category size adaptability, and hand-occlusion handling as critical evaluation dimensions.
The supplied Lamina protocol is a sound synthetic starting point: every garment gets front and back product cards, a material/detail crop, and an attribute sheet covering dominant color, closure, neckline, sleeve, hem length, print or logo, and intended silhouette. For a production buying decision, rerun that design on a consented holdout of real catalog garments and approved person images.
How do you run a fair AI virtual try-on benchmark?
Freeze the inputs and production settings
Make the garment cards, person reference cards, pose cards, and structured product attributes before generating a single image. Fix output resolution, aspect ratio, model version, reference-image order, inference settings, retry policy, and queue region, leaving the prompt variant as the only deliberate difference.

Use a fixed attempt count for every cell
Run the same number of attempts for every person-garment-pose case—for example, three fixed seeds. Log request time, delivery time, retries, success or error state, output resolution, source-image hashes, prompt, seed, model version, and output hash. Failed outputs stay in the record.

Blind the outputs, then score them separately
Randomize the filenames. Have at least two trained reviewers compare every output with both source inputs, scoring product fidelity apart from face or identity preservation, body proportions, pose, anatomy, garment-body alignment, and occlusion handling. Bring in a third reviewer only if the score gap exceeds a pre-set threshold.

Set release labels and publish the denominator
Give every completed output a Publish, Repair, or Reject label, with reason codes. Report safety blocks, malformed files, blank outputs, and timeouts in their own operational-failure line; reporting only successful generations hides them. Calculate bootstrap 95% confidence intervals over garment-model-pose cells, not repeated seeds treated as independent products.

Release the audit trail beside the table
Publish the benchmark pack, scoring rubric, anonymized rater-level data, run manifest, and a same-cell contact sheet. Readers should see where a variant wins, where it fails, and whether its result relies on easy frontal images.

How should reviewers score product fidelity in virtual try-on?
Treat product fidelity as a garment-specific release gate, with explicit critical errors that can override a respectable average. Use six 0-to-2 checks: color and material appearance; print or pattern alignment; logo or text preservation; texture and fine detail; construction details such as buttons, collars, lapels, and seams; and silhouette, sleeve, and hem geometry.
Flag wrong logo or text, materially wrong color or print, a missing major construction feature, or a misleading silhouette as a critical product error. A realistic face cannot excuse selling the wrong item.
Automated garment-consistency assessment can triage large batches, particularly for color and texture. Keep it out of release authority: a supplied study found VLM evaluation more sensitive to color and texture than shape and line dimensions, exactly where tailored, layered, and structured garments require human scrutiny.
What is AI virtual try-on’s publish-ready failure rate?
Publish-ready failure rate is the share of all completed outputs labeled Repair or Reject after blinded review: (Repair + Reject) divided by all completed outputs. Report critical-misrepresentation rate beside it as the stricter subset for assets whose product errors could mislead a shopper.
Keep this separate from API reliability. Timeouts, safety blocks, blank files, and malformed outputs belong in their own reporting and remain in the full run record; the publish-ready label asks whether your creative team can place the asset on a product page without correcting it.
The supplied Lamina run does not yet have an observed publish-ready failure rate. Any statement that Apparel-lock reduces failures is projected, not measured, until blinded labels are applied to frozen outputs and the full denominator is released.
What results table should an ecommerce team publish?
Publish one table that puts quality, robustness, speed, and release risk side by side, then break out the hard-case slice. Do not turn Lamina’s available latency readings into a product-fidelity leaderboard: no blinded fidelity, robustness, or publish-ready outcomes have been supplied.
Unit cost is the same across variants. The planned 2,160-output run therefore costs $86.40 per variant and $259.20 across all 6,480 outputs, before human review, revisions, or media spend. Those figures cover generated outputs rather than approved assets, so cost per approved image must wait for the missing publish-ready rate.
Lamina virtual try-on results table: what has been measured, and what remains pending?
The sole observed comparison is that Apparel-lock was slightly faster than Standard, while Detail-first took longer. Every quality result remains pending. These timings are single supplied measurements under the frozen protocol, not a general latency guarantee; P95 latency has not been reported.
| Variant | Outputs planned | Product fidelity /100 | Pose robust | Body/identity robust | Measured generation time | Publish-ready failure | Status |
|---|---:|---:|---:|---:|---:|---:|---|
| A: Standard | 2,160 | Pending | Pending | Pending | ~25 seconds | Pending | Baseline |
| B: Apparel-lock | 2,160 | Pending | Pending | Pending | ~24 seconds | Pending | Primary test |
| C: Detail-first | 2,160 | Pending | Pending | Pending | ~27 seconds | Pending | Robustness test |
Which virtual try-on model should an ecommerce team pick?
Pick the model with the lowest critical-misrepresentation and publish-ready failure rates among those meeting your latency service-level objective. Do not choose from a composite beauty score. Aggregate quality can hide poor results on logo-heavy, patterned, dark, sheer, tailored, or layered SKUs.
Speed still determines the iteration budget. As an external operational reference—not a cross-vendor comparison—Genlook reported a 9.2-second median and completion within 14.4 seconds for 95% of try-ons across more than 156,000 completed try-ons at 577 Shopify stores; treat that as context for production expectations, not proof of product fidelity.
AI virtual try-on can produce on-brand on-model imagery across concepts, body presentations, poses, and garment styling at a fraction of a conventional production cycle. Human art direction and approval remain the control point, especially on brand-critical hero images. Use a precise garment brief, frozen inputs, and a reviewable evidence trail to get the category’s speed without letting a plausible image misrepresent the SKU.
| Metric | Value | Source |
|---|---|---|
| OpenVTON-Bench scale | Roughly 100K high-resolution image pairs across 20 fine-grained garment categories | arxiv.orgas of 2026-01-30 |
| VTONQA evaluation set | 8,132 images from 11 representative VTON models and 24,396 mean-opinion scores | arxiv.orgas of 2026-01-06 |
| VTON-QBench evaluation set | 62,688 try-on images from 14 VTON models, with 431,800 quality annotations from 13,838 qualified annotators | arxiv.orgas of 2026-03-13 |
| Genlook operational latency reference | 9.2-second median generation time; 95% completed within 14.4 seconds | genlook.appas of 2026-07-17 |
| ISO scope limitation | ISO 20947-2:2020 covers virtual-garment pattern-cutting and clothing-simulation modules, not a complete generative image-based ecommerce VTO benchmark | iso.orgas of 2026-08-08 |
Methodology
Original Lamina experiment run 2026-08-08. Hypothesis: A garment-constrained Lamina virtual try-on prompt will improve product fidelity and reduce publish-ready failures versus a standard ecommerce try-on prompt, while adding modest generation latency. Protocol: create an original, frozen benchmark pack before testing: 18 Lamina-generated catalog garments (6 tops, 4 dresses, 4 outerwear pieces, 4 bottoms), each with a front product card, back product card, material/detail crop, and a structured attribute sheet (dominant color, closure, neckline, sleeve, hem length, print/logo, and intended silhouette). Create 8 distinct synthetic adult model reference cards in Lamina and 5 pose cards per identity (front-standing, 3/4 standing, arms raised, seated, and walking), for 720 garment-model-pose cells. Run every cell with 3 fixed seeds for each variant (2,160 outputs per variant; 6,480 total for three variants). Keep output resolution, aspect ratio, model version, reference-image order, inference settings, retries, and queue region fixed; record a manifest containing prompt, seed, source-image hashes, timestamps, Lamina model/version, and output URL/hash. Randomize output filenames and have three blinded raters independently score images; adjudicate only when ratings differ by more than 1 point. Pre-register exclusion rules: do not replace failed images, and count safety blocks, blank outputs, and malformed files as failures. Report bootstrap 95% confidence intervals over garment-model-pose cells, not individual reruns. Results table (populate only after the frozen run; do not present projected values as observed): | Variant | N outputs | Product fidelity /100 (95% CI) | Pose robust (%) | Body/identity robust (%) | End-to-end time, median / P95 s | Publish-ready failure (%) | Notes | |---|---:|---:|---:|---:|---:|---:|---| | A: Standard | 2,160 | pending | pending | pending | pending | pending | baseline | | B: Apparel-lock | 2,160 | pending | pending | pending | pending | pending | primary test | | C: Detail-first | 2,160 | pending | pending | pending | pending | pending | robustness test | Publish the benchmark pack, scoring guide, anonymized rater-level CSV, run manifest, and a contact sheet showing the same 24 randomly selected cells across all variants. This produces original Lamina imagery and an auditable evidence set rather than relying on vendor examples.. Measured 3 variant(s) for cost and latency on the Lamina image engine; numbers cited here are our own measurements.
Continue reading

AI ecommerce creative benchmark: fidelity, brand, time
A controlled benchmark framework for judging AI virtual try-on, product editing, and product reels on SKU truth, repeatable brand control, and approval-ready time.

Lamina Team
Product Team @ Lamina

How accurate is AI virtual try-on for ecommerce?
AI virtual try-on has no defensible single accuracy rate. Benchmark garment fidelity, model consistency, and campaign readiness by SKU and input condition.

Lamina Team
Product Team @ Lamina

Ecommerce virtual try-on benchmark: 100 images
A 100-image virtual try-on benchmark should separate garment fidelity, model consistency, and brand accuracy—and report failures, not just attractive averages.

Lamina Team
Product Team @ Lamina