Virtual Try-OnAug 4, 2026·Data as of Aug 4, 2026

Data report: Ecommerce virtual try-on benchmark — measuring garment fidelity, model consistency, and brand accuracy across 100 catalog images

A 100-image virtual try-on benchmark should separate garment fidelity, model consistency, and brand accuracy—and report failures, not just attractive averages.

Lamina Team

Lamina Team

Product Team @ Lamina

A merchandising reviewer compares virtual try-on outputs for patterned dresses, jackets, knitwear, and trousers in a catalog quality-control grid.

What did this 100-image ecommerce virtual try-on benchmark actually measure?

This benchmark checks whether virtual try-on outputs retain the garment, keep the virtual model internally consistent, and stay safe for retailer brand presentation. It runs 100 catalog garments, evenly divided across five garment classes, and generates two outputs per SKU against the same fixed virtual-model specification. The result is 200 paired images, not a handpicked vendor-demo reel.

The prompt direction is the only variable under test. Variant A anchors on the garment with explicit preservation constraints; Variant B uses a looser styled-retail prompt. Source-image order is randomized, and reviewers never see the variant label, keeping the comparison on prompt discipline rather than rater expectation.

There is a hard limit here: no human scores or outcome measurements were supplied. The benchmark can report operational measurements now. It cannot honestly name a winner for garment fidelity, brand accuracy, model consistency, occlusion handling, or scene realism.

A paired Lamina virtual try-on protocol using 100 ecommerce catalog garments, a fixed virtual model specification, and two prompt conditions per SKU.

Garment-fidelity verdict

Hypothesis onlyNo blinded rater scores supplied

over Current benchmark record

Brand-accuracy verdict

Hypothesis onlyNo attribute-preservation measurements supplied

over Current benchmark record

Operational cost comparison

Equal per-generation cost planned for both variantsEqual per-generation cost recorded for both variants

over 200 generated benchmark outputs

Operational latency comparison

No comparative latency assumptionGarment-grounded variant completed faster in this test

over One 100-image paired run

Benchmark facts and evaluation evidence
MetricValueSource
Catalog garments in the paired benchmark100uselamina.aias of 2026-08-04
Generated benchmark outputs200uselamina.aias of 2026-08-04
Generation cost per output for both prompt variants$0.040uselamina.aias of 2026-08-04
Average generation time, garment-grounded variant~24 secondsuselamina.aias of 2026-08-04
Average generation time, styled-retail variant~27 secondsuselamina.aias of 2026-08-04
Virtual try-on images in the VTONQA human-opinion dataset8,132arxiv.orgas of 2026-01-01
Human opinion scores in VTONQA24,396arxiv.orgas of 2026-01-01
Human annotations in the VTON-IQA benchmark431,800arxiv.orgas of 2026-03-01

What results did the paired prompt test actually measure?

There is one measured result: operations. Both prompts cost $0.04 per generated asset, while the garment-grounded prompt averaged about 24 seconds and the styled-retail prompt about 27 seconds. Across hundreds of SKUs, that roughly three-second edge gives the constrained condition a little more iteration room at the same generation spend.

That is not the cost of a published asset. Human review, rejected generations, revisions, creative-direction time, and any media spend sit outside it. Nor is this a future-latency promise: input complexity, queue conditions, and the requested output treatment can all move production time.

Do not treat the faster run as proof that Variant A holds a logo, hem, print, or silhouette more faithfully. The stated hypothesis predicts that result. A hypothesis is not a scorecard.

Why score virtual try-on quality across separate dimensions?

Virtual try-on needs separate scores because a plausible-looking image can still contain the wrong SKU, a distorted neckline, or an inconsistent person. VTBench frames real-world evaluation across image quality, texture preservation, background consistency, cross-category size adaptability, and hand occlusion, with human-preference annotations alongside those cases.

Use three commercial pillars. Garment fidelity asks whether the product facts survived generation. Model consistency covers whether the person, anatomy, pose, hair, hands, and scene stay intact; brand accuracy asks whether the result reads as the actual item and is fit for a catalog. VTEdit-Bench likewise separates model consistency, cloth consistency, and overall image quality.

One generic image metric cannot settle this. VTON-IQA research notes that real try-on work often has no ground-truth photo of the same person wearing the target garment, and dataset-level FID and KID cannot express the perceived quality of one unpaired output. Put blinded human review at the center of the call.

How do you run a defensible 100-image virtual try-on benchmark?

  1. Build a stratified garment set you own

    Choose 100 real catalog garments from the categories customers buy most, then deliberately cover colors, prints, fabrics, lengths, silhouettes, and awkward cases. Include layered pieces, loose garments, hands or hair crossing the product area, inconsistent source photography, and non-studio backgrounds. A vendor-evaluation guide specifically calls for testing actual retailer imagery—including poorly lit and inconsistent inputs—instead of curated demonstrations.

    Build a stratified garment set you own
  2. Freeze variables before you generate

    Run every candidate condition from the same garment input and a predeclared virtual-model or person set. Hold output requirements, product-category instructions, and the review rubric steady. In a prompt comparison, alter only the prompt strategy; otherwise, you cannot tell whether the result came from the system, the garment, or an unlogged instruction change.

    Freeze variables before you generate
  3. Blind-score every output against anchored criteria

    Give each image to at least three trained raters without its condition label. Score garment fidelity across color, print, logo or text, texture cues, material appearance, silhouette, hem, neckline, and sleeves; score model consistency across identity, anatomy, pose, hands, hair, and background; score brand accuracy for SKU recognizability, correct category and styling, no invented details, and catalog-safe presentation. Check fabric drape, texture, and color across body shapes. A polished full image proves very little about product integrity.

    Blind-score every output against anchored criteria
  4. Log critical failures apart from average quality

    Treat a wrong garment, corrupted logo or text, materially wrong color or pattern, anatomy or identity corruption, and unsafe or off-brand presentation as critical failures. Report the mean score, 4–5 top-box rate, critical-failure rate, category cuts, and confidence intervals. A strong average can still hide enough failures to make an output unusable on a PDP or in a campaign.

    Log critical failures apart from average quality
  5. Use paired analysis; keep automated checks secondary

    For each SKU, compare the two outputs with paired mean differences, Wilcoxon signed-rank tests, and bootstrap confidence intervals. Use garment-region or semantic automated measures for diagnosis, not release approval: OpenVTON-Bench found its multi-modal protocol aligned more strongly with human judgments than SSIM, while garment-consistency research reports automated assessment is more sensitive to color and texture than shape and line. Have human reviewers inspect structure and silhouette closely.

    Use paired analysis; keep automated checks secondary

Which launch gates should an ecommerce team set before publishing virtual try-on images?

Set launch gates before generation starts: no critical wrong-SKU or brand failures in priority categories, a predeclared garment-fidelity mean and top-box rate, a separate model-consistency threshold, and an acceptable p95 generation time. The retailer sets the exact threshold. Choosing it after seeing the outputs invites a result-shaped standard.

Put brand-critical hero images into a tighter review lane; do not abandon generation over them. Strong briefs can direct on-model styling, material detail, product preservation, and campaign context at a fraction of a traditional shoot's turnaround, while a human art director still approves what represents the brand. Weak instructions produce weak evidence.

Once the quality gate clears, test the business case in the storefront. Visitor-level randomization keeps each shopper in one experimental arm for the test window, making conversion, add-to-cart, and return-rate measurement cleaner. An image benchmark establishes visual readiness. It does not establish commercial impact.

What can this benchmark let you conclude today?

This benchmark shows that, in one paired 100-image run, the garment-grounded prompt was modestly faster at the same measured generation cost. It cannot support a claim that either prompt produced better garment fidelity, model consistency, or brand accuracy, because no rater outcomes were recorded.

That restraint is the whole point. A useful ecommerce benchmark publishes its test set, fixed conditions, scoring dimensions, failure definition, and operational limits, so a merchandising team can rerun it on its own assortment. A visually convincing sample does not become a performance ranking without the underlying review data.

Methodology

Original Lamina experiment run 2026-08-04. Hypothesis: For the same 100 ecommerce catalog garment images and the same fixed virtual model specification, a garment-grounded Lamina prompt with explicit preservation constraints will produce higher garment fidelity and brand-attribute accuracy than a visually styled, less constrained prompt, while the constrained prompt may slightly reduce scene realism. Generate two outputs per catalog image (200 original benchmark images total), randomized by source-image order and reviewed blind to variant label.. Measured 2 variant(s) for cost and latency on the Lamina image engine; numbers cited here are our own measurements.