Virtual Try-OnAug 5, 2026·Data as of Aug 4, 2026

AI virtual try-on benchmark for ecommerce: a test protocol and results table for product fidelity, body/pose robustness, generation time, and publish-ready failure rates

A reproducible AI virtual try-on benchmark for ecommerce: test product fidelity, body and pose robustness, latency, and publish-ready failures without inventing model rankings.

Lamina Team

Lamina Team

Product Team @ Lamina

Ecommerce team reviewing AI virtual try-on outputs in a benchmark dashboard, comparing garment details, body poses, and generation times.

How should ecommerce teams benchmark AI virtual try-on systems?

Benchmark AI virtual try-on with a fixed held-out set, then score product fidelity, body-and-pose robustness, generation operations, and publish readiness separately. Do not hide behind one “realism” number. On a product page, the costly failure is an image that looks believable while changing the SKU, misstating fit, or falling apart on a seated pose.

Give every system the same person–garment pairs, output requirements, and retry policy. Log the model version, endpoint, region, resolution, garment and person preprocessing, prompts, masks, pose controls, seed policy, safety filters, price, and every submission outcome. A vendor-recommended setup is fair enough. Different image dimensions or source-quality thresholds are not.

The research lands in the same place. VTBench breaks real-world virtual try-on into overall quality, texture preservation, complex-background consistency, cross-category size adaptability, and hand-occlusion handling; OpenVTON-Bench separates background consistency, identity fidelity, texture fidelity, shape plausibility, and overall realism. Put those dimensions in the benchmark as fields. Do not leave them as loose notes in an approval thread.

Evidence and measured operating results for a virtual try-on benchmark
MetricValueSource
OpenVTON-Bench agreement with human judgments (Kendall’s tau)0.833arxiv.orgas of 2026-01-30
SSIM agreement with human judgments in the OpenVTON-Bench comparison (Kendall’s tau)0.611arxiv.orgas of 2026-01-30
VTONQA generated outputs8,132doi.orgas of 2026-01-06
VTONQA mean-opinion scores24,396doi.orgas of 2026-01-06
Taobao-Tryon-Benchmark paired samples1,780huggingface.coas of 2026-08-04
Baseline standing catalog try-on measured latency~17 seconds per assetuselamina.aias of 2026-08-04
Motion-and-occlusion stress try-on measured latency~30 seconds per assetuselamina.aias of 2026-08-04
Fit-boundary stress try-on measured latency~18 seconds per assetuselamina.aias of 2026-08-04
Observed cost across all three measured try-on scenarios$0.04 per assetuselamina.aias of 2026-08-04

What belongs in a standardized virtual try-on test set?

Build the test set around garments, people, poses, occlusions, and scenes that expose commercial errors. A row of front-facing studio portraits is a smoke test. It is not a release gate.

Split held-out pairs by garment class: tops, outerwear, bottoms, and dresses, adding accessories where they are in scope. Inside each class, cover plain items as well as logos, readable text, geometric prints, high-texture fabrics, dark and light colorways, closures, seams, pockets, layered looks, flat-lay sources, and on-person sources. Taobao-Tryon-Benchmark spans five garment and three accessory categories, layered multi-item scenarios, and fine-grained attributes. That should cure anyone of trusting a single-shirt set, which makes nearly any model look better than it is.

Segment people separately by body-size band, gender presentation, skin tone, front and side view, arm-raised and seated positions, hands over the garment, hair overlap, bags, and complex backgrounds. Keep those cells identical across systems. If a tool handles standing catalog poses then mangles a logo under a hand, that subgroup result is what you need to act on.

A reproducible AI virtual try-on benchmark protocol

  1. Lock test conditions before you generate

    Make a run sheet covering model and version, endpoint or commit, region, resolution, inputs, prompts, masks, pose controls, seed policy, retry policy, safety filtering, and unit price. Apply one output-size rule and one input-quality rule to every system. Keep the submitted payloads and returned files; a rerun months later needs a paper trail for any difference.

    Lock test conditions before you generate
  2. Build and freeze a held-out SKU matrix

    Tag every pair by garment category, detail complexity, body-size band, pose, occlusion condition, and scene complexity. Keep that set out of prompt tuning and vendor-selection work. Run every contender through the same matrix, then publish the completion count in every cell.

    Build and freeze a held-out SKU matrix
  3. Generate repeat attempts and keep every result

    Generate 3–5 outputs per person–garment pair under each system’s documented policy. Show first-pass performance apart from a defined recovery policy, such as one automatic retry. Track request-to-delivery time, timeouts, safety rejections, transport errors, and completed outputs. Failed jobs stay in the denominator.

    Generate repeat attempts and keep every result
  4. Use an attribute ledger to score the garment

    Before generation, label visible attributes: silhouette, neckline or collar, sleeve, hem, length, closure placement, seams, pockets, fabric texture, primary and secondary color, print geometry, logo or text, and layer order. Have two trained blind raters mark each applicable attribute as preserved, minor deviation, materially altered, or not visible. Where visible, treat logo/text, print, colorway, closures, and silhouette as critical. A material change to any critical attribute fails product fidelity.

    Use an attribute ledger to score the garment
  5. Score the person and pose on their own

    Rate identity preservation, anatomy and body compatibility, garment-to-body alignment, drape and shape plausibility, pose preservation, and occlusion handling. VTONQA’s clothing-fit, body-compatibility, and overall-quality structure gives reviewers a compact human rubric. Break out every score by body-size band, pose, garment class, and occlusion condition. Include the worst subgroup.

    Score the person and pose on their own
  6. Set the publish-ready rule before anyone reviews

    Call an output publish-ready only if it has no material critical-attribute alteration, no misleading fit or body deformation, no visible anatomy, hand, layering, or occlusion artifact at PDP zoom, and it clears crop, resolution, brand-safety, rights, and consent requirements. Keep aesthetic failures separate from API or job failures. Human art direction and approval still own the release call, especially on hero imagery.

    Set the publish-ready rule before anyone reviews

How should product fidelity be measured in virtual try-on?

Measure product fidelity by whether source-SKU attributes survive, with critical attributes able to fail an otherwise attractive image. Pixel similarity is the wrong judge for an ecommerce decision. The intended output may use a different person, pose, or background than the input.

Use a blind two-rater ledger, and keep the row-level labels. Calculate critical-attribute pass rate, logo/text pass rate, print/color pass rate, and the share of visible applicable attributes marked preserved. Resolve rater disagreements, publish inter-rater agreement, and show failure examples at PDP zoom. The garment-consistency framework defines 20 garment attributes for vision-language-model evaluation and reports stronger sensitivity to color and texture than shape and line; silhouette, construction, and closure placement therefore need explicit human review, not one automated score.

Automated assessment still has a job in triage. OpenVTON-Bench’s interpretable multimodal evaluation reported stronger agreement with human judgment than SSIM in its comparison, though a retailer should set any release threshold against its own blind labels and check false accepts around text, logos, hands, boundaries, and uncommon categories. Let a machine score sort the queue. Do not let it quietly approve a PDP asset.

How do you measure body-shape and pose robustness?

Measure body-shape and pose robustness by publish-ready rate and human score inside each body-size, pose, and occlusion cell. A portfolio average conceals the problem. One system may preserve a striped sweater, then produce implausible drape, distorted anatomy, or a misleading silhouette on another body.

Report body compatibility, garment-to-body alignment, shape plausibility, pose preservation, and occlusion handling by subgroup. Then name the lowest-performing cell and its sample size. VTBench explicitly covers cross-category size adaptability and hand-occlusion handling. FitVTON warns that diffusion try-on methods centered on 2D texture preservation can fail to show authentic garment fit across diverse body shapes.

Do not turn an image-quality result into a physical-fit claim. ISO 20947-3:2023 defines a protocol for evaluating the gap between a virtual garment and a virtual human model in digital fitting. Generative imagery needs separate validation using garment and body measurements, or a validated fitting method, before it can claim size or fit accuracy.

Lamina’s available experiment measured three generated virtual try-on scenarios on 2026-08-04: a baseline standing catalog pose, motion with occlusion, and fit-boundary stress. It did not measure product-fidelity scores, publish-ready rates, anatomy pass rates, completion counts, or body/pose subgroup outcomes, so no system ranking or quality conclusion is reported.

Observed per-asset generation cost

Baseline standing catalog try-on: $0.04 per assetMotion-and-occlusion and fit-boundary stress: $0.04 per asset

over Three scenario measurements on 2026-08-04

Generation latency under motion and occlusion

Baseline standing catalog try-on: ~17 secondsMotion-and-occlusion stress try-on: ~30 seconds

over Three scenario measurements on 2026-08-04

Generation latency under fit-boundary stress

Baseline standing catalog try-on: ~17 secondsFit-boundary stress try-on: ~18 seconds

over Three scenario measurements on 2026-08-04

What do the available virtual try-on results show?

The available measurements show a material latency penalty for motion-and-occlusion inputs, while measured unit cost remained constant across all three scenarios. The motion-and-occlusion case took about 30 seconds, against about 17 seconds for the baseline—roughly 13 extra seconds for each submitted asset. At batch volume, that changes queue planning and how many reviewable iterations fit into a production window.

Fit-boundary stress came in at about 18 seconds, near the standing baseline. Treat these as scenario measurements, not a general service-level guarantee: the supplied experiment gives no run count, percentile latency, completion rate, or hardware and endpoint conditions. Human review, revisions, asset selection, and media spend are also excluded. So $0.04 per generation is not a cost per published asset.

No product-fidelity scores, failure counts, delivery counts, or publish-ready pass rates were supplied. Do not invent a winner from that gap. Run the frozen protocol, disclose completions and failure denominators, and publish confidence intervals for every rate.

What should an ecommerce virtual try-on results table include?

Put quality, failures, and operational performance beside each other in the results table, with enough subgroup detail to show where a system cracks. Leave untested fields blank. Marketing examples are not benchmark evidence.

Use paired comparisons: every system should process the same held-out person–garment pairs. For rates, report binomial or bootstrap 95% confidence intervals plus the number of completed evaluations. In unpaired real-world try-on, images of the exact person wearing the target garment are often unavailable; VTON-IQA argues that dataset-level FID and KID miss individual-output perceptual quality. That is why the release workflow needs human labels and reference-free assessment.

Publishable results-table template — do not populate without test data
MetricValueSource
System/versionRecord exact model, version, endpoint, and configurationarxiv.orgas of 2026-01-30
Test completions (n)Report completed first-pass outputs and all submitted jobs separatelyarxiv.orgas of 2026-03-13
Product fidelityCritical-attribute pass %, logo/text pass %, and print/color pass %link.springer.comas of 2026-06-11
Body and pose robustnessBody compatibility, shape plausibility, pose preservation, and occlusion pass rate by subgroupdoi.orgas of 2025-05-26
Publish readinessFirst-pass pass %, first-pass failure %, recovery rate, and failure reasondoi.orgas of 2026-01-06
OperationsLatency p50/p95, timeout rate, API/job failure rate, and cost per submitted assetuselamina.aias of 2026-08-04
Worst subgroupReport the lowest result by garment category, body-size band, pose, occlusion condition, or detail complexityhuggingface.coas of 2026-08-04

Methodology

Original Lamina experiment run 2026-08-04. Hypothesis: Virtual try-on systems that preserve garment construction in a simple standing pose will show materially lower product-fidelity scores and higher publish-ready failure rates when the same SKU is applied to matched bodies in motion or seated/occluded poses. A reproducible Lamina-generated benchmark can quantify the trade-off among garment fidelity, body/pose robustness, latency, and operational usability without relying on proprietary retailer imagery.. Measured 3 variant(s) for cost and latency on the Lamina image engine; numbers cited here are our own measurements.