AI virtual try-on benchmark for ecommerce
A practical benchmark for testing apparel, eyewear, and accessories from one shopper photo without confusing visualisation with a fit guarantee.

Lamina Team
Product Team @ Lamina

Do not buy one “virtual try-on” feature and assume it covers apparel, eyewear, and accessories. Benchmark photo-based generative apparel rendering separately from live AR fitting for rigid goods, using the same consenting shopper photo, three representative SKUs, and one hard rule: every result is a visualisation, never a promise of garment size or fit.
The product mechanics are different. Google Cloud’s Virtual Try-On API generates an image from a person image and a clothing product image; Camweara describes real-time eyewear placement with pupillary-distance measurement and interactive 3D assets. Uwear covers the adjacent commerce flow: turn one shopper photo into a reusable avatar, apply products, then send the resulting image back to the merchant app or storefront. Upload once. Evaluate each category against the interaction it actually requires.
| Metric | Value | Source |
|---|---|---|
| CHI virtual try-on study participants | 24 | dl.acm.orgas of 2026-01-01 |
| Apparel categories reported in Google’s consumer photo-upload try-on | 4: shirts, pants, dresses, and skirts | techradar.comas of 2025-05-21 |
| Inputs for Google Cloud apparel generation | 1 person image + 1 clothing product image | cloud.google.com |
| References accepted for a complete Runware outfit | 1 person image + separate garment photos by role | runware.ai |
| Eyewear asset inputs Camweara says it can convert | 2 product images or GLB files | camweara.com |
What should an ecommerce virtual try-on benchmark measure?
A credible ecommerce virtual try-on benchmark measures repeat use, visual fidelity, category-appropriate placement, storefront delivery, privacy handling, and whether the fit disclaimer is clear—not just whether an image appears. The 2026 CHI study of 24 participants found that virtual try-on cut exploration time and product-detail views while giving participants clearer expectations before delivery. They also raised fit accuracy and asked for reliability signals such as confidence scores.
Start with one clear, front-facing, full-body photo from a consenting shopper; keep that identity input fixed. Choose one patterned or structured apparel SKU, one eyeglass frame, and one jewelry or accessory SKU. Apparel reveals whether prints, seams, sleeves, layering, and body pose survive generation. Frames reveal face alignment. Accessories show whether the system can hold plausible hand, face, or body placement instead of pasting on a catalog cutout.
Record mobile upload-to-first-result time, then time the second-product flow without another upload. Review every output at PDP scale, then close up. Log skin tone, facial features, pose, garment geometry, logo placement, texture, frame position, and accessory attachment: do they stay credible? A quick first image fails the test if the next selection sends the shopper through onboarding again or alters their identity enough to wreck comparison.
Get the pass/fail language right. Describe apparel output as an AI visualisation of appearance, styling, and garment detail; never call it a size recommendation or fit prediction. CHI participants asked for fit accuracy and reliability cues, so a visible account of what the image can and cannot show belongs in the product experience, not buried in legal fine print.
| Category | Recommended interaction | What the live test must prove | Documented implementation reference | Source |
|---|---|---|---|---|
| Apparel | Photo-based generative try-on | The shopper’s identity and pose persist while garment detail appears plausibly on the body. | Google Cloud Virtual Try-On API: person image plus clothing product image | cloud.google.com |
| Multi-item outfit | Photo-based composited outfit generation | Separate garments tagged by role resolve into one coherent outfit rather than competing layers. | Runware P-Image-Try-On: person image plus garment references by role | runware.ai |
| Eyewear | Real-time AR fitting | Frame position remains stable on the face and the shopper can assess placement interactively. | Camweara: real-time try-on and automatic pupillary-distance measurement | camweara.com |
| Watches and selected accessories | Live AR or category-specific placement test | The product tracks the relevant body area and works from the merchant’s product assets. | TryOn Virtual: live AR try-on for eyewear and watches | apps.shopify.com |
| Cross-category catalog flow | Reusable photo/avatar workflow | One shopper image can be reused and the image returns cleanly to the merchant experience. | Uwear: reusable avatar from one shopper photo | uwear.ai |
How should apparel try-on be tested from one shopper photo?
Judge apparel try-on on identity preservation and garment-detail preservation at the same time. Google Cloud documents an apparel flow built from a person image and a clothing image. Breuninger’s “be your own model” work shifted from pre-selected models to a shopper-selfie approach after user feedback. Test first-person representation, not a generic model swap.
Lead with the structured or patterned garment. Have reviewers inspect the same five points every time: facial identity, body pose, silhouette, print or texture, and the edges where sleeves, hems, collars, or straps meet the body. Then run a plain garment in that category. The gap tells you whether the system only behaves when product detail is simple.
Runware’s P-Image-Try-On documentation sets out a separate outfit test: the same person image plus separate garment photos tagged by role. That matters for merchants selling coordinated looks, and it needs tougher review. Check that an outer layer does not wipe out an inner one, garments keep their intended role, and the output has not invented a different neckline, closure, or pattern.
Keep category coverage on the scorecard. Reporting on Google’s consumer photo-upload feature named shirts, pants, dresses, and skirts—a reminder that apparel availability can be category-limited even when the product is sold as broad try-on. Before promising this across the catalog, make the vendor name the supported SKU classes exactly.
Why should eyewear and accessories use a different pass bar?
Give eyewear a live-fitting pass bar. Frame placement, rather than a flattering static render, is the purchase-critical question. Camweara advertises real-time eyewear try-on, automatic pupillary-distance measurement, and conversion of two product images or GLB files into interactive 3D try-on. Test those controls on a phone under ordinary indoor lighting, during a head turn, and with a brief occlusion such as a hand passing near the face.
Use the same frame for every finalist. Inspect bridge position, lens alignment, temple placement, scale, and tracking recovery. Where the product exists as both product photography and a GLB asset, test both paths. Shoppers should be able to inspect placement without the system replacing their face or letting the frame slide after movement.
Test accessories where they are actually worn. Assess jewelry at the face, hand, neck, or wrist as appropriate; watches need wrist tracking, while bags need body-relative positioning. TryOn Virtual’s Shopify listing claims photo-based AI try-on across a catalog alongside live AR try-on for eyewear and watches, and says shopper data is not stored. Verify privacy retention, deletion, and asset-use terms in the merchant’s own account and agreement before treating them as acceptance criteria.
One engine may still make operational sense. It has to earn that convenience. One cross-category candidate mentioned in the market, 1Match, claims clothing, eyewear, and jewelry with face, hand, and body tracking; its limited review footprint makes it a finalist for testing, not proof that it leads every category.
Run the one-photo ecommerce virtual try-on test
Prepare one consented identity input
Capture a clear, front-facing, full-body shopper photo, then document consent, the intended retention period, and the deletion path. Run that exact photo through every apparel, eyewear, and accessory test. Otherwise you are comparing inputs, not products.

Choose three difficult but sellable SKUs
Use one structured or patterned garment, one eyeglass frame, and one accessory. Keep the original product images and any available GLB file. Every vendor gets the same SKUs.

Run the first-use mobile journey
On a phone, time upload through to the first usable result. Capture each required permission, crop, image-quality rejection, product-selection step, and result screen. Test a reusable-avatar workflow only after the first result exists.

Run the repeat-use journey
Apply the other two SKUs without uploading another shopper photo. Uwear’s described reusable-avatar workflow makes this a decisive check: confirm that the merchant experience reuses the initial shopper representation and returns generated images where shoppers browse.

Apply category-specific review criteria
For apparel, score identity, pose, garment shape, and detail. For eyewear, score bridge alignment, scale, and tracking during movement. For accessories, score body-area placement and recovery after partial occlusion. Save a screen recording and mark every criterion pass, conditional pass, or fail.

Review PDP integration and customer language
Confirm how try-on opens from the PDP, how a result is saved or shared, and whether the page says the output is a visualisation rather than a measurement of actual fit. Check shopper-image retention and deletion controls before any production launch.

Repeat on edge conditions before rollout
Repeat the identical six-run matrix with a second consented shopper photo and the same three SKUs. Use that run to spot input sensitivity, not to replace the first result. Give brand-critical hero moments art-direction review before publication.

| Tier | Price | Included | Best for |
|---|---|---|---|
| Photo-rendering pilot | Vendor quote required | 3 first-use apparel runs + 3 repeat-use runs | Merchants validating person-image and garment-image generation on a single apparel SKU |
| Cross-category validation | Vendor quote required | 6 runs per shopper photo across apparel, eyewear, and accessories | Teams comparing reusable-photo flow against category-specific fitting |
| PDP launch evaluation | Vendor quote required | Pilot matrix plus mobile recording, privacy review, and PDP integration review | Teams preparing a production commerce implementation |
One shopper photo, three product categories, first-use and repeat-use journeys
6 runs × vendor-quoted per-run or pilot price1 shopper photo × 3 SKUs × 2 journeys = 6 documented test runs
Two shopper photos for an input-sensitivity check
12 runs × vendor-quoted per-run or pilot price2 shopper photos × 3 SKUs × 2 journeys = 12 documented test runs
Outfit plus eyewear and accessory validation
6 category interactions × vendor-quoted price1 apparel outfit run + 1 eyewear live-fitting run + 1 accessory placement run, repeated after initial setup
What should a merchant ask vendors before paying for a pilot?
Make vendors demonstrate the exact input contract. Virtual try-on products do not all take the same assets. Google Cloud documents a person image plus a clothing product image for apparel. Runware documents a person image plus separate garment photos tagged by role for a complete outfit. Camweara describes an eyewear route built from two product images or GLB files. Those differences change catalog preparation, API design, and what your team must produce before a shopper sees anything.
Ask whether the result is generated imagery, real-time AR, or both. Make the vendor show each one on a mobile PDP. A photo-based garment result can support styling exploration from a shopper selfie. A rigid eyeglass frame needs interactive face placement. A static eyewear mockup does not prove live fitting just because both are called virtual try-on.
Get written answers on shopper-photo storage, deletion, reuse, and merchant control. TryOn Virtual claims shopper data is not stored; check that claim against the account configuration, privacy policy, and contract that govern your store. Ask whether output images can be retained, removed on request, and where image processing runs.
Make the vendor label capability boundaries SKU by SKU. Google’s consumer photo-upload coverage was reported across shirts, pants, dresses, and skirts; use that as the model for a rollout list. A credible launch plan names included classes, leaves unsupported ones out, and assigns a human reviewer to brand-critical outputs.
How should ecommerce teams interpret try-on engagement and conversion claims?
Treat engagement, conversion, and returns claims as hypotheses for your own controlled rollout, never as portable forecasts. Dopplr co-founder Sresht Agarwal connects try-on engagement with shopper adoption and reports client outcomes, yet the published figures come from Dopplr’s client base and implementation context. Compare exposed and unexposed PDP sessions in your pilot, then split apparel, eyewear, and accessory results instead of pooling products with different buying behavior.
Measure the decision path around the visualisation: try-on open rate, completion rate, repeat-product use after the original upload, result-save or share behavior, add-to-cart rate, purchase rate, and return reason. Add a short post-use prompt asking whether shoppers understood the image was visual, not a sizing guarantee. The CHI study’s findings on clearer pre-delivery expectations and demand for reliability signals justify that comprehension check.
Keep a review queue for product-detail failures. Apparel patterns, logos, jewelry attachment, and frame alignment can all dent trust, even when the overall image looks convincing. Generative try-on can produce complex styling and believable material detail. Human art direction and approval still control outputs used as brand-critical commerce assets.
“The most important metric that a VTO impacts is the engagement rate which indicates whether or not shoppers are taking to the try-on experience. On an average, we see a 3.5X engagement on products offering the try-on experience. Across our clients, we have impacted conversions by about 27 per cent and reduced returns rate up to 37 per cent,”
What is the right production decision after the benchmark?
Choose the implementation that clears the category-specific test, even if that means photo-based generation for apparel and live AR for eyewear. Google Cloud’s documented person-and-garment image flow, Breuninger’s shopper-selfie deployment, and Camweara’s stated real-time eyewear controls describe distinct jobs, not one universal interaction.
Launch with an explicit category map: supported apparel classes, supported frame types, supported accessories, accepted asset formats, image-retention rules, and the exact fit language shown to customers. Keep the benchmark recordings and scorecards as the acceptance baseline for catalog changes after the first production release.
The operational target is straightforward: shoppers upload once, see credible category-appropriate visualisations, and understand where appearance preview ends and physical fit begins. That beats a broad virtual-try-on claim that bundles apparel rendering, face tracking, and accessory placement into one untested promise.
FAQ: Does virtual try-on predict clothing size or fit?
No. Present virtual try-on as a visualisation of how a product may look, never a guarantee of size or physical fit. The 2026 CHI study found users valued clearer expectations before delivery, while also emphasizing fit accuracy and asking for reliability signals. PDP language needs to state that boundary plainly.
FAQ: Can one shopper photo be reused across a catalog?
Yes. Uwear describes a workflow where one shopper photo becomes a reusable avatar, products are applied to it, and generated images return to the merchant app or storefront. Verify repeat use without a second upload in the merchant’s actual mobile experience.
FAQ: Can one platform handle apparel, eyewear, and jewelry?
A cross-category workflow is possible, provided you test it by interaction type. Apparel is documented as photo-based person-and-product generation. Eyewear vendors such as Camweara describe real-time fitting and pupillary-distance controls. Evaluate cross-category claims with the same shopper photo and separate pass criteria for each category.
FAQ: What product assets should a merchant prepare?
Prepare clean apparel product images, frame imagery, and accessory assets for the selected SKUs; keep any available GLB files for eyewear tests. Google Cloud documents one clothing product image for its apparel API. Runware documents separate garment references by role for outfits, and Camweara describes conversion from two product images or GLB files for interactive eyewear try-on.
Continue reading

AI virtual try-on accuracy checks for ecommerce
A three-condition AI virtual try-on test shows what timing and cost data can prove—and why apparel brands still need SKU-level checks before publishing.

Lamina Team
Product Team @ Lamina

AI virtual try-on benchmark: protocol and results
A reproducible ecommerce virtual try-on benchmark protocol, with the measured Lamina cost and latency results separated from quality metrics that remain pending.

Lamina Team
Product Team @ Lamina