Virtual Try-OnPricing guideAug 21, 2026·Data as of Jul 17, 2026

AI virtual try-on A/B test for Shopify product pages

A 30-day Shopify PDP test plan for measuring AI virtual try-on conversion lift, return-rate movement, and fully loaded return-cost savings.

Lamina Team

Lamina Team

Product Team @ Lamina

Shopify apparel product page on a phone showing an AI virtual try-on widget beside the size selector and add-to-cart button

Run Shopify virtual try-on as a 50/50 randomized PDP experiment. Read conversion immediately, then measure returns only after a post-delivery cohort has matured. Never compare shoppers who opted into try-on with those who skipped it; intent sorted those groups before your analysis began.

A disciplined 30-day pilot answers two different business questions: does a visible try-on experience change completed-order conversion per eligible PDP session, and does it cut the expensive returns driven by fit or expectation? Get the conversion read quickly. Let the return data age properly.

Benchmarks that should shape the test—not predict its outcome
MetricValueSource
Try-ons completed on phones in Genlook’s Shopify-store dataset; make mobile rendering and reporting a launch requirement90%genlook.appas of 2026-07-17
Add-to-cart rate in sessions with a try-on in Genlook’s observational dataset; use as an engagement reference, not causal lift29.2%genlook.appas of 2026-07-17
Add-to-cart rate in sessions without a try-on in the same observational dataset9.4%genlook.appas of 2026-07-17
Conversion uplift reported by GANT’s 3D visualization A/B test across 274 products and 13 markets6.3%fibbl.comas of 2024-12-16
Return reduction reported in Faslet’s controlled DIDI virtual try-on case study13.1%faslet.meas of 2026-04-14
Year-over-year reduction in size-related returns reported by Under Armour Japan for October–December 202327%virtusize.comas of 2023-12-31

What experiment design works for Shopify virtual try-on?

Assign at the cookie level on the first eligible PDP view, then hold that assignment for 30 days. Someone first placed in treatment keeps seeing the try-on widget on every eligible return visit; someone assigned to control does not see it. That keeps returning shoppers from crossing between variants and muddying the result.

Set eligibility before launch. Choose stable, high-traffic apparel or footwear hero SKUs with sufficient inventory, no planned photography replacement, and no SKU-specific promotion scheduled during the test. Hold price, discounting, shipping promise, size chart, PDP copy, traffic acquisition, and fulfillment policy constant across both arms. Virtual try-on is the sole planned difference.

Make completed-order conversion per eligible PDP session the primary metric: completed purchases divided by eligible PDP sessions. It beats widget clicks because every assigned shopper stays in the denominator, including shoppers who never open try-on. Analyze the assigned arm, not actual try-on usage. That is the intention-to-treat result a merchandising team can use.

Keep the variant in a first-party cookie or app visitor ID, then write the assignment into the cart, checkout, and order record as identifiers become available. Device switching can still create two apparent visitors unless login or email joins the identity. Report that constraint; never reassign a known shopper after the first eligible view.

How do you run a 30-day Shopify virtual try-on test?

  1. 1. Lock the SKU cohort and state the hypothesis

    Select the eligible apparel or footwear PDPs, then freeze the list for the 30-day conversion read. Pre-register one primary hypothesis: the visible AI virtual try-on widget changes completed-order conversion per eligible PDP session. Set secondary metrics before traffic lands—add-to-cart rate, checkout-start rate, revenue per eligible PDP session, average order value, widget completion, generation errors, and mature return rate.

    1. Lock the SKU cohort and state the hypothesis
  2. 2. Put the widget where fit gets decided

    On current Shopify themes, add the try-on experience as a native product-page app block: Online Store, Themes, Customize, product template, Add block, app block, Save. Place the treatment beside the size selector or Add to Cart, never under reviews or behind separate navigation. Match that placement across desktop and mobile. Then confirm it does not sit over variant selectors, sticky purchase controls, or accelerated checkout buttons.

    2. Put the widget where fit gets decided
  3. 3. Randomize eligible visitors before they interact

    Give each first eligible PDP visitor an equal-probability control or treatment assignment. Control gets the unchanged PDP. Treatment gets that PDP with the visible try-on block added. Save the assignment before anyone opens the widget, retain it on eligible return visits, and send it into order data. Assigning after a photo upload or try-on start makes engagement the selection rule.

    3. Randomize eligible visitors before they interact
  4. 4. Instrument the PDP path and the widget path

    Keep Shopify and GA4 funnel events for view_item, add_to_cart, begin_checkout, and purchase. Add tryon_widget_view, tryon_start, photo_upload_success, tryon_generated, tryon_error, and tryon_add_to_cart. Every event needs variant, SKU, product type, device class, anonymous visitor ID, and timestamp. Get explicit photo consent, and settle uploaded-image retention rules before the first customer session.

    4. Instrument the PDP path and the widget path
  5. 5. Connect orders, returns, exchanges, and refunds

    Query or export Shopify and order-management records using the experiment assignment stored on the order. Capture returned units, return reason, exchange status, refund amount, cancellation status, and delivery status. Your reason codes should separate size or fit, style or expectation, quality or defect, bracketing, and other causes. Otherwise a real fit signal disappears inside all-return totals.

    5. Connect orders, returns, exchanges, and refunds
  6. 6. Freeze conversion, then let returns mature

    On day 30, freeze the eligible-session and order cohort for conversion analysis. For returns, include only orders with the same observation window in both variants—for example, 14 days after delivery or the merchant’s established return window. Keep reading returns after the PDP test ends. One arm cannot carry late returns while the other has not yet had time to produce them.

    6. Freeze conversion, then let returns mature

Which events belong in a Shopify try-on test?

Track the standard commerce funnel and a distinct try-on funnel. Shopify’s product and purchase events show where assigned shoppers leave the purchase path; custom widget events tell you whether relevance, photo upload, generation, or the final PDP handoff failed.

The minimum sequence is view_item, tryon_widget_view, tryon_start, photo_upload_success, tryon_generated, tryon_error, tryon_add_to_cart, add_to_cart, begin_checkout, and purchase. Put the experiment variant and SKU on every event. Without both fields, a top-line GA4 report cannot separate a weak widget flow on one shoe style from a wider treatment result.

Treat mobile as its own operating surface. Genlook’s Q2 2026 report found that 90% of its completed try-ons occurred on phones. Before paid traffic reaches the experiment, test photo permissions, camera and gallery behavior, widget load state, and the sticky Add to Cart interaction on common mobile devices.

What rule decides whether the test wins?

Call treatment a conversion winner only when its two-sided 95% confidence interval for the completed-order conversion difference sits above zero and clears the pre-registered minimum detectable effect. Set that effect at the smallest lift worth paying for after app fees, implementation work, and review time. A percentage lift that cannot cover those costs is not a commercial win.

Calculate required eligible sessions before launch from the control conversion baseline, minimum detectable effect, a two-sided 95% confidence level, and chosen statistical power. Do not end the test because a daily chart looks good. Watch the 30 days for sample-ratio mismatch, failed order linkage, photo-upload failures, and treatment errors, then make the winner call at the planned readout.

Set a pre-launch guardrail for worsening error rate and a separate PDP-performance threshold. Pause and repair if treatment technical failures exceed the agreed tolerance, even when add-to-cart rate climbs. Conversion does not excuse a widget that strands a meaningful share of shoppers after image upload.

Treat SKU, device, market, and new-versus-returning-customer cuts as exploratory unless each has adequate pre-planned sample. Every extra cut gives chance another shot at finding a flattering result by accident. The all-eligible-session result remains the decision metric; a mobile or footwear cut can shape the next test, not rewrite the first.

Why use completed-order conversion rather than try-on engagement?

Completed-order conversion leads because it measures shoppers assigned to see treatment, including those who never touch it. Try-on completion can expose adoption and friction. It cannot show that the experience caused an order.

GANT’s large 3D-First experiment is a useful readout standard: 274 products, 13 markets, more than 400,000 users, and a statistically significant conversion result. The experience was 3D visualization rather than necessarily photo-upload AI try-on. Use the figure as a hypothesis benchmark, never a Shopify forecast.

My team supports all eCommerce teams and other business functions in gaining a deeper understanding of our customers, the user experience and business performance through a continuous optimization process. For this test we used the conversion rate as the primary metric, and our data confirms a 6.3% uplift in conversions with 95% statistical significance,
Andrej ObloginHead of Growth, GANT

How should return rate and return-cost savings be measured?

Pull returns from Shopify or order-management records, not the GA4 purchase funnel alone. A Shopify and GA4 instrumentation review notes that refund does not fire natively in that implementation. Analysts relying only on funnel events inherit a purchase-to-refund gap.

Define a return as merchandise physically sent back and received or approved through the merchant’s return workflow. A refund is the financial transaction that can follow a return, cancellation, price adjustment, or undelivered order. Keep those measures separate.

For the main return analysis, remove pre-shipment customer cancellations and undelivered or lost parcels from merchandise-return rate; report them separately as fulfillment outcomes. Count partial returns at unit level and order level. Count an exchange as a returned unit plus a replacement order, then report exchange rate apart from cash-refund rate. Returning one item from a three-item order cannot become a full-order return in a unit-rate calculation.

Calculate return rate per order and per unit. Then isolate size and fit returns: Under Armour Japan’s Virtusize case study used size-related returns, not every return cause, in its reported outcome. A virtual try-on tool may change fit confidence while defects, delivery damage, and buyer remorse remain untouched.

Build a fully loaded avoided-return model: avoided returned units multiplied by fully loaded cost per return. Include return label and transportation, receiving and inspection, order picking, repackaging, customer-service labor, payment or handling fees, markdown or disposal loss, tied-up capital, and inventory holding. A Scandinavian footwear-retailer study explicitly modeled handling, tied-up capital, inventory holding, transportation, and order-picking costs. Return shipping alone is far too narrow.

GANT’s ecommerce team puts the customer case plainly: visualization can reduce the need to order first just to see how footwear looks when worn. Split fit and expectation reasons from every other return cause for exactly that reason.

The greatest value I see in Fibbl’s technology is that it allows us to offer our customers a more engaging and enjoyable experience beyond the usual on our site. Additionally, it enables our customers to visualize how the shoes look when worn without having to order them first. This helps address one of the biggest challenges we face in digital commerce: the difficulty of experiencing products—touching, feeling, and trying them on—through a screen,
André LagoGlobal E-commerce Director, GANT

Returns involve far more than sending money back. Zach Footwear founder and CEO Robbie Thompson describes the manual work behind labels, repacking, customer communication, and damaged packaging. Put those cost categories in the merchant’s return-cost worksheet.

It’s a nightmare! Aside from the routine process of handling a return, there are additional tasks like replacing ruined boxes, repacking shoes for shipping, customer communication, creating shipping labels, and providing drop-off instructions and so on. It is possible to reduce costs by handling returns manually this way, but it’s very time-consuming,
Robbie ThompsonFounder and CEO, Zach Footwear
TierPriceIncludedBest for
Conversion pilotApp subscription + theme setup + analytics implementationUse the vendor’s contracted allowanceA 30-day PDP conversion decision
Conversion and returns pilotConversion-pilot costs + return-data joining + mature-cohort analysisUse the vendor’s contracted allowanceMerchants that need a return-cost decision
Scale decisionPilot cost + rollout support + creative and QA reviewForecast from eligible PDP traffic and expected try-on startsMulti-SKU or multi-market expansion
Model the pilot as an incremental business case. These are planning inputs, not vendor price quotes.

Conversion value calculation

Incremental contribution before app and implementation costs

Incremental completed orders × contribution margin per completed order

Return savings calculation

Gross return-cost savings before app and implementation costs

Avoided returned units × fully loaded cost per return

Net pilot decision

Net economic result for the mature cohort

Incremental contribution + gross return-cost savings − app fee − setup − analytics and review labor

What should Shopify teams do after the 30-day readout?

Roll out only when the primary conversion result clears the pre-registered commercial threshold, technical guardrails hold, and the later mature return cohort shows no offsetting cost. Treatment can drive more purchases while bringing more size bracketing or expectation returns. The order-to-return link exposes that trade-off.

If conversion stays flat while try-on completion is healthy, test placement, call to action, supported product types, or output presentation before rejecting the category. Fix mobile first if photo uploads or generation errors cluster there. The experiment should yield a specific next action, not a foggy verdict on AI virtual try-on.

Human review remains necessary. Check generated outputs against product color, silhouette, material cues, and brand standards, especially on high-traffic hero SKUs. The widget can add useful detail to a Shopify PDP at scale. Weak product inputs, unclear consent, or a broken mobile handoff will still contaminate a perfectly randomized test.

FAQ: How long should a virtual try-on return test run?

Run conversion assignment for 30 days, then wait until every included order has the same post-delivery return-observation window. The maturity period depends on store return policy and delivery timing. Never compare early treatment orders with later control orders that have had less time to come back.

FAQ: Should the test use session-level or visitor-level assignment? Use cookie-level visitor assignment beginning with the first eligible PDP session and retain it across return visits. Session-level reassignment can show one shopper control and treatment, weakening the causal comparison.

FAQ: Can GA4 alone measure return-rate change? No. Use GA4 for PDP and checkout behavior, then join Shopify or order-management return and refund records to the stored experiment assignment. Refunds cannot stand in for returned merchandise, exchanges, or reason-coded return outcomes.