Virtual Try-OnPricing guideAug 26, 2026·Data as of Aug 22, 2026

How to test AI virtual try-on for ecommerce fashion in 30 minutes

Run a controlled nine-output virtual try-on benchmark before scaling fashion imagery. Score garment fidelity, model retention, brand compliance, and PDP readiness.

Lamina Team

Lamina Team

Product Team @ Lamina

Fashion ecommerce team reviewing AI virtual try-on images of three garments on the same model beside garment reference images and a scoring rubric

Don’t spend the first session chasing conversion claims. Run a controlled 30-minute creative-quality gate instead: three representative garments, one approved model reference, one art direction, three outputs per garment, plus a hard-fail rubric for product detail and PDP readiness.

Lamina defines Virtual Try-On as garment-on-model generation from shopper photos. The first question for a fashion team is practical: can it turn the garment references merchandising already uses into PDP images that retain the colorway, print, logo, neckline, sleeves, hem, trim, and styling rules?

Keep this benchmark narrow. It tests visual compositing and creative governance, not whether a shopper gets an exact fit. A 2026 CHI study on ecommerce virtual try-on found that shoppers value fit accuracy and reliability signals, including confidence scores; those belong in a later storefront experiment with separate product and measurement work.

Benchmark facts to set before the first run
MetricValueSource
Recommended test window30 minutesuselamina.aias of 2026-08-08
Representative SKU types3uselamina.aias of 2026-08-08
Controlled outputs in the first matrix9uselamina.aias of 2026-08-08
Lamina model routes15+uselamina.ai
Brand-aware output types, including virtual try-on6uselamina.ai
AI assets generated on Lamina (last 30 days)314Lamina platform telemetryas of 2026-08-22
Median time to generate an asset223sLamina platform telemetryas of 2026-08-22
90th-percentile generation time472sLamina platform telemetryas of 2026-08-22

What should a 30-minute virtual try-on test actually prove?

A 30-minute virtual try-on test should show whether Lamina can make publishable, on-brand visual candidates for a defined catalog slice. It does not prove that a garment fits every shopper, predicts size, or reduces returns.

This gate makes merchandising inspect the PDP failure modes before a broad catalog enters the workflow. A beige sweater with warped rib texture fails. So does a dress that renders cleanly yet changes the selected colorway, blurs a logo, drops a fastening, shifts the model’s face, or invents a layer beneath a sheer garment.

Use the score to set eligibility. A strong result authorizes a constrained pilot by category and input standard; a weak one gives you a useful repair list—cleaner garment references, clearer attributes, fewer permitted poses, or exclusions for difficult constructions. Better that than a gallery of pretty, unscored images.

Run the 30-minute on-brand creative benchmark

  1. Minutes 0–5: Pick three SKUs and write the pass/fail rubric

    Choose one control SKU with a simple silhouette, one brand-signature SKU carrying a logo, print, trim, or distinctive construction, and one hard case: dark, loose, layered, or heavily patterned. Score every output from 1 to 5 across five areas: garment and color fidelity; neckline, sleeve, hem, and coverage accuracy; model identity, body, and pose retention; lighting and brand styling; and PDP readiness. A changed logo, illegible print, wrong colorway, missing component, anatomical artifact, or implausible layering is a hard fail. Lamina’s benchmark protocol calls for garment, person, pose, and structured-product-attribute cards before generation.

    Minutes 0–5: Pick three SKUs and write the pass/fail rubric
  2. Minutes 5–10: Assemble one fixed input pack

    For each SKU, pull the cleanest garment reference available, the SKU and colorway identifier, and a short list of what must stay and what must not change. Pick one model reference, one approved pose, and one background. Attach the same brand kit to each run. Lamina defines that kit as a structured palette, typography, voice, do and don’t rules, reference shots, and product-fidelity constraints; use it as a fixed control, not a loose moodboard.

    Minutes 5–10: Assemble one fixed input pack
  3. Minutes 10–20: Generate a controlled matrix of nine outputs

    Generate three outputs for each of the three SKUs. Lock resolution, aspect ratio, model version, reference-image order, inference settings, retry policy, queue region, brand kit, model reference, and pose. Change one predeclared creative variable only, such as a tightly specified styling direction. Lamina’s protocol is right on this: move several settings at once and nobody can tell whether a garment failure came from the prompt, the inputs, or the generation setup.

    Minutes 10–20: Generate a controlled matrix of nine outputs
  4. Minutes 20–27: Blind-score images against the source files

    Have the reviewer score outputs beside the original garment and model references without knowing which styling variant was meant to win. Mark exact defects: print or logo mutation, color drift, silhouette change, missing fastening or trim, sleeve or hem displacement, face or pose drift, hand or hair occlusion, lighting mismatch, and off-brand styling. Don’t score “fit accurate.” An image can preserve a silhouette cue without establishing real-world fit.

    Minutes 20–27: Blind-score images against the source files
  5. Minutes 27–30: Make a bounded pilot call

    Green-light a limited creative pilot only when the control and signature SKUs clear the preset threshold with no hard failure. If the hard case misses, keep the useful result from the other two categories: mark that SKU type excluded for now, then rerun it later with stronger references or a tighter brief. Save the input standards, score sheet, excluded SKU types, and winning settings as the operating spec for the next batch.

    Minutes 27–30: Make a bounded pilot call

Which garments belong in the first virtual try-on benchmark?

Start with one easy control, one unmistakably branded item, and one construction likely to expose errors. That three-SKU mix catches generic visual polish while forcing product fidelity—not aesthetic preference—to decide the result.

Use a solid-color tee, knit, or simple shirt with a clean front-facing reference as the control. It tells you whether the chosen model, pose, lighting, and background work with the brand’s basic catalog language. If it misses color, sleeve placement, or silhouette, quit fiddling with styling prompts and inspect the input pack and fixed conditions.

Your signature item needs the detail a shopper uses to identify the product: a wordmark, repeat print, contrast piping, embroidery, hardware, or an unusual collar. This is the credibility check. A polished virtual try-on image that drops that brand-signature cue can mislead on a PDP even if the rest of the composition looks good.

Use the hard case to draw boundaries early. Loose layers, black-on-black construction, sheer overlays, dense prints, long hair over a shoulder, and complex fastenings need closer inspection. AI generation can still serve these categories, though the approval rule should be stricter and the creative brief more explicit.

What should the rubric score before an image reaches a PDP?

Score garment fidelity first, model and pose retention second, brand compliance third, then publication readiness. A high average cannot save product misrepresentation. Altered branding, a wrong colorway, a missing garment element, or implausible layering means automatic rejection.

Start with the garment beside its source reference. Check hue under approved lighting; logo or print placement and legibility; neckline shape; sleeve length; hem position; closures; pockets; seams; and visible trim. With patterns, inspect the repeat. “Feels similar” is useless here.

Then check the person reference. Face, body, pose, and framing should stay recognizably consistent across all three variants, especially when the images sit together in a campaign or collection. Hair, hands, bags, and outerwear often block the garment; label the defect location so the next brief can prevent it.

Apply the brand kit last. Lamina’s definition covers palette, typography, voice, do/don’t rules, reference shots, and product-fidelity constraints. For virtual try-on, ask one concrete question: does the background, lighting, styling, crop, and visual restraint read as approved ecommerce imagery rather than some unrelated fashion editorial?

Use one scorecard for every output
Review areaPass conditionHard failureAction if it failsSource
Garment and color fidelitySelected SKU, colorway, print, logo, and construction match the referenceWrong colorway, altered logo, illegible print, missing componentReject; improve garment reference and preservation rulesuselamina.aias of 2026-08-08
Silhouette and coverageNeckline, sleeves, hem, layers, and garment coverage remain plausibleChanged silhouette, misplaced hem or sleeve, implausible layeringReject; narrow pose or category eligibilityuselamina.aias of 2026-08-08
Model and pose retentionApproved model reference and pose remain consistentFace, body, or pose drift that breaks the creative systemRevise the fixed reference packuselamina.aias of 2026-08-08
Brand complianceBrand-kit rules, reference-shot cues, lighting, and styling are followedOff-brand styling or visual treatmentRevise the art-direction instructionuselamina.ai
PDP readinessImage is an honest visual composite suitable for normal approvalArtifact or a presentation that implies exact fitReject; retain a clear distinction between imagery and fit guidancedl.acm.orgas of 2026-08-22

How do you keep the experiment fair and reproducible?

Change one intentional creative variable at a time. Fix everything else. Lamina’s benchmark protocol specifically names resolution, aspect ratio, model version, reference-image order, inference settings, retry policy, and queue region as conditions to lock before comparing prompt variants.

That discipline beats generating a giant gallery. Different poses, backgrounds, model references, and prompts across outputs tell a reviewer almost nothing about whether the garment held up. A small matrix with a named variable yields an approval rule creative operations can rerun next week.

Where the implementation exposes a seed, keep it with the brief and source assets for comparable reruns. Use one file-name pattern: SKU, colorway, model reference, pose, variant, score, and failure label. That record turns a 30-minute exercise into a reusable eligibility map, rather than a one-off demo.

Human art direction still earns its keep. The evaluator should closely approve brand-critical hero moments, compare every candidate with the source garment, and reject vague prompts before production. The generator follows the discipline of the brief it gets.

TierPriceIncludedBest for
First benchmarkFree credit allowance for new accountsUse the current allowance to cover a small controlled testA three-SKU, nine-output creative-quality gate
Limited pilotCredit-based usageBudget credits against the selected routed model and approved output countEligible categories, selected colorways, and internal PDP review
Catalog productionCredit-based usagePlan generation credits separately from human review and revisionsRepeatable SKU batches using documented input standards
Lamina uses credit-based usage and provides a free credit allowance for new accounts. Confirm the active credit rate for the selected model route before running a production batch.

Initial creative-quality gate

9 output credits at the current applicable rate

3 SKUs × 3 controlled outputs = 9 generated candidates; 9 × current credit rate for the selected route

Approved-category pilot

75 output credits at the current applicable rate

25 eligible SKUs × 3 candidates each = 75 generated candidates; 75 × current credit rate for the selected route

Scale after selecting one approved candidate per SKU

300 output credits at the current applicable rate, excluding review and revisions

100 SKUs × 3 candidates = 300 generated candidates, plus human review of the candidates; 300 × current credit rate for the selected route

How should ecommerce teams read generation time and cost?

Treat generation credits and generation time as an iteration budget, not the cost of a published asset. Lamina routes across more than 15 image, video, and try-on models, so check the active model route and its credit rate before approving a pilot.

Lamina telemetry reports a median asset generation time of 223 seconds and a 90th-percentile time of 472 seconds. Generate the controlled matrix early in the 30-minute window, then leave room for side-by-side scoring. Edge cases will not all finish together.

A credit calculation leaves out the work that decides whether an image can publish: selecting references, writing preservation rules, reviewing defects, revising a failed brief, and approving the final PDP candidate. Keep that work in the budget. It is the quality control that holds the asset on brand.

When should a team move from a creative benchmark to a storefront test?

Move to storefront testing after the creative benchmark establishes which categories and inputs pass reliably. Then measure assisted and control sessions by category across a full return cycle. Clicks and engagement alone don’t settle it.

Kleep recommends separating assisted from control sessions and evaluating virtual try-on by category over the return cycle. That avoids treating an entertaining try-on experience as a proven commercial result, particularly where fit expectations vary across knitwear, dresses, outerwear, and structured garments.

Carry the benchmark’s failure taxonomy into the storefront test. If shoppers see virtual try-on for a printed blouse, product teams should trace complaints to a scored criterion—color drift, print change, silhouette mismatch, or an expectation problem—instead of falling back on “the image looked good.”

What is the practical decision after 30 minutes?

After 30 minutes, approve a limited Lamina virtual try-on pilot for the SKU categories that passed, or narrow the brief and rerun the failures. Nine images are not a catalog-wide commitment.

A green-light needs clean passes on the control and signature SKUs, zero hard product-fidelity failures, and an approved record of the model reference, pose, brand kit, source-image standard, and prompt variant. The hard case can stay excluded while proven categories move ahead. That is a sensible operating boundary.

If the signature SKU fails, improve the garment reference and preservation instructions before raising volume. If the control fails, revisit the fixed setup—reference order, pose, background, and generation conditions—before changing the creative idea. Find the cause. Don’t keep rerolling until one output happens to look acceptable.

FAQ: Can AI virtual try-on prove garment fit?

No. AI virtual try-on can produce a garment-on-model visual composite and retain useful silhouette cues, yet it does not establish exact shopper fit. The CHI research identifies fit accuracy and reliability cues as shopper needs; use clear product information and later assisted-versus-control measurement instead of treating a generated image as sizing evidence.

FAQ: How many virtual try-on images should the first test generate?

Generate nine candidates: three representative SKUs with three controlled attempts apiece. That is enough to compare one deliberate creative variable while keeping the review small enough for a single working session.

FAQ: What is the most important virtual try-on failure to reject?

Reject any output that changes what a shopper believes they are buying: a wrong colorway, mutated or illegible logo or print, missing component, changed silhouette, or implausible layer. A polished background cannot make a misleading garment acceptable.

FAQ: Should the first test use every fashion category?

No. Begin with the three-SKU mix, then expand only from categories that meet the rubric. Log difficult constructions and weak source imagery as exclusions; revisit them with a more specific brief and stronger references.