EcommerceAug 10, 2026·Data as of Jun 12, 2026

Data report: AI product image editing for ecommerce—benchmarking pattern direction, texture fidelity, lighting match, and approval rate across AI-only vs. designer-controlled workflows for mattresses and other soft goods

No supplied head-to-head study proves AI-only or designer-controlled editing wins for soft goods. This report provides a blinded benchmark and cost model to find out.

Lamina Team

Lamina Team

Product Team @ Lamina

Designer reviewing AI-edited mattress and bedding product images beside fabric swatches, pattern overlays, and lighting references on a large monitor

What does this AI product-image editing benchmark actually prove?

There is no supplied evidence that AI-only editing or designer-controlled editing earns a higher approval rate for mattresses, bedding, or other soft goods. The defensible result is narrower: give both workflows the same real-SKU inputs, blind the reviewers, and let final approval rate—the business metric—settle the comparison.

That gap matters. A polished lifestyle image can still misrepresent mattress ticking, rotate a stripe, enlarge a repeat, or blur a piping edge; it can look credible and still be unusable for that SKU. Score publishability and named product-preservation failures, then. Do not ask reviewers only whether an image looks realistic.

Use AI generation where it earns its keep: fresh scenes, complex styling, controlled background swaps, on-model concepts, and catalog variants at volume. Keep a designer in the loop when the product itself is the proof—directional quilting, damask, embroidery, ribbing, nap, fringe, seams, gussets, labels, and branded hero imagery. That is how generated output becomes dependable enough to ship.

Evidence that shapes the benchmark design
MetricValueSource
Human-feedback examples used in an inpainting evaluation framework44,000doi.orgas of 2025-04-11
Precision reported for that framework against compared open-source visual-quality models96.4%doi.orgas of 2025-04-11
Kendall’s tau against human judgments for OpenVTON-Bench’s multimodal protocol0.833arxiv.orgas of 2026-01-01
Kendall’s tau against human judgments for SSIM in the same benchmark0.611arxiv.orgas of 2026-01-01
Reported pattern-scale error in a textile mockup workflow2–3×blle.coas of 2026-01-01

Which metrics determine whether an AI-edited soft-good image is publishable?

Use final approval rate to compare workflows. It tells you whether a submitted image can actually clear your product and channel standards: approved deliverables divided by submitted deliverables. Report it by product class and task type; blending easy cleanup jobs into texture-critical hero images will hide the result you need.

Track first-pass approval, attempts per approved image, designer minutes per approved image, rework rate, time to approved asset, and cost per approved asset. Each catches a different leak. A workflow may generate quickly, then burn money while reviewers reject output over and over. Keep human review, revision time, and any media spend out of generation cost unless your calculation explicitly includes them.

Score each output in separate fields: silhouette preservation; pattern direction and repeat scale; texture and material fidelity; lighting match; colorway accuracy; seams, piping, quilting, fringe, labels, and logos; composition and channel compliance; and overall publishability. OpenVTON-Bench separates background consistency, identity fidelity, texture fidelity, shape plausibility, and realism, and reports stronger agreement with human judgments for its multimodal protocol than for SSIM. Use automated signals to triage. They should not be the release gate.

How do you run a fair AI-only versus designer-controlled benchmark?

Run a stratified, blinded experiment. Give both workflows identical real-SKU source files, creative briefs, crops, resolution requirements, and generation or revision limits. Randomize output order and conceal which workflow made each asset; process familiarity can contaminate an approval call.

Build the set around prints and construction that tend to fail: mattresses, bedding, throws, pillows, and upholstered furniture; stripes, checks, directional quilting, damasks, embroidery, ribbing, velvet or other nap, fringe, seams, gussets, labels, and multiple colorways. Include background replacement, lifestyle staging, retouching, and colorway-extension tasks. Passing a plain white pillow says very little about a patterned mattress.

In the AI-only arm, use predefined generation settings and automated quality gates only. Let the designer-controlled arm choose references, protect product areas with masks, adjust controls, reject and regenerate, and make constrained finishing edits. Generic textile generation can invent grain and warp bedding patterns. Keep the supplied product source as the anchor rather than asking the model to rebuild a SKU from a loose visual impression.

Reproducible soft-goods benchmark protocol

  1. Lock the inputs before either workflow begins

    Build a real-SKU source set. Record product dimensions, visible repeat counts, colorways, required crops, final resolution, the creative brief, and the maximum generation or revision budget. Pair every output with its exact source image and brief so reviewers can inspect the product delta.

    Lock the inputs before either workflow begins
  2. Set up two operating arms

    Allow only predefined automated generation and automated gates in the AI-only arm. In the designer-controlled arm, allow reference selection, product protection, control adjustment, rejection and regeneration, and constrained finishing edits. Neither arm gets extra source photography, looser deliverable standards, or a bigger revision budget.

    Set up two operating arms
  3. Measure pattern direction against the product reference

    For every visible patterned area, annotate expected angle, repeat size, motif order, and continuity across seams, folds, and edges. Log absolute direction error, repeat-scale error, repeat-count error over a defined span, registration error, plus any warped, missing, duplicated, or discontinuous motif area. One reported pillow mockup case rendered stripes 2–3 times larger than the real fabric; repeat counts tied to product dimensions catch that defect.

    Measure pattern direction against the product reference
  4. Score texture and lighting on separate lines

    Give reviewers calibrated crops of the actual ticking, weave, knit, pile, quilting, tufting, or upholstery. Have them check local structure, sheen, pile direction, loft, and construction detail, then separately inspect key-light direction, softness, color temperature, contact shadow, highlight behavior, and the product boundary. Inpainting can alter product pixels near a generated mask boundary. Boundary review is mandatory.

    Score texture and lighting on separate lines
  5. Blind-review, publish the rubric, calculate approval economics

    Qualified reviewers should approve, reject, or flag each deliverable without seeing its workflow. Publish the rubric, defect-severity rules, task mix, input set, and scoring examples with the result. Calculate approval rate and cost per approved asset by product class, then use automation scores to send suspicious images to review—not to overrule a reviewer.

    Blind-review, publish the rubric, calculate approval economics

How do you measure pattern direction on mattresses, bedding, and upholstery?

Measure pattern direction against the SKU reference, not a reviewer’s sense that the output looks plausible. Record expected direction in degrees, repeat dimensions or a dimension-relative equivalent, motif order, and continuity where fabric crosses a seam, fold, edge, or gusset.

Direction error is the absolute difference between the expected and rendered angle. Repeat-scale error is the percentage difference between expected and rendered repeat size; repeat-count error compares visible repeats across a defined span, while phase or registration error catches a motif that does not meet correctly at a seam. That turns “the stripe feels wrong” from a vague rejection into a defect record someone can repair.

Texture diagnostics can support this work. They cannot replace it. Research on print analysis treats direction angle as an explicit parameter and uses orientations including 0°, 45°, 90°, and 135°; use those signals as secondary checks. The SKU reference remains the authority on the actual pattern, scale, and construction.

Why should people review texture fidelity and lighting match?

People need to review texture fidelity and lighting match because automated visual assessment has uneven sensitivity across product attributes, and it cannot establish that a plausible textile depiction belongs to the correct SKU. A VLM-based garment-consistency study found stronger sensitivity to color and texture than to shape and line dimensions. Useful for triage; insufficient for independent approval.

Use calibrated reference crops under standardized lighting for ticking, weave, knit, pile, quilting, tufting, and upholstery. Review fiber-facing cues, local scale, sheen, pile direction, loft, seam detail, and repeated-patch artifacts, then compare them with material response in the generated scene. A fictional-throw test warns that believable texture does not establish the correct repeat, fiber, weave, color, dimensions, or other real-SKU attributes.

Treat lighting as a separate criterion. Annotate intended key-light direction, softness, color temperature, shadow direction and length, contact-shadow presence, and expected highlight response; inspect mattress piping and upholstered edges closely. Inpainting analysis warns that edits can change product content near mask boundaries, so the release workflow needs a product-delta check against the supplied base image.

The honest answer: it depends entirely on what you use AI for, how you implement it, and whether the people running the process understand fashion well enough to catch what AI gets wrong.
Kamil CzajaFounder & CEO, GoPackshot

What does a designer-controlled workflow add to AI image generation?

A designer-controlled workflow adds source protection, selective regeneration, and accountable approval decisions while retaining AI generation’s speed for scenes and variants. The designer decides which product regions must stay anchored to the reference, catches defects automated checks miss, and rejects a handsome image when it changes the item for sale.

That control matters most when the source view forces the system to infer hidden construction or material behavior. Fynn Badgley’s warning about a single front-facing garment image applies to soft goods as well: incomplete references create too much guesswork. Supply detail crops and additional views for product-critical surfaces. A weak source image is not a prompt-writing problem.

AI-only workflows still have a clear job. Supplied practitioner guidance assigns them to low-risk, controlled, repeatable assets, with complex textures and hero work routed to people. Test that as a hypothesis in your own blinded dataset; do not market it as a measured approval-rate advantage before the experiment returns results.

A single front-facing photo of a garment isn't enough — the model has to guess too much.
Fynn Badgley
TierPriceIncludedBest for
AI-only benchmark armGeneration cost plus automated QA cost; enter measured valuesLow-risk, controlled, repeatable cleanup and catalog tasks
Designer-controlled benchmark armGeneration cost plus designer review and constrained finishing; enter measured valuesDirectional patterns, texture-critical SKUs, and brand-critical hero assets
Release-cost comparisonCost per approved asset; calculate from your test dataChoosing the workflow by publishable output rather than raw generation cost
No supplied source establishes a market price or a measured cost winner for either workflow. Use this worksheet to compare fully comparable production paths after your blinded test.

AI-only cost per approved asset

Enter measured study values

(AI generation spend + automated QA spend + rework spend) ÷ number of approved AI-only deliverables

Designer-controlled cost per approved asset

Enter measured study values

(AI generation spend + designer review time cost + constrained finishing cost + rework spend) ÷ number of approved designer-controlled deliverables

Workflow decision by product class

Select the workflow that meets the release standard for that class

Compare approval rate, time to approved asset, and cost per approved asset separately for mattresses, bedding, throws, pillows, and upholstery

What is the practical decision for ecommerce teams?

Use AI-only processing where your controlled benchmark shows it hits the same approval standard at lower time or cost per approved asset. Use designer-controlled generation where texture, directionality, construction, or hero status raises rejection risk. Raw image cost, demo realism, and one aggregate score are poor grounds for the call.

Start small, with difficult real SKUs. Publish the input protocol, scoring rubric, product-class breakdown, defect examples, approval denominator, and cost formula beside the results so a merchandising lead, creative director, and procurement team can inspect the finding. If a benchmark cannot show what was rejected—and why—it cannot run production.

The operating split is already clear enough to test: automate repeatable, low-risk work; keep designer control on source-anchored soft goods that must survive close inspection. Generation remains the production engine. Human art direction and approval hold the quality line around it.