Brand & Creative OpsAug 5, 2026·Data as of May 27, 2026

Data report: Do AI-generated ecommerce product ads feel less authentic than scripted human ads? A blind-test framework for measuring emotional connection, trust, and purchase intent across on-brand product video and image campaigns.

A blind-test framework for measuring whether AI-generated ecommerce ads change authenticity, trust, emotional connection, and purchase intent.

Lamina Team

Lamina Team

Product Team @ Lamina

Creative team reviewing matched AI-generated and human-created ecommerce product ads on a testing dashboard

Do AI-generated ecommerce ads come off as less authentic than scripted human ads?

Yes. AI-generated ads carry a real authenticity risk, though it is conditional rather than automatic. The strongest evidence comes largely from general advertising rather than ecommerce specifically, and it draws an awkward but useful line: viewers may struggle to identify how an ad was made while still scoring an ad they believe was human-made more favorably.

“AI versus human” is the wrong launch question. Test the exact product, offer, brand codes, channel, and audience that will get the campaign, then pull apart what the asset actually is from what viewers believe it is. A polished product clip may clear a quick source check and still bleed trust when the styling, claims, or emotional cue lands outside the brand.

AI generation is well suited to the matched concepts, involved styling, on-model variants, and material detail this experiment needs. The hard part is the brief and the review. Lock product facts, art direction, claim language, and the approval bar before you compare executions.

What the current evidence means for an ecommerce blind test
MetricValueSource
Average accuracy at identifying whether ad images were AI- or human-made64%faculty.haas.berkeley.eduas of 2025-06-29
Ipsos performance result for matched human-made ads versus its benchmark11 points above benchmarkphys.orgas of 2026-05-27
Ipsos performance result for matched AI-made ads versus its benchmark5 points below benchmarkphys.orgas of 2026-05-27
Ad-day observations in the large-scale display-image studyMore than 2 millioncolumbia.eduas of 2025-01-14
Trust score reported for human-made versus AI-made marketing imagery5.63 vs. 4.24doi.orgas of 2026-04-29
Median time to generate an asset207sLamina platform telemetryas of 2026-08-04

What should a blind test track beyond click-through rate?

A useful ecommerce blind test tracks emotional connection, authenticity, trust, product realism, brand fit, and purchase intent, then puts those responses beside observed choice behavior. Click-through by itself cannot show whether an asset feels genuine or whether a shopper believes the product depiction; large-scale display-ad evidence found a click-through edge only when AI images did not appear AI-generated.

Use a seven-point scale for every diagnostic: “This ad feels genuine for this brand,” “This ad makes me feel something,” “I trust the product portrayal and claims,” and “I would consider buying this product.” Add a forced-choice preference question, an “AI or human?” attribution question, and confidence in that attribution. Keep those last two. Perceived origin can affect performance independently of actual origin.

Put the same creative in an experimental storefront where you can. Under controlled traffic allocation, record product-page clicks, time on page, add-to-cart, and completed checkout intent; these are behavioral proxies, not proof of long-term sales. Human review, revisions, and media spend sit outside asset-generation cost, and they still belong in the operating plan.

Why test perceived source separately from actual source?

Test perceived source separately because believing an ad was human-made can change a viewer’s ratings even if AI actually generated it. Berkeley researchers found that ads perceived as human-made scored higher across effectiveness measures, while participants identified image origin only modestly above chance.

That finding changes the experiment. A clean blind wave estimates the effect of the actual production method without priming respondents; a separate disclosure wave shows what changes once the label is visible. Do not roll those conditions into one average. An East Tennessee State University experiment found source differences when participants correctly identified the source, while disclosure alone did not show strong effects.

Category matters as well. Research suggests functional and hedonic products can draw different interest responses, so do not toss a replacement filter and a fragrance launch into one pool, then call the outcome a creative verdict. Split results by product type, purchase risk, existing versus new customer, brand familiarity, and image versus video.

How to run a blinded ecommerce ad authenticity test

  1. Write a production-neutral creative brief

    Set the SKU, offer, audience, placement, duration or aspect ratio, mandatory product details, brand palette, prohibited claims, and success threshold. Hand the same brief to the human creative route and the AI-generation route; production method should be the variable, not strategy. For generated work, keep prompts, input imagery, versions, human edits, and approvals.

    Write a production-neutral creative brief
  2. Make matched image and video executions

    For each planned concept, produce at least one human-created execution and one AI-generated execution using the same product facts and message. Get production value as close as practical. AI can turn out multiple on-brand candidates fast, though the final pick needs the same brand and product-accuracy review you would apply to every ad.

    Make matched image and video executions
  3. Randomize a blinded audience wave

    Give each respondent only one execution for a given concept, and hide how it was made. Randomize asset order across concepts, hold audience eligibility steady, and keep participants from seeing alternate versions. Measure the seven-point diagnostics, forced choice, source attribution, attribution confidence, and experimental-storefront actions.

    Randomize a blinded audience wave
  4. Run a separate disclosure wave

    Repeat the randomized test with clear “AI-generated,” “human-created,” or “AI-assisted” labels wherever those labels accurately describe the work. This separates the disclosure effect from the visual execution. Give AI-assisted its own condition: research on social content and emotional marketing found less negative reaction where AI supported human creation rather than fully replacing it.

    Run a separate disclosure wave
  5. Decide on practical thresholds, not one p-value

    Before fielding, pre-register the minimum acceptable AI-minus-human difference for authenticity, emotional connection, trust, and purchase intent. Report confidence intervals, then cut results by format, category, risk, customer status, and source-attribution accuracy. Scale only executions that clear your statistical rule and your business-materiality rule; one test is evidence for that brief and audience, never a universal guarantee.

    Decide on practical thresholds, not one p-value

What belongs in a defensible scorecard for AI-generated ecommerce product ads?

A defensible scorecard shows the AI-minus-human difference for every core outcome alongside actual source, perceived source, and disclosure condition. That structure stops a good click result from masking a trust deficit. It also stops a weak survey result from hiding a segment where the creative performs.

Use one row for each product category and format, including sample size, field dates, audience definition, brief identifier, and the exact asset versions tested. For every outcome, show the human mean, AI mean, difference, uncertainty interval, and pre-registered pass/fail threshold. Keep raw respondent data and asset audit records ready for internal review.

For governance, name who approves product realism, who approves disclosure treatment, and who owns exceptions. The IAB’s transparency framework flags consumer mistrust, inconsistent operations, and regulatory exposure as AI-advertising risks. Treat it as a policy reference, not evidence that any particular disclosure label improves campaign performance.

What does existing ad-testing evidence show about emotional connection and sales?

Existing evidence says AI work can be hard to spot yet still lag on some brand and sales-predictive measures. Ecommerce teams should test the response, not grade the pixels. Ipsos reported a gap between matched pre-2021 human ads and AI counterparts built from the same strategic brief, while separate experiments linked AI use or perceived AI authorship with lower authenticity and less favorable follower, loyalty, or word-of-mouth responses.

Generated product advertising can earn trust. Human art direction and approval still belong in the production system, especially for a brand-critical hero asset. A tight brief, believable product texture, accurate claims, and a deliberate disclosure policy give a generated execution its best shot at meeting brand expectations.

Why does the ad Turing-test perspective matter?

The Turing-test perspective matters because experienced ad reviewers can judge whether an execution works without being told how it was produced. That is the right approach for a blind wave. Measure the ad first, then measure what source assumptions do to the result.

We put together a jury with probably about 300 years of combined advertising experience – people who have seen many thousands of ads
Noah Brierco-organiser of the ad Turing test, BrXnd

What standard should an ad-scoring framework use?

An ad-scoring framework should judge the audience response to the creative while retaining the conditions needed to explain that response. Put source belief, disclosure, and brand fit beside the headline commercial metrics, or the score leaves you guessing why a result moved.

outperforming the norm and pretty damn good
John Kearonquoted by New Scientist regarding the ad-scoring framework, System1 Group

What should ecommerce teams do in practice?

Treat authenticity as a campaign outcome you can measure, not an assumed AI flaw or an assumed virtue of human production. Build matched, on-brand ads. Blind the first audience wave, disclose source in the second, then use the segment-level evidence to decide which concepts get media budget.

Start with a small set of high-volume or strategically important SKUs, not a mixed catalog. If AI-generated product imagery or video matches the human control on trust, authenticity, and purchase intent while passing product-accuracy review, you have evidence to scale that workflow under the tested conditions. If it misses, inspect the brief, styling, source cues, category, and disclosure treatment before you change the production approach.

Make this a repeatable creative-ops practice. Every new product category, format, and audience can shift the outcome, so retain the test design and benchmark future launches against the same scorecard.