Brand & Creative OpsHow-toAug 8, 2026·Data as of May 18, 2026

Do AI ecommerce ads feel less authentic?

A practical blind-test framework for measuring whether AI ecommerce ads affect authenticity, trust, emotional connection, and purchase intent.

Lamina Team

Lamina Team

Product Team @ Lamina

Creative team reviewing matched AI-generated and human-made ecommerce product ads on a split-screen dashboard with survey results

Do AI-generated ecommerce ads feel less authentic than scripted ads made by people?

AI-generated ecommerce ads are not inherently less authentic. They do tend to lose ground once viewers spot AI, see an AI disclosure, or get thin human storytelling. The question is not whether a model or a production crew made the file; it is whether the finished asset earns trust, depicts the SKU truthfully, and pushes a buyer toward a qualified purchase.

Ipsos found that people often could not tell where an ad came from, though human-made work still outperformed on emotional and business measures in a controlled comparison. Harvard Business School researchers saw a different result with display images: AI-generated images could win clicks when consumers did not perceive them as AI-made. That difference is real. A click may be curiosity; trust and purchase intent tell you whether the creative is building a sale you actually want.

Keep production origin, perceived origin, and AI disclosure separate. A blind test judges the asset. A labeled test captures the audience response to being told how it was made. Both count, and each answers a different buying question.

What does current research say about AI ad authenticity and effectiveness?
MetricValueSource
Ipsos study sample: 20 ads shown to U.S. consumers; the study found AI creative lagged human creative on emotional engagement and business-outcome measures, so brands should test outcomes beyond recognition or visual plausibility.20 ads; 3,000 U.S. consumersipsos.comas of 2026-05-18
In the Ipsos and Syracuse comparison, viewers were no more likely to confidently identify AI ads than human ads as AI-made. Recognition alone is therefore a poor proxy for ad effectiveness.13%eurekalert.orgas of 2026-05-18
Matched human ads exceeded Ipsos’s sales-effectiveness benchmark, while AI versions fell below it. The result supports using a shared brief and a pre-set commercial endpoint.Human: +11 points; AI: -5 pointseurekalert.orgas of 2026-05-18
Harvard Business School researchers analyzed display-ad behavior at substantial scale and found that AI images can outperform on click-through rate when they do not appear AI-generated. Use post-click quality metrics to determine whether those clicks are valuable.More than 16 billion impressions; 116 million clicks; 4,633 matched sibling adshbs.eduas of 2025-01-14
A Berkeley working paper found that people identified an ad’s origin only modestly better than chance. Text, logos, and people can reveal perceived authorship, so inspect those elements before testing.64% average origin-identification accuracyfaculty.haas.berkeley.eduas of 2025-06-29

Why should brands separate creative quality from AI disclosure?

Separate creative quality from AI disclosure because a strong blind result does not mean a labeled AI campaign will draw the same reaction. Two experiments involving disclosed AI found that disclosure lowered perceived authenticity; that drop explained negative effects on purchase intention, brand image, and brand trust. Treat disclosure as a message effect you can test, not as fine print.

The split is clearest in categories where effort and lived experience sit inside the promise. Luxury-advertising research found that disclosed AI imagery could seem lower effort and less authentic, though more creative AI imagery reduced that penalty. A jewellery-advertising study found much the same: hidden AI authorship could earn favorable ratings, while human-created work was preferred overall after authorship was disclosed.

Use the blind cell to assess execution. Use the labeled cell to decide whether, how, and where to explain AI involvement to that audience.

What does an expert say about evaluating AI ad authenticity?

NIQ’s Ramon Melgarejo says brands need evidence-led creative evaluation because consumers respond to authenticity consciously and unconsciously. So do not let a team meeting decide whether an asset looks convincing. Pair direct survey responses with behavioral outcomes.

Brands and agencies are innovating at a rapid pace, leveraging AI-generated content in their advertising. They need to be cautious, as our study reveals that consumers are quite sensitive to the authenticity of ad creatives, both at the implicit (nonconscious) and explicit (conscious) levels. Brands must prioritize insights-led creative evaluation to produce effective ads.
Ramon MelgarejoPresident, Strategic Analytics & Insights, NIQ

Why measure emotional connection alongside stated purchase intent?

Measure emotional connection alongside stated purchase intent. People do not assess advertising through conscious reasoning alone, and University of Toronto psychology professor Paul Bloom’s observation warns against treating one rational survey question as the full buyer response.

For video, Kantar recommends validated questions on branding, enjoyment, and persuasion alongside facial coding that records emotional response second by second. You do not need facial coding on every test. You do need some way to catch a trust drop at the product reveal, a human expression, or the final brand frame when the overall score papers over it.

Psychologists, philosophers, neuroscientists often argue that we’re prisoners of the emotions, that we’re fundamentally and profoundly irrational, and that reason plays very little role in our everyday lives.
Paul BloomProfessor of Psychology, University of Toronto

How do you run a valid AI-versus-human ecommerce ad test?

Hold the product proposition, audience, placement, offer, and landing page fixed; change only the production condition. The MADE framework for generative-AI experiments requires you to map and control differences beyond the intended experimental factor. Change the offer, shorten the AI edit, or give it a more generous media audience, and the result tells you nothing useful about production.

Create separate image and video cells. A static PDP image and a 15-second paid-social video make viewers process different cues; lumping them together leaves you with an answer too mushy to guide production.

How can you build a 2 × 2 blind-test framework?

  1. Write one locked creative brief for each angle.

    Spell out the SKU, product facts, substantiated claims, target audience, offer, aspect ratio, placement, duration, CTA, landing page, and brand assets. Produce a human-scripted, human-produced execution and an AI-generated execution from that same brief. Log prompts, model settings, seeds, retouching, and human edits.

    Write one locked creative brief for each angle.
  2. Run an unlabeled, source-blind cell.

    Randomly show each respondent one matched asset without naming its origin. Ask about emotional response, perceived authenticity, product accuracy, credibility, brand linkage, clarity, trust, purchase consideration, and suspected source, including confidence. Randomize asset order. Where possible, keep a respondent from seeing both versions of the same concept.

    Run an unlabeled, source-blind cell.
  3. Add a source-by-label module.

    Test four versions of the same creative: a human-made asset with no label, an AI-generated asset with no label, a human-made asset labeled by origin, and an AI-generated asset labeled by origin. That setup separates the effect of the visual execution from the effect of telling viewers AI was involved.

    Add a source-by-label module.
  4. Check product truth before fielding.

    Review every asset at mobile size and, for video, frame by frame. Flag wrong product geometry, changed colors, unreadable packaging, implausible human anatomy or motion, mismatched shadows, and claims the product page does not support. Pull broken assets before they poison the results.

    Check product truth before fielding.
  5. Validate survey findings in balanced media.

    Run live A/B or incrementality cells with equal audiences, bid strategy, budget, frequency caps, campaign dates, landing page, and measurement windows. Track qualified clicks, PDP engagement, add-to-cart, conversion, revenue or ROAS, returns, cancellations, and sentiment. CTR by itself cannot separate purchase interest from curiosity.

    Validate survey findings in balanced media.

Which measures should decide whether an AI ad is ready to scale?

Scale an AI ad only if product accuracy holds and it clears the pre-set thresholds for authenticity, trust, and commercial quality against its matched human control. Pick one primary outcome before responses come in: purchase intent in the survey, for example, or conversion rate in media. Use emotional connection, trust, and source suspicion as diagnostics for why it won or lost.

Cut results by category, product risk, AI familiarity, technology trust, and whether respondents believed the asset was AI-made. Berkeley researchers found that perceived human origin improved effectiveness scores, including for ads that were actually AI-generated. The audience’s read can matter more than the production record.

For a usable non-inferiority rule, a brand could require an AI asset to sit no more than 3 percentage points below the human control on both trust and purchase intent, with no increase in product-accuracy failures. That is an example operating threshold, not a universal research benchmark. Set the margin before fieldwork, based on category commercial risk and the test’s sample size.

What is the practical decision for ecommerce teams?

Start AI on product-forward variants, localized creative, background changes, and high-volume campaign adaptation. Then prove each winner with matched testing. Ipsos found AI performed better on straightforward product briefs than on work relying on storytelling, emotion, or a distinctive point of view. That is where I would put production capacity first, not a ban on ambitious AI creative.

Review luxury, health and beauty, high-consideration products, testimonials, and founder- or creator-led ads more closely. Buyers in those cases may treat the effort and human experience behind the message as part of the brand’s credibility. AI can produce the creative. A human art director should write the brief, verify product truth, and approve what customers see.

An AI asset that wins CTR while losing trust, purchase intent, conversion quality, or return-rate performance has not won. It got attention without enough confidence to justify scale. Keep the asset log, rerun the test on the next controlled variant, and let the evidence place AI production where it belongs in the campaign.