Product PhotographyPricing guideAug 25, 2026·Data as of Feb 15, 2026

Use 3D models to lock Stable Diffusion product scenes

Keep SKU geometry, logos, and labels intact by treating the 3D render as the product truth and using Stable Diffusion for controlled scene generation.

Lamina Team

Lamina Team

Product Team @ Lamina

A 3D-rendered skincare bottle locked in place while AI-generated ecommerce lifestyle scenes and vertical video frames surround it

For ecommerce, keep the approved 3D product model as the source of truth and build the scene around its rendered pixels. Do not hand Stable Diffusion a loose reference and ask it to redraw a sellable SKU when the 3D asset already preserves the exact logo, label, silhouette, materials, and camera view.

That shifts the work. A GLB or USD model provides the approved product transform, lens, materials, decals, alpha matte, depth, and normal data; Stable Diffusion takes on the variable parts—surfaces, props, environments, transition zones, reflections, contact shadows, and restrained relighting. You get more usable PDP variants, campaign stills, and short vertical reels this way than by repeatedly asking a model to remember which side of a bottle has the label or how many buttons a garment has.

What the control signals actually protect
MetricValueSource
Img2img denoising at 0Adds no noise to the inputstable-diffusion-art.comas of 2025-11-11
Img2img denoising at 1Fully replaces the input with noisestable-diffusion-art.comas of 2025-11-11
Black regions in a background-replacement maskPreserve the productagentbus.shas of 2026-02-15
White regions in a background-replacement maskGenerate the new sceneagentbus.shas of 2026-02-15
Depth ControlNet inputEnforces spatial relationships and product-to-background separationdocs.comfy.org
Combined IP-Adapter and ControlNet workflowUses an image reference for art direction plus a structural signal such as depth or Canny edgesdeepwiki.comas of 2025-04-19

Why lock the 3D model as the subject?

Lock the 3D model because it renders the exact approved SKU before generation starts. In a practical 3D-to-photo setup, the operator loads a.glb, sets its position and orientation, then uses Stable Diffusion 1.5 inpainting to construct the virtual product-photo scene around it.

This is a production distinction, not theory. One-step AI viewpoint changes can drift a logo, distort proportions, change feature counts, or invent back-view details: a black sneaker gets a new outsole pattern, a cosmetics carton loses a regulatory line, a sofa acquires an arm profile that no longer matches its listing. Stage the re-rendering workflow so fidelity and styling stay separate; then a bad environment can be rerun without disturbing the approved product.

Render the product at its intended final framing whenever you can. Need a three-quarter campaign hero? Set that camera in Blender or your 3D package first instead of generating a front-on packshot and asking diffusion to rotate it. The render becomes a controlled subject plate, not a suggestion for a fresh image.

Which render passes belong with every SKU?

For every approved camera view, include beauty RGB, alpha or matte, shadow catcher, depth, and normal passes. The RGB plate holds the product pixels. Alpha isolates them; the shadow catcher gives you a credible grounding start, depth fixes placement, and normal data maps surface direction for controlled finishing.

Tie each output package to a specific SKU, colorway, camera, and material revision. Your structured truth pack should also log protected decals and labels, the lens and framing rule, permitted crops, approved reference images, palette, acceptable props, prohibited claims, and positive and negative prompt templates. An oxblood handbag and that same handbag in sand are separate production assets, not nearly identical folders.

Keep brand styling instructions apart from product-fidelity rules. Product rules are binary: preserve the wordmark, button count, label copy, silhouette, and colorway. Styling can move—limestone pedestal, sharp morning sidelight, warm cream wall, restrained botanical prop—so an operator can fix a scene mask or styling prompt without reopening product approval.

Ever struggled to maintain visual consistency across multiple product images? Here’s a game-changer: using ControlNet and IP-Adapters in Stable Diffusion, you can take one product photo and generate it in 50 different lifestyles while keeping the product dimensions 100% accurate.
Aryan NateSolutions Architect Executive, Netweb Technologies India Ltd.

How should ControlNet and IP-Adapter split the job?

Put geometry in ControlNet and approved visual direction in IP-Adapter. Depth ControlNet uses a depth map as its structural condition, making it useful for product-showcase layouts where the object must stay separated from its background at a defined spatial position.

IP-Adapter is an image-prompt adapter for pretrained text-to-image diffusion models. Feed it campaign references that show the visual language you need: a specific studio palette, tabletop surface, prop density, grading, or lighting character. It works alongside text prompts and existing controllable-generation tools, directing atmosphere while a depth or edge map stops the scene from disregarding the product’s physical placement.

Use both only when each supplies a separate constraint. A depth pass can keep a bottle placed on its plinth; an approved campaign still can call for cool brushed steel, hard lateral light, and sparse composition. Leave the product masked, and neither control has to perform the impossible job of reproducing tiny type or a proprietary contour.

How to generate locked-SKU ecommerce stills and reels

  1. Approve the 3D source and shot list

    Begin with the approved GLB or USD asset, including correct materials, dimensions, decals, and colorway. Lock product transforms, camera position, focal length, crop, and brand lighting for every PDP, campaign, or social shot. Name outputs by SKU, view, and revision; a later material change should never quietly overwrite an approved scene.

    Approve the 3D source and shot list
  2. Render a product truth pack

    Export beauty RGB, alpha matte, shadow catcher, depth, and normal passes at the delivery aspect ratio. Add a Canny or line-art pass for categories with unmistakable edges, including eyewear, footwear, or hardware. Keep the render and its metadata. That is what lets you replace a rejected background without generating the product again.

    Render a product truth pack
  3. Write the brand pack before prompting

    Set approved reference images, palette, surfaces, props, camera rules, positive prompt language, negative prompt language, and prohibited product claims. State protection rules literally: do not change logo, label text, closure count, stitching, colorway, or silhouette. “Keep it accurate” is weak; a concrete SKU checklist holds.

    Write the brand pack before prompting
  4. Generate the environment on its own

    First generate a background plate, or composite the untouched product over a generated scene. For background replacement, make an inpainting mask where the preserved product is black and the generated scene is white. Stable Diffusion gets room to build the set while the product body stays out of redraw.

    Generate the environment on its own
  5. Add structural and style conditioning only where it earns its place

    Send the product depth render through Depth ControlNet to preserve placement and separation. Use IP-Adapter with approved campaign imagery to steer art direction. If you need one unified finishing pass, keep img2img denoising low and limit higher-change work to masked background or transition areas; more denoising means bigger changes.

    Add structural and style conditioning only where it earns its place
  6. Batch stills with settings you can repeat

    For each scene family, hold the prompt template, reference set, control inputs, seed policy, resolution, and workflow graph steady. Change one styling variable per run: set surface, prop, lighting mood, or crop. Save every accepted seed and conditioning setup—merchandising will ask for that same image in a new colorway.

    Batch stills with settings you can repeat
  7. Build reels from deterministic 3D motion

    Animate the camera, object, or lighting in 3D, then export depth and normal sequences frame by frame. A ComfyUI Blender integration can load Blender EXR depth and normal sequences as ControlNet conditioning for temporal video generation. Keep prompt and conditioning rules identical across the sequence, then inspect for flicker, drifting reflections, and product details that change.

    Build reels from deterministic 3D motion
  8. Run visual QA at SKU level before publishing

    Reject any output where the logo, label text, feature count, colorway, silhouette, or product-to-surface contact has shifted. Check shadows and reflections as well. A product floating one millimeter above a stone plinth does not belong on a PDP; fix the mask edge, contact shadow, or background plate before regenerating the whole composition.

    Run visual QA at SKU level before publishing

What denoising strength fits a locked product?

Use the lowest denoising strength that delivers the requested finishing change, and keep the approved product body outside the redraw region whenever possible. Img2img denoising runs from 0 to 1: 0 adds no noise, 1 fully replaces the input with noise, and higher settings create larger departures from the original.

Treat denoising as a risk dial. Low-denoise finishing can blend a reflection, soften a transition edge, or bring a scene’s color cast into a protected product plate. It does not authorize regenerating label typography, packaging claims, stitching, or a recognizable physical feature. Put major changes—new foliage, wall texture, table material, broad atmospheric lighting—inside the white inpainting region.

There is one useful exception: concept exploration before SKU approval. At that point, high-change generation can test surfaces, lighting worlds, and shot directions. Once an actual product variant reaches merchandising, go back to render-first and lock the subject.

How should you QA a generated product image or reel?

Compare every published frame against the SKU truth pack, not against whether the image simply looks convincing. Check logos and label text at useful zoom; count buttons, seams, closures, ports, stitches, and other category-specific features; verify the exact colorway; then match the outer silhouette against the approved render.

Then check physical integration. The contact point must meet the surface, cast shadows must run in a plausible direction, glossy materials should reflect the environment without inventing a second object, and background occlusion cannot cut into the protected asset. For a reel, scrub frame by frame for wordmark flicker, moving highlights, repeated props, and changes in product scale.

Use a repair-first rule. A failed background, matte edge, shadow, or reflection is usually a local production issue, so rerun that stage or composite a correction instead of regenerating a full hero image. Separate standardization, an executable brief, and final composition in the workflow, and selective reruns remain possible.

TierPriceIncludedBest for
Still-image proof of conceptVaries by compute providerNot applicableTesting one approved SKU, one camera view, and several background directions
SKU scene productionVaries by compute providerNot applicableBatching approved stills with saved seeds, masks, and ControlNet inputs
Short reel productionVaries by compute providerNot applicableFrame-sequence rendering, temporal conditioning, review, and edit
Stable Diffusion is commonly deployed as a self-hosted workflow, so total cost depends on the chosen GPU provider, runtime, storage, and the amount of human review. Build estimates from the actual workflow rather than treating a generated frame as a fixed-price published asset.

One ecommerce still from an existing approved 3D camera view

Varies by GPU provider, workflow settings, and review effort

3D render time + diffusion background runtime + any local inpainting runtime + human QA time

A five-shot product reel using deterministic 3D camera motion

Varies by frame count, resolution, GPU provider, and revision rounds

Per-frame 3D pass generation + temporal diffusion runtime across every frame + flicker review + final edit

What does this workflow cost in practice?

Production cost is compute plus review, not a universal per-image fee. A still needs a 3D product render, scene-generation or inpainting runtime, possible local repairs, and SKU QA; a reel adds per-frame passes, temporal conditioning, flicker review, and editing. GPU pricing, model choice, resolution, batch size, and revision count set the compute portion.

Estimate at the shot-family level. One family might cover a SKU in a fixed three-quarter camera over six seasonal surfaces, or a five-shot vertical reel with camera motion already authored in 3D. Reuse the render pack, mask, depth pass, normal sequence, and approved style pack across that family so you are not paying to solve product fidelity again and again.

Keep human review as a real line item. Generation can produce new concepts, complex styling, on-model contexts, and believable material detail with far less turnaround than a conventional production cycle. A brand-critical hero still requires truth-pack approval before it goes to a marketplace, PDP, or paid placement.

What should you do if the output still changes the SKU?

If Stable Diffusion changes a SKU detail, stop pushing prompt emphasis and shrink the area it can alter. Replace the product plate with the original RGB render, tighten the alpha matte, put the changed region back under a protected mask, or rerun only the background and transition zone.

A changed logo, missing seam, added switch, or altered label is not a minor style defect. It shows that probabilistic generation was given responsibility for an exact product fact. Put that responsibility back in the 3D render, then let ControlNet, IP-Adapter, and the prompt vary the environment safely.