Use 3D models to lock Stable Diffusion product scenes
Keep SKU geometry, logos, and labels intact by treating the 3D render as the product truth and using Stable Diffusion for controlled scene generation.

Lamina Team
Product Team @ Lamina

For ecommerce, keep the approved 3D product model as the source of truth and build the scene around its rendered pixels. Do not hand Stable Diffusion a loose reference and ask it to redraw a sellable SKU when the 3D asset already preserves the exact logo, label, silhouette, materials, and camera view.
That shifts the work. A GLB or USD model provides the approved product transform, lens, materials, decals, alpha matte, depth, and normal data; Stable Diffusion takes on the variable parts—surfaces, props, environments, transition zones, reflections, contact shadows, and restrained relighting. You get more usable PDP variants, campaign stills, and short vertical reels this way than by repeatedly asking a model to remember which side of a bottle has the label or how many buttons a garment has.
| Metric | Value | Source |
|---|---|---|
| Img2img denoising at 0 | Adds no noise to the input | stable-diffusion-art.comas of 2025-11-11 |
| Img2img denoising at 1 | Fully replaces the input with noise | stable-diffusion-art.comas of 2025-11-11 |
| Black regions in a background-replacement mask | Preserve the product | agentbus.shas of 2026-02-15 |
| White regions in a background-replacement mask | Generate the new scene | agentbus.shas of 2026-02-15 |
| Depth ControlNet input | Enforces spatial relationships and product-to-background separation | docs.comfy.org |
| Combined IP-Adapter and ControlNet workflow | Uses an image reference for art direction plus a structural signal such as depth or Canny edges | deepwiki.comas of 2025-04-19 |
Why lock the 3D model as the subject?
Lock the 3D model because it renders the exact approved SKU before generation starts. In a practical 3D-to-photo setup, the operator loads a.glb, sets its position and orientation, then uses Stable Diffusion 1.5 inpainting to construct the virtual product-photo scene around it.
This is a production distinction, not theory. One-step AI viewpoint changes can drift a logo, distort proportions, change feature counts, or invent back-view details: a black sneaker gets a new outsole pattern, a cosmetics carton loses a regulatory line, a sofa acquires an arm profile that no longer matches its listing. Stage the re-rendering workflow so fidelity and styling stay separate; then a bad environment can be rerun without disturbing the approved product.
Render the product at its intended final framing whenever you can. Need a three-quarter campaign hero? Set that camera in Blender or your 3D package first instead of generating a front-on packshot and asking diffusion to rotate it. The render becomes a controlled subject plate, not a suggestion for a fresh image.
Which render passes belong with every SKU?
For every approved camera view, include beauty RGB, alpha or matte, shadow catcher, depth, and normal passes. The RGB plate holds the product pixels. Alpha isolates them; the shadow catcher gives you a credible grounding start, depth fixes placement, and normal data maps surface direction for controlled finishing.
Tie each output package to a specific SKU, colorway, camera, and material revision. Your structured truth pack should also log protected decals and labels, the lens and framing rule, permitted crops, approved reference images, palette, acceptable props, prohibited claims, and positive and negative prompt templates. An oxblood handbag and that same handbag in sand are separate production assets, not nearly identical folders.
Keep brand styling instructions apart from product-fidelity rules. Product rules are binary: preserve the wordmark, button count, label copy, silhouette, and colorway. Styling can move—limestone pedestal, sharp morning sidelight, warm cream wall, restrained botanical prop—so an operator can fix a scene mask or styling prompt without reopening product approval.
Ever struggled to maintain visual consistency across multiple product images? Here’s a game-changer: using ControlNet and IP-Adapters in Stable Diffusion, you can take one product photo and generate it in 50 different lifestyles while keeping the product dimensions 100% accurate.
How should ControlNet and IP-Adapter split the job?
Put geometry in ControlNet and approved visual direction in IP-Adapter. Depth ControlNet uses a depth map as its structural condition, making it useful for product-showcase layouts where the object must stay separated from its background at a defined spatial position.
IP-Adapter is an image-prompt adapter for pretrained text-to-image diffusion models. Feed it campaign references that show the visual language you need: a specific studio palette, tabletop surface, prop density, grading, or lighting character. It works alongside text prompts and existing controllable-generation tools, directing atmosphere while a depth or edge map stops the scene from disregarding the product’s physical placement.
Use both only when each supplies a separate constraint. A depth pass can keep a bottle placed on its plinth; an approved campaign still can call for cool brushed steel, hard lateral light, and sparse composition. Leave the product masked, and neither control has to perform the impossible job of reproducing tiny type or a proprietary contour.
How to generate locked-SKU ecommerce stills and reels
Approve the 3D source and shot list
Begin with the approved GLB or USD asset, including correct materials, dimensions, decals, and colorway. Lock product transforms, camera position, focal length, crop, and brand lighting for every PDP, campaign, or social shot. Name outputs by SKU, view, and revision; a later material change should never quietly overwrite an approved scene.

Render a product truth pack
Export beauty RGB, alpha matte, shadow catcher, depth, and normal passes at the delivery aspect ratio. Add a Canny or line-art pass for categories with unmistakable edges, including eyewear, footwear, or hardware. Keep the render and its metadata. That is what lets you replace a rejected background without generating the product again.

Write the brand pack before prompting
Set approved reference images, palette, surfaces, props, camera rules, positive prompt language, negative prompt language, and prohibited product claims. State protection rules literally: do not change logo, label text, closure count, stitching, colorway, or silhouette. “Keep it accurate” is weak; a concrete SKU checklist holds.

Generate the environment on its own
First generate a background plate, or composite the untouched product over a generated scene. For background replacement, make an inpainting mask where the preserved product is black and the generated scene is white. Stable Diffusion gets room to build the set while the product body stays out of redraw.

Add structural and style conditioning only where it earns its place
Send the product depth render through Depth ControlNet to preserve placement and separation. Use IP-Adapter with approved campaign imagery to steer art direction. If you need one unified finishing pass, keep img2img denoising low and limit higher-change work to masked background or transition areas; more denoising means bigger changes.

Batch stills with settings you can repeat
For each scene family, hold the prompt template, reference set, control inputs, seed policy, resolution, and workflow graph steady. Change one styling variable per run: set surface, prop, lighting mood, or crop. Save every accepted seed and conditioning setup—merchandising will ask for that same image in a new colorway.

Build reels from deterministic 3D motion
Animate the camera, object, or lighting in 3D, then export depth and normal sequences frame by frame. A ComfyUI Blender integration can load Blender EXR depth and normal sequences as ControlNet conditioning for temporal video generation. Keep prompt and conditioning rules identical across the sequence, then inspect for flicker, drifting reflections, and product details that change.

Run visual QA at SKU level before publishing
Reject any output where the logo, label text, feature count, colorway, silhouette, or product-to-surface contact has shifted. Check shadows and reflections as well. A product floating one millimeter above a stone plinth does not belong on a PDP; fix the mask edge, contact shadow, or background plate before regenerating the whole composition.

What denoising strength fits a locked product?
Use the lowest denoising strength that delivers the requested finishing change, and keep the approved product body outside the redraw region whenever possible. Img2img denoising runs from 0 to 1: 0 adds no noise, 1 fully replaces the input with noise, and higher settings create larger departures from the original.
Treat denoising as a risk dial. Low-denoise finishing can blend a reflection, soften a transition edge, or bring a scene’s color cast into a protected product plate. It does not authorize regenerating label typography, packaging claims, stitching, or a recognizable physical feature. Put major changes—new foliage, wall texture, table material, broad atmospheric lighting—inside the white inpainting region.
There is one useful exception: concept exploration before SKU approval. At that point, high-change generation can test surfaces, lighting worlds, and shot directions. Once an actual product variant reaches merchandising, go back to render-first and lock the subject.
How should you QA a generated product image or reel?
Compare every published frame against the SKU truth pack, not against whether the image simply looks convincing. Check logos and label text at useful zoom; count buttons, seams, closures, ports, stitches, and other category-specific features; verify the exact colorway; then match the outer silhouette against the approved render.
Then check physical integration. The contact point must meet the surface, cast shadows must run in a plausible direction, glossy materials should reflect the environment without inventing a second object, and background occlusion cannot cut into the protected asset. For a reel, scrub frame by frame for wordmark flicker, moving highlights, repeated props, and changes in product scale.
Use a repair-first rule. A failed background, matte edge, shadow, or reflection is usually a local production issue, so rerun that stage or composite a correction instead of regenerating a full hero image. Separate standardization, an executable brief, and final composition in the workflow, and selective reruns remain possible.
| Tier | Price | Included | Best for |
|---|---|---|---|
| Still-image proof of concept | Varies by compute provider | Not applicable | Testing one approved SKU, one camera view, and several background directions |
| SKU scene production | Varies by compute provider | Not applicable | Batching approved stills with saved seeds, masks, and ControlNet inputs |
| Short reel production | Varies by compute provider | Not applicable | Frame-sequence rendering, temporal conditioning, review, and edit |
One ecommerce still from an existing approved 3D camera view
Varies by GPU provider, workflow settings, and review effort3D render time + diffusion background runtime + any local inpainting runtime + human QA time
A five-shot product reel using deterministic 3D camera motion
Varies by frame count, resolution, GPU provider, and revision roundsPer-frame 3D pass generation + temporal diffusion runtime across every frame + flicker review + final edit
What does this workflow cost in practice?
Production cost is compute plus review, not a universal per-image fee. A still needs a 3D product render, scene-generation or inpainting runtime, possible local repairs, and SKU QA; a reel adds per-frame passes, temporal conditioning, flicker review, and editing. GPU pricing, model choice, resolution, batch size, and revision count set the compute portion.
Estimate at the shot-family level. One family might cover a SKU in a fixed three-quarter camera over six seasonal surfaces, or a five-shot vertical reel with camera motion already authored in 3D. Reuse the render pack, mask, depth pass, normal sequence, and approved style pack across that family so you are not paying to solve product fidelity again and again.
Keep human review as a real line item. Generation can produce new concepts, complex styling, on-model contexts, and believable material detail with far less turnaround than a conventional production cycle. A brand-critical hero still requires truth-pack approval before it goes to a marketplace, PDP, or paid placement.
What should you do if the output still changes the SKU?
If Stable Diffusion changes a SKU detail, stop pushing prompt emphasis and shrink the area it can alter. Replace the product plate with the original RGB render, tighten the alpha matte, put the changed region back under a protected mask, or rerun only the background and transition zone.
A changed logo, missing seam, added switch, or altered label is not a minor style defect. It shows that probabilistic generation was given responsibility for an exact product fact. Put that responsibility back in the 3D render, then let ControlNet, IP-Adapter, and the prompt vary the environment safely.