Independent benchmark of Whisper Thunder for text-to-video: quality, speed, controls, safety, and pricing claims
A source-bounded benchmark review of Whisper Thunder (Runway Gen-4.5): measured latency and cost, plus the evidence gaps for quality, controls, safety, and pricing.

Lamina Team
Product Team @ Lamina

Is Whisper Thunder good for text-to-video generation?
Whisper Thunder had credible evidence of top-tier text-to-video preference quality at launch, but the supplied evidence does not establish speed, fine-grained control, safety, or pricing as production guarantees. Artificial Analysis identified Whisper Thunder as Runway Gen-4.5, so assess Gen-4.5 rather than similarly named third-party sites.
The available Lamina results cover only nominal per-run cost and end-to-end latency under three input conditions. They include no blind quality ratings, prompt-adherence scores, temporal-defect counts, safety outcomes, completion rates, retries, or published reproducibility artifacts. Treat them as early operational measurements, not a full model benchmark.
Three single reported 5-second, 16:9, 24-fps Whisper Thunder conditions were compared: text-only, text plus a Lamina reference image, and text plus a reference image with explicit shot controls. No manual post-production was reported.
Nominal per-run cost
over Single reported run per condition
Generation latency
over Single reported run per condition
Generation latency
over Single reported run per condition
Quality and control-fidelity evidence
over Current supplied experiment data
| Metric | Value | Source |
|---|---|---|
| Runway-reported Artificial Analysis Elo at Gen-4.5 launch | 1,247 Elo | runwayml.comas of At release |
| Current supplied leaderboard leader in the audio-output category | Gemini Omni Flash, 1,240 Elo | artificialanalysis.aias of Supplied current leaderboard snippet |
| Text-only measured run | $0.040; 43 seconds | uselamina.aias of 2026-07-20 |
| Reference-image measured run | $0.040; 52 seconds | uselamina.aias of 2026-07-20 |
| Reference image plus explicit shot-controls measured run | $0.040; 48 seconds | uselamina.aias of 2026-07-20 |
What do the measured speed and cost results show?
The reported runs show the same nominal $0.040 price for text-only, reference-image, and reference-plus-shot-control inputs, though the reference-image condition took longer than the text-only baseline. The current data supports one narrow purchasing finding: these added inputs did not raise the measured per-run charge, but they may affect turnaround time.
The reference-image run added 9,010 ms over text-only. The reference-plus-controls run added 5,043 ms over text-only and was faster than reference-only. Each condition, however, is a single reported measurement rather than the planned repeated runs across a balanced prompt suite. You cannot infer median latency, queue sensitivity, failure-adjusted cost, or statistically reliable differences. Use these readings to size a pilot, not write an SLA.
The previously anonymous Video Arena model “Whisper Thunder” was Runway Gen-4.5.
Does Whisper Thunder currently lead text-to-video quality?
Whisper Thunder was reported as the leading model in the Artificial Analysis text-to-video benchmark at Gen-4.5’s release, but the supplied record does not support a durable present-tense “number one” claim. Runway reported the launch result, while the supplied current Artificial Analysis snippet names Gemini Omni Flash as the leader in an audio-output category. Rankings change, and category definitions matter.
This is meaningful evidence of broad human preference at a specific point in time, not independent proof of every vendor-level quality claim. The supplied material does not establish physical accuracy, temporal consistency, or prompt adherence on a defined, published test set. One user report describes frontier-level style and adherence but warns that consistency is not flawless; treat it as a practitioner signal, not benchmark evidence.
Do reference images and shot controls improve Whisper Thunder outputs?
The supplied experiment cannot show that reference images or explicit shot controls improve Whisper Thunder outputs because it records no quality, similarity, or camera-motion evaluation. It shows only that these input conditions ran at the same reported nominal price and produced different observed latencies.
A product-description source says Gen-4.5 supports text-to-video, image-to-video, and keyframes, with 5-to-10-second output and no audio. That supports the existence of basic creative inputs, not reliable control fidelity. Before committing a creative workflow, score a fixed prompt set for subject continuity, scene preservation, requested camera movement, temporal artifacts, and prompt compliance, then review results blind to condition.
Are Whisper Thunder safety and pricing claims verified?
No. The supplied sources provide no reliable primary documentation for Whisper Thunder safety policies, moderation behavior, rights protections, enterprise governance, commercial-use terms, credit consumption, or official pricing. Do not treat claims such as free access or no watermark on third-party and look-alike sites as Runway product terms.
Before production use, procurement should require current documentation from the actual Runway purchase or account flow. Confirm the applicable model, region, account tier, content restrictions, data handling, output rights, watermark policy, retention policy, and charges for retries or failed generations.
What should you validate before using Whisper Thunder in production?
Validate Whisper Thunder through a repeatable, account-specific acceptance test before relying on it for deadlines, regulated content, or customer-facing work. Keep duration, aspect ratio, resolution, generation mode, account tier, region, and submission window fixed. Randomize run order, and test text-only, reference-image, and controlled-shot variants separately.
Publish the exact prompts, seeds where available, reference files, timestamps, output hashes, UI or API receipts, raw outputs, and an unedited contact sheet. Run each condition multiple times and report medians, completion rate, retry rate, and failure-adjusted cost. Add safety/refusal cards and a blinded scoring rubric so quality and control claims are measured rather than inferred from a leaderboard or one generation.
Methodology
Original Lamina experiment run 2026-07-20. Hypothesis: Whisper Thunder’s advertised text-to-video quality, generation speed, control fidelity, safety behavior, and effective cost can be independently measured with a fixed, public test protocol. For identical 5-second, 16:9, 24-fps requests, reference-assisted prompting should improve subject/scene consistency and requested camera-motion adherence versus text-only prompting, while potentially increasing latency and cost. Create original benchmark imagery in Lamina first, publish its exact prompts, seeds, output files, timestamps, and hashes, then run every Whisper Thunder condition three times in a fresh account/session with no manual post-production. Use a 12-card balanced prompt set: 3 photoreal product/physics cards, 3 human-action continuity cards, 3 stylized/world-consistency cards, and 3 safety/refusal cards. Hold duration, aspect ratio, resolution, generation mode, region, account tier, and submission time window constant; randomize run order. Archive raw outputs, screen recordings of submission/completion, API/UI receipts, and an unedited contact sheet so the results are reproducible and independently auditable.. Measured 3 variant(s) for cost and latency on the Lamina image engine; numbers cited here are our own measurements.