Moderation that doesn't wait for the pixels
A lightweight probe reads the diffusion model's denoised latent tensor — before a single frame is decoded — and returns a safety score inline with generation.
Five orders of magnitude smaller
11.3M parameters vs. ~8B for a pixel-space guard model
Sub-10ms overhead
Moderation stops being a bottleneck in the generation pipeline, not an added queue
Pre-decode intervention
Catch violations before spending compute on the decode step, not after
Drop-in ↓
Details below
Real-time moderation API
Safety score returned alongside — or instead of — decoded output, pluggable into existing T2V/I2V/V2V serving stacks
Early-abort inference mode
Skip decode entirely on flagged generations, cutting both compute cost and exposure window
Custom category probes
Extend beyond binary adult-content detection to platform-specific policy categories, trained on your own moderation taxonomy
On-prem / air-gapped deployment
No external API call or content leaving the inference host, for platforms with strict data-handling requirements
Probes are trained on latents derived from real-world video passed through the encoder, then applied at inference to latents the diffusion model itself produces — a domain gap we haven't yet fully characterized. Adversarial robustness against latent-space evasion is also untested, and results to date are validated on CogVideoX specifically.
See it on your own pipeline
Talk to us about running a pilot on your video generation stack — no architecture changes required.
Talk to us