Generative video and media

Moderation that doesn't wait for the pixels

A lightweight probe reads the diffusion model's denoised latent tensor — before a single frame is decoded — and returns a safety score inline with generation.

97.29% F1
Adult-content detection on CogVideoX
4–6ms
Latency per 10s clip, vs. 3–5s for pixel-space baselines
>1,000×
Faster than decode-then-classify moderation

Five orders of magnitude smaller

11.3M parameters vs. ~8B for a pixel-space guard model

Sub-10ms overhead

Moderation stops being a bottleneck in the generation pipeline, not an added queue

Pre-decode intervention

Catch violations before spending compute on the decode step, not after

Drop-in ↓

Details below

Real-time moderation API

Safety score returned alongside — or instead of — decoded output, pluggable into existing T2V/I2V/V2V serving stacks

Early-abort inference mode

Skip decode entirely on flagged generations, cutting both compute cost and exposure window

Custom category probes

Extend beyond binary adult-content detection to platform-specific policy categories, trained on your own moderation taxonomy

On-prem / air-gapped deployment

No external API call or content leaving the inference host, for platforms with strict data-handling requirements

Probes are trained on latents derived from real-world video passed through the encoder, then applied at inference to latents the diffusion model itself produces — a domain gap we haven't yet fully characterized. Adversarial robustness against latent-space evasion is also untested, and results to date are validated on CogVideoX specifically.

See it on your own pipeline

Talk to us about running a pilot on your video generation stack — no architecture changes required.

Talk to us