
Probing the Prefill: Can an LLM's Own Activations Tell You Its Code Is Vulnerable?
Companion writeup to our paper "Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations" (arXiv:2608.16970)
Insights and updates on AI safety, model security, and enterprise AI deployment best practices.

Companion writeup to our paper "Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations" (arXiv:2608.16970)

*A reproducibility study of latent-space safety probes across model families*

Video generation models can be coaxed into producing explicit or harmful content with little effort. We trained lightweight probes on the internal latents of CogVideoX to detect unsafe content before a single pixel is decoded — matching 8B guard models at over 1,000× lower latency.

We kept running into the same frustration with external guardrails: they're blind. They see tokens — what went in, what came out — but nothing in between. So we asked a simpler question: does the model already know when a prompt is harmful? We trained lightweight probes on LLaMA-3.1-8B's hidden states and found that it does — matching 7B guard models at a fraction of the cost.

Wrynx secures enterprise foundation models using Latent Space Probes — real-time model-layer defenses that detect unsafe concepts before they become harmful outputs. Built for CISOs and AI leaders, Wrynx enables secure, scalable generative AI deployment with runtime protection and executive-grade risk visibility.

Wrynx introduces latent space probes that analyze LLM activations in real time to detect harmful and sensitive concepts. In a case study on Llama 3.1, we show how internal probing matches guard-model accuracy while delivering 1000× lower latency and cost.