Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
Khatri, A.
Companion writeup to our paper "Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations" (arXiv:2608.16970)
Peer-reviewed papers and preprints from the Wrynx team.
Khatri, A.
Companion writeup to our paper "Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations" (arXiv:2608.16970)
Khatri, A., Chan, D. L.
A reproducibility study of latent-space safety probes across model families.
Khatri, A., Prabhu, C.
Video generation models can be coaxed into producing explicit or harmful content with little effort. We trained lightweight probes on the internal latents of CogVideoX to detect unsafe content before a single pixel is decoded — matching 8B guard models at over 1,000× lower latency.
Khatri, A., Prabhu, C., Neogi, O.
We kept running into the same frustration with external guardrails: they're blind. They see tokens — what went in, what came out — but nothing in between. So we asked a simpler question: does the model already know when a prompt is harmful? We trained lightweight probes on LLaMA-3.1-8B's hidden states and found that it does — matching 7B guard models at a fraction of the cost.