Research

Peer-reviewed papers and preprints from the Wrynx team.

Accepted — IEEE DSN-W 2026

Safety Beyond the Interface: Detecting Harm via Latent LLM States

Khatri, A., Prabhu, C., Neogi, O.

We kept running into the same frustration with external guardrails: they're blind. They see tokens — what went in, what came out — but nothing in between. So we asked a simpler question: does the model already know when a prompt is harmful? We trained lightweight probes on LLaMA-3.1-8B's hidden states and found that it does — matching 7B guard models at a fraction of the cost.