Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerics?

Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerics?

Wrynx Team
8 min read
AI SafetyInterpretabilityLLM MonitoringResearch

Activation probes hold up across the batch sizes and precisions a serving stack uses. The text they read does not. Aggregate accuracy hides that difference by up to 9x.

The mismatch nobody measures

Activation probes have quietly moved from interpretability tool to production monitor. These small classifiers on a model's hidden states now flag deception, high-stakes requests, sleeper-agent behaviour and misuse. They are cheap because they reuse activations the model computes anyway.

But a probe is trained once, on activations extracted under one inference configuration. In deployment, the same model runs at whatever batch size the scheduler assembles, often at a different precision. Neither change means anything semantically. Both change the numbers.

Common GPU kernels for matrix multiplication, normalization and attention are not batch-invariant. Their reduction order depends on batch shape, and so does their rounding. bfloat16 and float16 round differently again. Prior work showed these effects change greedy LLM outputs and benchmark scores. Nobody had checked whether they break the probes reading those models' internals. So we did.

What we did

We ran a controlled study on three models, with a fixed software stack and deterministic PyTorch settings throughout.

Axis Values
Models Llama-3.1-8B, Qwen3-8B, Gemma-3-4B
Batch size 4, 8, 16
Precision float32, bfloat16, float16
Depth block 3, 50%, 75%, 100%
Token position last prompt token; generated tokens 5, 10, 20
Probes logistic regression, 1-layer MLP, 3-layer MLP
Tasks Rotten Tomatoes sentiment, SMS Spam (TruthfulQA for activation analysis)

We trained 768 probes on one configuration and evaluated each on every other, for 5,568 transfers in all. We also went beyond accuracy. We ran each probe over both runs' activations for the same examples and counted how many individual verdicts changed.

One caveat up front. The 8B models' float32 weights did not fit on our L4, so those runs used an A100. Every batch-size comparison and every Gemma comparison holds the GPU fixed. The same-GPU subset alone supports every conclusion below.

Finding 1: at the prompt, probes don't care

Across 1,392 prompt-position transfers, probe accuracy never moved by more than 0.47 percentage points. Only 0.076% of individual verdicts changed, and gains and losses were balanced. In float32 with only the batch size varied, not one verdict changed in 201,960 evaluations.

The activations do move, by roughly the precision of each number format:

  • float32 is batch-invariant at the prompt for Llama and Qwen on the A100. Across batch sizes, 99.9–100% of prefill vectors are bitwise identical.
  • bfloat16 batch-size noise is about 8x larger than float16's. The median relative ℓ2 is about 10⁻² versus 10⁻³, matching float16's three extra mantissa bits (2³ = 8).
  • On Gemma, where both runs share a GPU, changing only the batch size in bfloat16 moves activations slightly more than the entire float32-to-bfloat16 conversion. In that sense, a probe trained at one bfloat16 batch size is already off-distribution at another.
  • The noise is structureless. It has no directional drift and no exponential amplification with depth. Global geometry is preserved, with CKA ≥ 0.995.

Finding 2: during decoding, rounding becomes different text

Once the model starts generating, the picture changes. Decoding is greedy, so we reconstructed each run's realised tokens from stored logits without re-running anything.

A bfloat16 batch-size change flips the very first generated token for 2.1% of rows. By token 20, the two runs are on different tokens in 25% of rows. Under float16 the figures are 0.4% and 3.6%. Under float32, not one of 5,610 rows per model diverges anywhere.

We observed the mechanism directly rather than inferring it. A perturbation the size of rounding error breaks a near-tie between the top next-token candidates. From then on the model writes a different sequence. Divergence also starts early: a third of diverging rows already differ by token 5.

Probe accuracy at decode positions changes in 74% of transfers, by up to 2.35 pp. The size of the change tracks how many rows diverged (Spearman ρ = 0.73). That correlation holds within same-GPU transfers, within each precision, and in the strictest subset we could construct.

Finding 3: accuracy hides the churn, but the probe is fine

Accuracy doesn't register which examples are correct. A probe can hold its accuracy while quietly changing thousands of decisions. Measured example by example, verdicts change 2x to 9x more often than the accuracy change suggests.

Probe reads Verdicts flipped, all rows Same tokens Tokens diverged Accuracy change Cohen's κ
Prompt 0.08% 0.07% 0.11% 0.04 pp 0.999
Token 5 0.87% 0.12% 3.75% 0.18 pp 0.981
Token 10 1.64% 0.13% 7.48% 0.23 pp 0.965
Token 20 2.76% 0.15% 12.87% 0.31 pp 0.941

We split rows by whether the two runs generated the same text. On rows whose tokens matched, the flip rate stays flat at 0.12–0.15% at every position. On rows whose tokens diverged, it climbs to 12.9%. Verdicts move because the text changes, not because the arithmetic does.

The churn is also benign in the ways that matter:

  • Flips are symmetric. At token 20, 6,671 verdicts moved from class 0 to 1 and 6,813 moved the other way.
  • Ranking survives. AUROC moves by at most 0.05 pp, and TPR at 1% FPR by at most 0.34 pp.
  • Diverged text isn't harder to read. Probe accuracy on diverged rows is neither consistently higher nor lower than on matched rows; the gap scatters around zero. Probes reading later tokens simply start closer to their decision boundary.

Why the probe holds when the model doesn't

The same perturbation on the same prompt activations flips 0.14% of probe verdicts. It flips 2.1% of the model's own next-token choices, a 15x difference.

The reason is margin. A binary probe with a well-separated boundary absorbs a 10⁻² relative nudge. The model's next-token choice picks the top candidate from roughly 10⁵ tokens, and the top two are often nearly tied. A nudge that size is enough to swap them. Probes are robust because their decision tolerates much more noise, not because the representation is stable in absolute terms.

This gives us a simple way to think about any numerical perturbation:

  1. Benign: the representation changes but no decision does. That's the prompt position.
  2. Behaviourally consequential: a discrete choice flips, so the model's future input changes. That's decoding.
  3. Validity failure: the format can't represent the activations at all. That's Gemma-3-4B in float16, where hidden states overflow to NaN. We excluded it entirely.

What this means if you deploy probes

Probes on the prompt can travel. This covers input classifiers, pre-generation monitors and most truthfulness probes. They can be trained under one batch size and precision and served under another. Our worst case was a 0.47 pp accuracy change and about one verdict in 1,300 flipping.

Probes on generated tokens inherit the serving config. The probe itself is fine, but the text it judges depends on batch size and precision. You have three options. Use batch-invariant decoding, re-encode generated text deterministically before probing, or measure your own flip rate.

Need exact reproducibility for audits? Use float32. In our setup, float32 reproduced exactly across batch sizes. That held for prompt activations, every generated token and every probe verdict.

Change how you evaluate monitors. We have three recommendations for robustness evaluations of activation monitors:

  • Report per-example agreement, not just aggregate accuracy.
  • Separate representational noise from input change.
  • State the serving configuration alongside the result.

Limits, and what's next

The study has limits:

  • We used one software stack and batch sizes of 4–16. Production serving uses larger, ragged and continuously batched workloads.
  • We tested two binary tasks on dense models of 4B–8B parameters. Safety probes run at low false-positive rates on harder data, where margins are thinner.
  • We trained each probe once, so we still lack a baseline for how much probes vary across training seeds.

We'd expect harder settings to be more sensitive, not less. That is why per-example agreement should be the default measurement.

Read the full paper: Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism? (arXiv:2609.31796). We will release the extraction and analysis pipeline, the per-row metric caches, and the scripts behind every table and figure.

Questions, or want to compare notes on probe monitoring? Reach us at research@wrynx.com.