UndercurrentProduct

Activation probing for live LLM inference

Undercurrent runs probes on a model's internal activations while it generates — stop a generation mid-flight, or observe asynchronously without adding latency.

Open source
Apache-2.0, on GitHub and PyPI
vLLM + HF
Runs inside your existing serving stack
YAML
Declarative probe specs, no model surgery

Inline abort

Stop a request before the next token is produced

Async observation

Off the hot path, with no added response latency

Declarative specs

A few lines of YAML choose tensors, layers and positions

Runs where your model runs

vLLM with continuous batching, or Transformers on a laptop CPU

Declarative YAML specs

Say which tensor to capture, at which layers and token positions, and which probe receives it. No model surgery, no forked serving code.

Inline abort mid-generation

An inline probe runs on the generation path and can stop a request before the next token is produced.

Production-ready async execution

Observe-only probes run on bounded per-binding queues with overflow policies, and report to metrics and to file or webhook sinks with retries, dead-lettering and redaction.

Runs where your model runs

Inside your existing vLLM deployment with continuous batching, or with Hugging Face Transformers on CPU. Install with pip install undercurrent.

The vLLM adapter is experimental in v0.1 and currently supports vLLM 0.28 only. See the compatibility docs for tested combinations.

Open source, with enterprise support when you need it

Undercurrent is free to use. If you're running it in production, we can help.

Open source

Free · Apache-2.0

Install with pip, run it on vLLM or Hugging Face Transformers, and get help from the community on GitHub.

View on GitHub

Enterprise

Support from the team that builds Undercurrent

Talk to us about enterprise support, deployment help and licensing for running Undercurrent on your models and infrastructure.

Contact us

See the undercurrent before it surfaces.

Undercurrent is open source on GitHub, with a quickstart that runs on CPU.

View on GitHub