Activation probing for live LLM inference
Undercurrent runs probes on a model's internal activations while it generates — stop a generation mid-flight, or observe asynchronously without adding latency.
Inline abort
Stop a request before the next token is produced
Async observation
Off the hot path, with no added response latency
Declarative specs
A few lines of YAML choose tensors, layers and positions
Runs where your model runs
vLLM with continuous batching, or Transformers on a laptop CPU
Declarative YAML specs
Say which tensor to capture, at which layers and token positions, and which probe receives it. No model surgery, no forked serving code.
Inline abort mid-generation
An inline probe runs on the generation path and can stop a request before the next token is produced.
Production-ready async execution
Observe-only probes run on bounded per-binding queues with overflow policies, and report to metrics and to file or webhook sinks with retries, dead-lettering and redaction.
Runs where your model runs
Inside your existing vLLM deployment with continuous batching, or with Hugging Face Transformers on CPU. Install with pip install undercurrent.
The vLLM adapter is experimental in v0.1 and currently supports vLLM 0.28 only. See the compatibility docs for tested combinations.
Open source, with enterprise support when you need it
Undercurrent is free to use. If you're running it in production, we can help.
Open source
Free · Apache-2.0
Install with pip, run it on vLLM or Hugging Face Transformers, and get help from the community on GitHub.
View on GitHubEnterprise
Support from the team that builds Undercurrent
Talk to us about enterprise support, deployment help and licensing for running Undercurrent on your models and infrastructure.
Contact usSee the undercurrent before it surfaces.
Undercurrent is open source on GitHub, with a quickstart that runs on CPU.
View on GitHub