Our open inference-serving platform. Apache-2.0.
Northbeam sits between your models and your users: continuous batching, prefix caching, per-tenant rate limits and a Prometheus-native metrics surface. One binary, one config file, no control plane to babysit.
~38 tok/s/GPU sustained, mixed chat workload410ms at 60% load< 900MBRelease v2.4.1 (May 2026) added speculative decoding and OpenAI-compatible
/v1/chat/completions. Changelog and binaries on the blog.