Applied AI & software, shipped to production since 2019.
Machine learning that survives contact with production
We are a small engineering studio. We design, build and operate AI-powered
backends: retrieval pipelines, inference serving, evaluation harnesses and the
unglamorous plumbing (queues, caches, observability) that keeps them alive at
3am. No slideware — every engagement ships running code with tests.
Currently accepting Q3 projects. Two slots left for
retrieval-augmented search and inference-cost reduction work.
Get in touch →
What we do
- RAG & search — document pipelines, hybrid retrieval, grounded answers with citations.
- Inference serving — GPU scheduling, batching, quantization; p99 latency budgets we sign.
- Backend platforms — Go services, Postgres, queues; boring technology, exciting uptime.
Latest from the blog: Quantizing a 70B model without losing the plot (Jun 2026).