Q4_K_M vs Q5_K_S on three eval suites: MMLU dropped 0.4 points, GSM8K 1.1.
The surprise was long-context recall — keep the KV cache at higher precision
and quantize weights aggressively. Full tables and scripts inside Northbeam
v2.4.1.
Recall@k on golden sets, faithfulness scoring with an LLM judge, and a regression harness in CI. If your eval fits in a screenshot, it isn't an eval.
Why naive request batching wastes 60% of FLOPs and how iteration-level scheduling fixes it. With diagrams drawn in 20 minutes, as tradition demands.