All Topics

Quantization & Efficiency

Making models fit. Low-bit weight and activation formats, quantization-aware training versus post-training quantization, what each bit width costs in quality, and the memory arithmetic behind running large models on small hardware.

Articles

Diagram of a fragmented KV cache reorganized into reusable fixed-size pages, letting the same GPU serve more requests
Inference OptimizationQuantization18 min read

KV-Cache Engineering: The Memory Wall of LLM Serving

The KV cache is the real bottleneck in LLM serving: 70-90% of GPU memory at long context. PagedAttention, prefix caching, MLA, and quantization explained.

Roei ZJUL 9, 2026

Tutorials

Key Terms

The Intelligence Briefing.

Every Friday, we distill the noise of the AI world into a single, actionable briefing for researchers and engineers. No hype, just data.

Privacy focused. One-click unsubscribe.