Articles

Technical Deep Dives.

In-depth analysis of AI architectures, deployment patterns, and the research shaping the field.

Diagram of a weight precision ladder from 16 down to 1.58 bits and an activation distribution with one outlier spike, feeding a low-bit model that fits on one GPU
Quantizationinference optimization

Quantization Deep Dive: FP8 Training, FP4, and the Outlier Problem

A technical guide to LLM quantization: FP8 training, NVFP4 and MXFP4, W4A4 inference, the outlier problem, and where low-bit precision quietly breaks accuracy.

RayZ·19 min read·

No articles found for "Quantization".