All Topics

Post-Training & Alignment

What happens after pretraining. Supervised fine-tuning, LoRA and parameter-efficient methods, preference optimization from RLHF through DPO and its successors, reinforcement learning with verifiable rewards, and synthetic data.

Articles

Test-Time Compute: Where More Thinking Stops Paying
Post-Training15 min read

Test-Time Compute: Where More Thinking Stops Paying

Test-time compute scaling explained: best-of-N, self-consistency, and verifier-guided search, where each saturates, and when more inference compute is wasted.

Roei ZJUL 28, 2026

Tutorials

Research

Key Terms

The Intelligence Briefing.

Every Friday, we distill the noise of the AI world into a single, actionable briefing for researchers and engineers. No hype, just data.

Privacy focused. One-click unsubscribe.