Articles

Technical Deep Dives.

In-depth analysis of AI architectures, deployment patterns, and the research shaping the field.

Kimi K3 3:1 hybrid layer stack: three KDA linear-attention layers with fixed-size state per full-attention layer with growing KV cache
attention mechanismsLLM architecture

Linear Attention at Frontier Scale: Kimi K3's KDA Claim, Fact-Checked

Kimi K3 runs linear attention in 3 of 4 layers and claims it beats full attention. What the 48B evidence shows, and what stays unverified at 2.8T scale.

RayZ·13 min read·
DeepSeek V4 and the Hybrid Attention Bet
attention mechanismsLLM architecture16 min read

DeepSeek V4 and the Hybrid Attention Bet

Inside DeepSeek V4: hybrid attention (CSA + HCA), 1.6T MoE, 1M context, and the lineage from MLA to NSA to DSA that made it possible.

RayZAPR 27, 2026
LLM Inference Optimization: The Engineering Behind Fast, Cheap AI
LLM architectureinference optimizationdeep learningAI engineering18 min read

LLM Inference Optimization: The Engineering Behind Fast, Cheap AI

Master LLM inference optimization: speculative decoding, KV-cache compression, quantization, FlashAttention, and serving frameworks compared for fast, cost-effective AI.

RayZAPR 6, 2026
Understanding Transformer Architectures from Scratch
LLM architectureattention mechanismsdeep learningmodel training22 min read

Understanding Transformer Architectures from Scratch

Master the transformer architecture from first principles: self-attention, multi-head attention, positional encodings, encoder-decoder design, and modern innovations like RoPE, GQA, and SwiGLU, with code.

RayZAPR 6, 2026
Vibe Coding and the New AI-Assisted Development Stack
AI agentsAI engineering16 min read

Vibe Coding and the New AI-Assisted Development Stack

Explore vibe coding: the AI development paradigm coined by Karpathy. Compare Cursor, Claude Code, Google Antigravity & Copilot — with honest takes on which tools actually deliver.

RayZAPR 6, 2026