The Third Wave of Artificial Intelligence Scaling
For years, progress was governed by the Chinchilla scaling laws: more pre-training parameters plus more web tokens equaled higher accuracy. However, as high-quality human text exhausted and training runs surpassed $100M, returns began plateauing.
Inference-Time Compute: Thinking Before Speaking
The breakthrough achieved by OpenAI (o1, o3) and DeepSeek (R1) unlocked a new dimension: inference-time compute. Instead of generating the next token in milliseconds, models spend seconds or minutes searching reasoning trees, self-critiquing, and evaluating alternative hypotheses before delivering the final answer.