Scaling AI and Getting More Efficient
Kaplan/Chinchilla scaling laws provide a broad description of the resources needed to improve AI. The measured relationship is that loss falls as a power law in compute: L ≈ L∞ + A·C^(-α), with α somewhere around 0.05–0.07 for language models. Each halving of excess loss costs roughly 10–1000x more compute depending on the exponent. Converting …