Discussion about this post

User's avatar
Adam Zachary Wasserman's avatar

Great debugging, and the "honestly" framing invites one more scope condition that even Chinchilla inherits: both laws were fit on English.

Chinchilla fixed the data-vs-size bug, but look at your own Step 3. Kaplan's conclusion, you say, is "entirely accurate given a maximum number of tokens, but doesn't apply to the infinite-data limit scaling laws aim to model." Another similar limitation remains to be named: Chinchilla's law is entirely accurate given English, and a law about "language models" is implicitly claiming to model language, not one language.

My pre-registered experiment ran the same architecture, same optimizer, same compute, with only the training language changed. A 125M-parameter transformer on French reaches grammatical competence (100% on agreement probes) at about 197M tokens. The identical model on English is still at chance past 3 billion. That is more than a 15x gap in emergence threshold from the language alone, and against Pythia's English curve it puts French at roughly 50 to 100x more training-efficient on identical hardware.

English is morphologically impoverished, one of the least efficient languages in the twelve-language set I tested. It forces the model to infer from distribution what richer languages mark explicitly on the word, so English is unusually data-hungry. Which means Chinchilla's ~20-tokens-per-parameter optimum is really English's optimum. A morphologically richer language sits on a different, lower-token frontier. "How much data is compute-optimal" turns out to be partly a question about which language.

And this one cuts against your closing advice, in a good way. The Kaplan bug was something the big labs quietly knew; the language dependence is something that the field mostly hasn't flagged, and it is exactly the kind of question a non-big-lab researcher can settle, because the controlled version is cheap: a 125M model on a 92.5M-word French corpus reaches 86% on a native Quebec-French grammar benchmark for 65 dollars of compute. Pre-registered and reproducible: OSF SJ48B and CN75D; deposit "The Scaling Hypothesis Is Language-Contingent" (doi.org/10.5281/zenodo.19423151).

Scaling laws, honestly, are English scaling laws.

No posts

Ready for more?