Scaling Laws and Interpretability of Learning from Repeated Data
ID: 06a6e892-559d-5ea3-9250-47bd13c48820
STIX ID: report--06a6e892-559d-5ea3-9250-47bd13c48820
Feed Name: Anthropic Research
This research paper studies the effects of repeating small fractions of training data on large language models, revealing a strong double-descent phenomenon where certain repetition frequencies cause large drops in test performance by inducing memorization and damaging internal generalization mechanisms such as induction heads.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
