A new study used chess as a controlled experimental platform to systematically explore the interaction between pretraining and reinforcement learning (RL)...
论文研究HuggingFace Daily Papers(社区热门论文)
Today AI Intelligence Brief
A new study used chess as a controlled experimental platform to systematically explore the
interaction between pretraining and reinforcement learning (RL) in reasoning tasks. The study found that the post-training performance given the RL amount of computation can be accurately predicted by the pretraining loss, and the slope of the RL reward curve increases approximately linearly with the number of pretrained tokens. On a 1B-parameter mathematical language model, checkpoints with longer pretraining perform better and improve faster under RL.
Original Article Excerpt
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
A new study used chess as a controlled experimental platform to systematically explore the interaction between pretraining and reinforcement learning (RL)...