HOST: So why have the smaller model write answers before learning from the teacher? EXPERT: The teacher can then guide it on words it actually chose, rather than only on teacher-written answers. HOST: And what does this study change about that process? EXPERT: Its replay buffer version saves older student answers and uses them again for training updates. HOST: So writing a new batch and updating the model aren't the same thing. EXPERT: Right, one step collects a batch of answers. In the authors' Figure 2 replay curve experiment, they then make 256 updates using stored answers. That is not an update count inferred from the main table score. HOST: So what is that Scrotalin ordinary reader? EXPERT: For the Qwen3-4B teacher and Qwen3-1.7B base student on AMC 23, the replay version scores 36.60% after 10 collection steps. That score averages correctness over 16 attempts per problem. The on-policy baseline, which learns from newly collected student answers, scores 34.64% in the main table. HOST: Does that mean the method works for every kind of question? EXPERT: No such claim follows from these tests. Training uses math questions, and the authors' additional tests outside math cover four benchmarks. A practical first look is the repository's dry run commands; they print proposed training launches without training a model.