HOST: So if you're helping train a smaller language model, why should you care about this paper? EXPERT: It may help you decide which feedback from a larger model to test when the two split text differently. The paper studies training results, not a deployed tool. HOST: So what does splitting text differently actually change? EXPERT: Each model predicts tokens, pieces of text. One might see a word as one piece, and the other as several. It's easier to compare predictions when their pieces line up exactly. HOST: So, did the authors use only those matching places? EXPERT: Right, so they tested that approach and also added feedback for mismatched stretches. Strict full compared probabilities over every shared token at matching places, and strict top 16 used the 16 shared tokens the student rated highest there. HOST: So what happened when they included the mismatched stretches? EXPERT: For each of the three named model pairs, every tested positive weight on that added feedback lowered the full average benchmark score against its strict-only comparison. That score combines mathematics and code accuracy. HOST: So, does that mean feedback on mismatches is always harmful? EXPERT: No, they tested a particular way of scoring those stretches with particular models and tasks. Their practical distinction is that reaching more positions isn't the same as showing the added feedback actually helps learning.