HOST: So, if I help train a search assistant, why should I care about this? EXPERT: If it writes its own practice questions, you might want to know whether a better training score means better answers, or just repeated mistakes. HOST: How could a repeated mistake earn a better score? EXPERT: One part drafts a question and answer. Another learns from that answer. If the draft was wrong, the learner may later repeat it, and their agreement can earn credit. HOST: So what did the authors change? EXPERT: So what they did was split their source documents into two groups. Questions from either group get scored by an answering model trained only on the other group. They call that method CrossFit. HOST: What would that look like with an ordinary question? EXPERT: Imagine a draft incorrectly says a port is in Harbor A. A scorer that learned that draft might repeat it. A scorer kept away from that document cannot repeat it simply because it trained on that document's draft. HOST: So what did the tests actually show? And does making that swap really make the answers trustworthy? EXPERT: For Qwen 3.5 to 4B, the authors report fewer matching wrong answers and a higher average on their seven search benchmarks than with the standard training loop, but their audit didn't resolve every case, and models can still share mistakes from other sources. The practical point is to check where a scorer learned its answers, not treat agreement alone as proof.