HOST: So, if I manage a coding assistant, why should I care about this? EXPERT: It suggests a possible way to use the assistant's task records to choose later training material. It does not show that this will work in your team. HOST: What counts as a task record? EXPERT: It's not just the answer that counts. It can include the request, any tool calls, what the tools returned, and the outcome, like keeping the test result alongside a code change. HOST: And how do the authors decide what the model should practice? EXPERT: Their routing harness estimated how much capability each task needed. They used that estimate to introduce harder training examples in stages. HOST: Did training change the measured results? EXPERT: The authors report that NeoHorse-1-4B scored 64.87 on the average of their ten text-based benchmarks, versus 58.94 for its Qwen3.5-4B base model. That average combines test scores; it's not a real-world completion rate. HOST: So did the assistant keep improving itself? EXPERT: That remains open. The authors tested a single pass, not repeated cycles. The practical takeaway is to keep actions and outcomes together when reviewing task records, without assuming those records will automatically improve a deployed assistant.