HOST: If I build tools for data scientists, why should I care about this? EXPERT: It points to a possible way to train an agent that writes modeling programs without having to run every proposed program. The paper tests that idea on benchmarks, not in a workplace. HOST: Why is running them such a big part of the job? EXPERT: Each program needs its own isolated workspace and machine time. The agent can propose many solutions together, but the authors say those runs become the training bottleneck. HOST: So what replaces a run? EXPERT: So in this case, the authors call it a world model. You can picture an agent writing a program to classify comments, and the language model steps in to estimate the result without actually running that program. HOST: Well, how does it know when the estimate is wrong? EXPERT: Summit times are still run for real. Their predicted and real scores help correct a persistent scoring error and reduce the influence of unpredictable error. HOST: So what did the authors actually find, and what should I be careful not to assume from that? EXPERT: On their held-out competition tasks, their 4B agent used fewer GPU hours to train than its same-scale RL execution counterpart. A GPU hour just means one graphics processor running for an hour. On both benchmark averages, their 4B agent outscored Kimi 48B, a 3B, while their 9B agent outscored Nemotron 120B, a 12B. They also note possible exposure to older public competition material, and these tests don't establish how the method would work in your own setting.