HOST: So if I'm working on making an AI assistant better at my job, why should I care about this paper? EXPERT: You might get a request like, make it better at research. The paper helps you see why choosing what to practice and how to check it are separate jobs from making an update. HOST: So what did the researchers ask the AI to do? EXPERT: They gave it a broad capability goal. It had to choose learning material and a method, while the researchers kept the final checking tasks out of sight. HOST: Why hide those tasks? EXPERT: A self-chosen practice test can reward the wrong thing. Imagine practicing meet article summaries when the real need is spotting unsupported claims. That is only an illustration, not one of their results. HOST: So, did the agents at least manage to train new versions? EXPERT: Often, yes. In the adaptive feedback setting, 28 out of 30 configuration goal cells produced a saved model that could be evaluated, but only the Terra-directed mathematics cell retained a version above its starting Qwen 3.5 4B model under the study's rule. HOST: Does that retained result settle whether it learned math better? EXPERT: No independent repeat is reported for that result. Terra selected it using repeated overall scores on the same hidden tasks, without a separate confirmation set. The practical takeaway is to keep your own practice checks distinct from an independent check against the starting system.