HOST: If I help improve an AI assistant at work, why should I care which practice tasks it sees next? EXPERT: Choosing tasks as its problems change might make improvement time more useful. The paper tests that possibility on benchmarks, not in a workplace. HOST: So what exactly is being improved in these tests? EXPERT: The harness is basically the set of prompts, tools, and rules around the model. Active Sadler steps in to pick tasks that actually give you evidence, and then an existing optimizer makes those changes to the harness itself. HOST: So how does it decide which task deserves attention? EXPERT: It groups failures that seem to share a weakness. Imagine several fictional calendar tasks failing because attachments are omitted. It may revisit that group or try a task it hasn't seen to look for another weakness. HOST: What did the tests show? EXPERT: So what the authors saw was a bump in passive one when they added active Sadler to AutoSadler compared to using a fixed random task order. We're talking about gains of 4.4 percentage points on Gaia 2 and 7.5 on Terminal Bench 2.0. HOST: Does that mean we can put it straight into our assistant? EXPERT: No, the authors tested public and synthetic benchmark tasks, not real users, and they say they didn't establish production safety or privacy readiness. The practical lesson is to consider whether your practice schedule should change as failures change.