HOST: So, if I manage a coding assistant, why should I care about this paper? EXPERT: It may help you think about the routine around the model, how it inspects files, checks changes, and handles failure, not only which model you choose. HOST: So the paper tests that routine. EXPERT: Yes, the authors call it a harness. They asked models to build one from a weak starting system, then froze it and tested it on tasks withheld during development. HOST: So what would a harness actually do during an ordinary task? EXPERT: In a hypothetical code fix, it could guide the assistant to inspect files, edit, run a check, and try again if the check fails. That example is an illustration, not a reported test case. HOST: Did the models build useful ones? EXPERT: The authors found runnable harnesses, with results varying by task family. Code and search trailed their selected human-engineered references, writing approached its reference, and machine learning experimentation exceeded its reference when creators ran their own harnesses. Those references used different model-harness pairs. HOST: And did revising a harness after feedback make it better. EXPERT: All five self-runtime creators did better when they had visible coding feedback, but those gains got smaller on tasks they'd never seen before. When they fixed Gemini as the model in the harness, only the Opus lineage still improved on those held-out tasks; the other three actually went backwards. And because that held-out test was just coding tasks, the takeaway is that you need to check unseen tasks and the runtime model you care about before assuming a feedback gain will transfer.