HOST: So why should someone supervising a long coding job care about this? EXPERT: They might find it really useful to keep the jobs current stage and the required checks visible, so an AI agent doesn't just announce it's done. The paper tests that idea on benchmarks, not in an actual workplace. HOST: What actually keeps the agent on track? EXPERT: State M uses a shared runbook, so it's a written procedure that both the agent and a person can inspect. It records the current phase and can check conditions before the agent moves to the next one. HOST: So what would that actually look like during a task? EXPERT: In one reported terminal task, the agent had to check a live web service setup from the user's end before handoff. If that check failed, the run could stay open for repair. HOST: And what did the broader test count? EXPERT: Terminal bench has 89 tasks, run five times each. The authors report 424 successes in 445 trials for GPT-5.6, sole x-high, with a frozen state M runbook. That is a raw public submission score, review still pending. HOST: So can the same runbook travel to any model or job? EXPERT: No. Used unchanged, the GPT-developed runbook lowered the reported DeepSeek V4 flash score. The authors adapted it before their later gained. Results across business bench task families were mixed too. The practical lesson is to specify the handoff you need to check, not assume one procedure fits everything.