HOST: So why should someone using a coding agent for research experiments care about this? EXPERT: They might give the agent practical package instructions before it starts, what to run and how to check it. The paper tests that idea on research benchmarks, not at a workplace. HOST: Is that different from just handing the agent a long manual? EXPERT: Yes, Disco turns source material into these small skills. The agent opens an entry instruction and then follows links to deeper detail only when the task calls for it. HOST: What would that look like for a real question? EXPERT: Here's just an invented illustration, not a reported result. Ask an agent to compare two model serving packages. A skill could remind it to keep the workload the same and check its measurements. HOST: So what did the authors actually measure? EXPERT: On MLEBench's machine learning competitions, they report a higher Any-Medal score for GPT-5.5 Codex with task-oriented skills than without them. Any-Medal records whether a run earned a medal. They also report a higher average reproduction score on PaperBench with selected skills from related papers. HOST: Did the skills help every paper reproduction? EXPERT: No, two paper bench tasks actually scored lower with skills. The authors suggest a retreat skill might have distracted the agent from a better approach, but they don't actually establish that as the cause. Skill creation also had a separate budget, so the practical takeaway is to inspect what a skill tells an agent to do and see if it actually fits the task.