HOST: So if I handle the same sorts of messages every day, why should I care? EXPERT: You might be able to describe the sorting once and reuse a small local helper. That's the possibility this paper explores, not a tested promise for your inbox. HOST: So what would I actually describe in this context? EXPERT: You could say that a request for a signature today is urgent, while a newsletter can wait. The authors call that written task description a specification. HOST: So, how does a description become something that handles new messages? EXPERT: Larger teacher models make practice messages and answers. Those examples train a small adapter, a task-specific set of model settings used with a shared fixed language model. HOST: Did the authors test whether the answers were right? EXPERT: On FuzzyBench-Hard, their trained compiler scored 0.836 versus 0.224 for the earlier fast compiler. Each score is the share of outputs a judge considered correct for the written task, not just identical to one sample answer. HOST: So what is the catch for someone trying a new task? EXPERT: Right, so building takes longer, and teacher-made examples can carry errors. The authors say it's important to validate outputs when correctness matters. Their website, avatar, and translation examples show applications, but they didn't report systematic user studies.