HOST: Have you ever taken a sensible-looking route and only later realised it was a detour? EXPERT: That is the problem behind Taste-Bench. It pauses an AI agent at a real choice and asks which next step is better. HOST: What sort of choice? EXPERT: In a coding task, should the agent run a small diagnostic test or rewrite a big section immediately? The later task record tells us which route worked HOST: How hard is the test? EXPERT: The authors made 502 questions. Their best tested model reached 59.7 percent on the benchmark’s scoring rule. HOST: Does that score mean the model has good judgment everywhere? EXPERT: No. These examples mainly come from software and research tasks, and the labels rely on recorded outcomes. HOST: What can teams do with the idea? EXPERT: Save key decisions and later outcomes. Then review where an agent made a costly detour before giving it more autonomy.