HOST: So why should someone training language models care about this? EXPERT: It could help them investigate cheaper ways to generate training examples. The authors tested a way to use less precise calculations without the same loss of benchmark score they saw with some comparison methods. HOST: So, what goes wrong when the calculations are less precise? EXPERT: Generation and training may round nearly identical values differently. Imagine two people choosing tiles for almost the same spot and picking opposite tiles. HOST: So how does Trace bring their choices together? EXPERT: It’s a small clue about the choice made during generation. Training uses that clue when rounding the matching value. The authors call the low-precision format FP4, meaning four-bit floating point. HOST: Did that make the tested model as fast as the headline suggests? EXPERT: For Qwen 3.5-35B-A3B, the authors report up to 5.4x the BF16 references decoding throughput at 128K output length on four GB200 graphics processors. That's a specified generation test, not a speed claim for every training job. HOST: And what should a team not assume from the result? EXPERT: They say local rounding agreement doesn't guarantee agreement across the whole model. Their compact clue is approximate, and it can't erase differences caused by an older generated example. So really, treat this paper as more of a guide for a training experiment rather than something you deploy as a tool.