HOST: If I compare language models for my team, why should I care about this paper? EXPERT: It offers another training choice to consider. A model might learn from groups of text pieces as well as individual pieces. The paper tests training and benchmark scores, not a workplace rollout. HOST: What counts as a group here? EXPERT: The model combines internal information from neighboring tokens. Those are like small pieces of text. Think of several pieces in a travel itinerary. Its learned representation of that group is not necessarily a readable idea. HOST: So does it stop predicting the next piece of text? EXPERT: Now, NCP Arch Preview still predicts the next token. What's new is an added module that predicts a future group-level representation and feeds that back into text generation. HOST: So what did the authors see when they trained it? EXPERT: In Stage-1, they report reaching OLMo-3-7B's final training loss using 51.3 percent of its training tokens. Their summary of downstream task scores was also 2.45 points higher for NCP-ArchPreview. HOST: Did that pattern continue in stage two? EXPERT: Right, the reported overall task score gain was smaller, and NCP-ArchPreview's code category average was lower. They also found that some Stage-2 versions had lower training loss but actually worse task scores, and they say long-context training is still outside this preview. So the practical takeaway is to check the tasks you need, not just the loss.