Choosing a generation schedule may improve short-run AI output
The authors ask whether they can predict which generation order will work better for one model before using each order to produce an…
ISSUE 14/2026 · 10-10-2026
A useful correction needs a follow-up: did the coding agent fix the problem, or merely follow the advice? Imagine an agent repeatedly searching for a file with a command that does not work. Opera, the authors’ verbal critic, can point out the failed search, keep that…

6 paper · Original sources linked in every story
5 MINUTES ON PtoP
A coding assistant takes a correction, changes course and announces it is done. But did it solve the problem, or just follow the advice? That distinction runs through several items today. In a preprint about Opera, a critic that checks a coding agent’s visible work, the authors keep track of whether advice was attempted separately from whether the original issue was resolved. On one benchmark, they report a higher task-completion rate with the critic than without it. They also found cases where feedback made things worse. A helpful nudge, it seems, still needs a follow-up question. See: Follow coding-agent corrections until the problem is…
The same question gets harder when instructions spread. The authors of a study of shared coding-agent “skills” traced how instruction files were copied between public repositories. In their held-out historical evaluation, reviewing 100 repositories chosen by a copy-history model would have prevented 14.9% of later adoptions of skills flagged as high-risk, compared with 0.5% for the 100 most-starred repositories. Those are results of an evaluation rule, not real-world audits. The authors also report that copies rarely pick up later edits to their source. Fixing an original, then, need not fix what people already use. See: Copy histories can guide reviews of shared coding-agent…
There is a more personal version of this problem: an assistant remembering something you said. In conversations from ten consenting AI-companion users, the authors of another preprint found that only 3.4% of scoreable chat probes needed information outside the current thread under their recorded reading; a stricter reading put it at 1.3%. They also found that labelling earlier material “memories” made a generator bring up the past more often, including when it was unnecessary. Remembering is not just a question of whether an assistant *can* find an old exchange. It is also a question of whether this reply calls for it. See: Know when an AI companion needs to recall a conversation
That suggests a useful habit for bigger AI projects, too: decide what success would look like before declaring it. In Aspire, researchers let agents choose how to pursue broad improvement goals, then checked the resulting systems on evaluation tasks the agents had not seen. Improvements over the starting systems were uncommon in the reported settings. Meanwhile, Anthropic says it plans to commit $150 million over three years to help US government scientists try its tools; its announcement offers no results from those proposed projects. Access, effort and a completed plan are different from an independently checked outcome. See: Vague AI improvement goals need independent checks · Building on our commitment to American scientific discovery
None of this means every small task needs a formal test. It does suggest pausing between “the AI did something” and “the thing I needed is now true.” If you use an assistant this week, could you pick one answer or change it makes and ask yourself what simple check would tell you whether it actually helped? See: Follow coding-agent corrections until the problem is… · Vague AI improvement goals need independent checks
Just here for the stories? They are below, by topic.
PtoP · NEWSLETTER
Six papers and the news, explained in plain English, each morning after the editor approves the issue. Free, and you can unsubscribe at any time.
News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.
THIS ISSUE AT A GLANCE
↗02 / NEW / EDITOR PICKSFollow coding-agent corrections until the problem is resolvedA useful correction needs a follow-up: did the coding agent fix the problem, or merely follow the…
↗03 / NEW / EDITOR PICKSCopy histories can guide reviews of shared coding-agent skillsThe authors ask how to find the sources from which coding-agent skills spread, and whether reviewing…
↗04 / TOP VOTED · 6 MONTHSKnow when an AI companion needs to recall a conversationMost messages in these conversations did not need a distant memory, but some did. The authors ask…
↗05 / TOP VOTED · 6 MONTHSUnrelated Word Choices May Reveal What Model Training ChangedA model’s choice between ordinary words can carry information about training it received for an…
↗06 / TOP VOTED · 6 MONTHSVague AI improvement goals need independent checksAn AI agent can carry out steps toward a broad improvement goal, but the authors find that completing…
↗PODCAST / ENGLISH EDITION
The front-page episode covers the papers and news; each paper also has its own short conversation.
Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.
The authors ask whether they can predict which generation order will work better for one model before using each order to produce an…
A useful correction needs a follow-up: did the coding agent fix the problem, or merely follow the advice? Imagine an agent repeatedly…
The authors ask how to find the sources from which coding-agent skills spread, and whether reviewing those sources could limit the…
Most messages in these conversations did not need a distant memory, but some did. The authors ask whether an AI companion can tell the…
A model’s choice between ordinary words can carry information about training it received for an unrelated task. The authors tested…
An AI agent can carry out steps toward a broad improvement goal, but the authors find that completing those steps rarely produces an…
FROM THE NEWS DESK
Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.
Tools and building Hugging Face Blog · 09-10-2026 · company announcement
Ai2 says it changed how it shares computing machines among its research teams: managers now assign time budgets, and software decides which jobs run next. In its own 30-day test, Ai2 says teams received 98% of the GPU time they were owed, while occupancy held steady at 98%.
Imagine several teams sharing a kitchen. Each has a promised share of cooking time, but another team can use an empty stove rather than leave it idle. This is an illustration, not a reported Ai2 example.
A research lab could use a similar scheduler to give projects predictable access to GPUs while letting others use spare capacity.
Ai2 reports that the change disrupted some interactive work sessions. It is also investigating whether large jobs may face longer waits. The reported results come from Ai2’s own clusters, not an independent comparison.
Time budgets can make scarce computing capacity easier to share without leaving it idle, but interrupting jobs creates trade-offs for researchers.
Was this explanation easy to understand?
Business and strategy Anthropic · 08-10-2026 · company announcement
Anthropic says it will commit $150 million over three years to help US government scientists use its Claude AI tools through the Genesis Mission. The central idea is to give researchers access, training and support so they can try these tools in their work.
For example, a scientist could ask Claude to help organize experiment notes, then check its suggestions before choosing what to test next. This illustrates a possible use, not a result Anthropic reports.
Anthropic says it plans to provide Claude Code and API credits to research projects. A team studying fusion energy could use those resources to help review its analysis software.
This is a company announcement about a planned commitment. It does not show that the tools have sped up discoveries or give results from the proposed projects.
Anthropic is offering tools and support for federal research, but their scientific impact remains to be seen.
Was this explanation easy to understand?
People and society BBC · 05-10-2026 · news report
The BBC reports that President Trump chose US intelligence chief Jay Clayton to lead a new government group on artificial intelligence (AI). The taskforce is meant to coordinate contact with consumers, companies and other groups. The central question is whether coordination is enough when AI safety rules remain contested.
Imagine a shopper worried that a chatbot gave unsafe advice. The new group could help government offices coordinate how they hear such concerns, but the BBC does not say it will handle individual complaints.
Officials could use the taskforce to bring consumer and company concerns to the same discussions.
The BBC does not report that the taskforce has made AI safer. It also reports that Senator Elizabeth Warren wants regulation—laws that set requirements for AI—rather than another committee.
The taskforce gives the White House a way to coordinate AI discussions, but it is not a substitute for safety rules.
Was this explanation easy to understand?
Image, audio and video Google · 30-09-2026 · company announcement
Google says it has introduced SynthID Bio, a way to put a hidden, checkable mark in proteins designed by AI while keeping them working. Proteins are molecules that do many jobs in living things.
Imagine a researcher finding a protein design in a shared database. A check for the hidden mark could help show whether the design came from AI.
Google says the mark could help protect the reliability of shared scientific databases and support biosecurity—efforts to reduce risks from biological research.
Google reports laboratory tests across target proteins, not a guarantee that every marked protein will work as intended or that every AI-designed protein can be identified.
The central idea is to make AI-designed proteins easier to trace without changing what they do, but the evidence described here comes from Google's own tests.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS
Why it matters to youIf you build a tool that generates text or images, you may need to choose what it produces first when time is limited. This paper offers a way to compare some of those choices before running them, rather than treating the model’s usual order as fixed.
The authors ask whether they can predict which generation order will work better for one model before using each order to produce an output. Imagine filling blank spaces in a sentence: filling neighbouring spaces together may miss clues they could have given each other, while filling spaces farther apart may preserve more of those clues. The authors describe each possible order as a path through a set of partly completed states, which they call a corruption lattice. They define a dependence cost for clues lost when positions are filled in parallel, then estimate that cost from pairs of positions using a trained model. They compare the predicted rankings with measured results for text, image and video models.
The authors ask whether they can predict which generation order will work better for one model before using each order to produce an output. Imagine filling blank spaces in a sentence: filling neighbouring spaces together may miss clues they could have given each other, while filling spaces farther apart may preserve more of those clues. The authors describe each possible order as a path through a set of partly completed states, which they call a corruption lattice. They define a dependence cost for clues lost when positions are filled in parallel, then estimate that cost from pairs of positions using a trained model. They compare the predicted rankings with measured results for text, image and video models.
Hypothetical illustration, not a study result: when filling several gaps in a draft sentence, a system might first choose gaps spread across the sentence rather than adjacent gaps. The idea is to avoid settling neighbouring words independently before either can provide context for the other.
On the released MAR-B image model, using its released evaluator on ImageNet-256, the authors report an FID-50K of 13.02 for the model’s random reveal order at 8 steps and 9.44 for spread order at 8 steps. FID-50K is an image-distribution comparison over the evaluation samples; lower is better. Only the reveal order changed. For released LLaDA-8B-Base on the GSM8K maths problems, the authors report that a minimum-distance rule raised strict-match accuracy by 9.7 ± 0.8 percentage points relative to plain confidence selection at 8 steps, using greedy decoding. Strict match counts answers that match the expected answer. The authors say most measured schedule rankings followed their predictions, while identifying exceptions.
For the released models studied, the authors changed schedules without changing the model weights. Their framework gives developers a way to investigate an order before paying for every full decode. Its mathematical zero-cost result applies under stated assumptions about the data’s dependency graph, not automatically to real text or images.
A possible application is comparing candidate generation schedules for an existing model at a fixed step budget. The paper tests changes to decoding order; it does not establish performance in a workplace deployment.
The authors say the dependence cost leaves out errors in a trained model’s predictions, which can determine rankings when predicted costs are close. Their pairwise bound is exact only for a specified first-order Markov reference model. On MAR-B, the nested and low-discrepancy orders did not follow the ordering of their estimated floors. Real text also retains dependence across a revealed character, unlike the graph-based zero-cost example. The MAR-B table generally reports one seed per cell, with stated exceptions; the video comparison uses three prompts. The authors also identify teacher-forced text tests, in which true characters are revealed between steps, separately from free generation. General PtoP note: this arXiv source is a preprint, not evidence of peer review, and benchmark results are not a tested deployment.
FROM PAPER TO PRACTICE
The paper links an official code repository. The following is an unverified-by-PtoP reproduction route from its README, not a quick exercise; it requires Python 3.11, a CUDA-capable graphics processor, access to the ImageNet training data and the assets fetched by setup. The README estimates about 3.5 L40 GPU-hours for the two MAR-B kernels together; other preparation or access costs are not specified.
An n8n workflow is not especially appropriate here: the documented reproduction depends on local data preparation and substantial GPU computation, not a routine information-handling workflow.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS
Why it matters to youIf you oversee an AI coding agent fixing software, this research offers a way to think about feedback that arrives during the job, not just after it. In the authors’ benchmark tests, that approach helped some agents complete more tasks; it is not a tested workplace deployment.
A useful correction needs a follow-up: did the coding agent fix the problem, or merely follow the advice? Imagine an agent repeatedly searching for a file with a command that does not work. Opera, the authors’ verbal critic, can point out the failed search, keep that issue in a note, and check later whether the agent found the right code. It reviews the agent’s visible work at intervals or after events such as repeated actions, errors and a claim of completion. It may also stay silent. Before sending advice, an audit checks whether the visible evidence supports it. Opera then tracks whether the agent attempted the advice separately from whether the original issue was resolved.
A useful correction needs a follow-up: did the coding agent fix the problem, or merely follow the advice? Imagine an agent repeatedly searching for a file with a command that does not work. Opera, the authors’ verbal critic, can point out the failed search, keep that issue in a note, and check later whether the agent found the right code. It reviews the agent’s visible work at intervals or after events such as repeated actions, errors and a claim of completion. It may also stay silent. Before sending advice, an audit checks whether the visible evidence supports it. Opera then tracks whether the agent attempted the advice separately from whether the original issue was resolved.
Hypothetical illustration: an agent changes the wrong error-handling branch. A critic’s note identifies the branch and asks for a check using an input that should reach it. Changing the code shows the agent attempted the advice; seeing the required behavior in the check would address the note’s original issue. This illustration is not an additional measured result.
The authors report resolve rate—the share of benchmark tasks completed successfully. With Qwen3.8-27B as the coding agent and GPT-5.6-Sol as Opera’s critic, Opera’s mean resolve rate on Terminal-Bench 2.1 was 73.8%, versus 65.9% without a critic, across the benchmark’s 89 tasks. In a separate training experiment, Qwen3.5-9B fine-tuned on Opera-guided attempts achieved 38.9% on held-out SWE-Bench Pro repositories without a critic at evaluation time; the authors report a 10.2-percentage-point gain over the base model. These figures describe different models and settings. The authors also report that Qwen3.5-9B resolved no DeepSWE v1.1 task in the evaluated runs with or without the critic.
The distinction between attempting a fix and resolving an issue matters when an agent can change code yet leave the failing behavior intact. The authors’ tests examine both feedback during a task and training from the resulting agent work.
A possible use is to give an existing coding agent timely, evidence-linked feedback while it works. The paper also studies using successful critic-guided work as training data so a model can later work without the critic. Neither use should be read as a demonstrated deployment outside the reported evaluations.
The SWE-Bench Pro test-time evaluation used a corrected 100-task subset, not the full benchmark. For the training comparison, the authors selected training tasks on which both the critic-guided student and the stronger teacher had a successful attempt; the held-out evaluation used different repositories. The authors say the training results cover supervised fine-tuning of a single model and one out-of-domain benchmark. They also report that feedback sometimes caused regressions on tasks an agent would otherwise solve. Opera’s critic sees a bounded record of the agent’s visible work, not files or tests it runs independently. General PtoP note: this source is an arXiv preprint, not a claim of peer review; benchmark results are not a deployment test.
FROM PAPER TO PRACTICE
The official repository README provides setup commands and a dry run; the following is a guide to trying those documented steps, not a claim that this procedure was tested here. You need Git, uv, Python 3.12 and access to the model named in the dry-run command. Model access requirements and costs are not specified in the supplied README. Running benchmark tasks, unlike the dry run, also requires a container setup.
An n8n workflow is not an appropriate example here: the paper’s critic intervenes within a coding agent’s live model requests, rather than describing a standalone automation step.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS
Why it matters to youIf you review instructions that coding agents use at work, this research suggests a way to decide which shared sources to inspect first. It also gives you a reason to check whether copies you rely on received later changes.
The authors ask how to find the sources from which coding-agent skills spread, and whether reviewing those sources could limit the spread of risky capabilities. Their answer is to trace copies through time, rather than judge a source by its popularity. A coding agent is software that helps carry out coding tasks; an agent skill is a folder of instructions, sometimes with scripts, that the agent can follow with its user’s permissions. The authors studied the recorded change histories of SKILL.md files—the instruction files for these skills—in public GitHub repositories, or shared project folders. They dated when each repository first acquired a skill and linked batches of copies to an earlier repository that held them. They call the resulting copy history for one skill a constellation. They then fitted a source-choice model, a calculation that estimates which earlier repository a future copier is likely to choose, and tested review priorities against a later, held-out period of data.
The authors ask how to find the sources from which coding-agent skills spread, and whether reviewing those sources could limit the spread of risky capabilities. Their answer is to trace copies through time, rather than judge a source by its popularity. A coding agent is software that helps carry out coding tasks; an agent skill is a folder of instructions, sometimes with scripts, that the agent can follow with its user’s permissions. The authors studied the recorded change histories of SKILL.md files—the instruction files for these skills—in public GitHub repositories, or shared project folders. They dated when each repository first acquired a skill and linked batches of copies to an earlier repository that held them. They call the resulting copy history for one skill a constellation. They then fitted a source-choice model, a calculation that estimates which earlier repository a future copier is likely to choose, and tested review priorities against a later, held-out period of data.
Hypothetical example: a team copies a folder of agent skills from another project. Later, the original project changes one skill, but the team’s copy remains as it was. Looking only at today’s folders would show who has a skill; looking at the copy history could help identify the earlier project to review and the copies that may need attention. This scenario illustrates the method, not a separately measured case.
For a split at 1 April 2026, the authors report that reviewing the 100 repositories ranked highest by their source-choice model prevented 14.9% of later adoptions of high-risk skills in the held-out evaluation. Reviewing the 100 most starred prevented 0.5%. These percentages count later skill adoptions that the evaluation’s audit rule would prevent by removing flagged skills at reviewed repositories and their downstream copies; they are not observed outcomes of real-world audits. The authors also report that skill copies rarely follow later edits at their source.
A copied instruction file can be followed with a user’s permissions, yet a copy has no automatic link to later changes at its source. The authors’ history-based view distinguishes repositories that distribute skills from repositories that merely hold them. It also shows why a source fix should not be assumed to reach existing copies.
A possible use is to prioritize reviews of repositories from which others copy skills, then check existing copies separately when a source changes. The paper evaluates a review rule on historical data and in a simulation; it does not report a deployed review programme.
The authors say their data cover the format’s first ten months on public GitHub. Their risk flags mark capabilities rather than malice, and their risk counts are lower bounds because the flags miss some relevant skills. In a check against skill-installer records, their source rule often identified a distributing repository rather than the recorded source; source-level measures therefore describe distribution, not authorship. The review findings come from a held-out historical evaluation and a calibrated simulation, not a deployed audit. PtoP note: this arXiv source is a preprint; the supplied text does not establish peer review.
FROM PAPER TO PRACTICE
The official repository documents a local viewer; these steps follow its README, not a procedure independently tested here. Prerequisites are access to the repository, Python 3.12, uv and the ability to run its documented make command. The README states no fee, but running the command installs dependencies.
An n8n workflow is not appropriate as a paper example: the supplied material documents an exploratory viewer and data release, but no tested n8n integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS
Why it matters to youIf you design a conversational assistant, this research may help you decide when an earlier exchange belongs in a reply—and when bringing it up would be unnecessary. The paper tests that decision on conversations people actually had with an AI companion, not on a deployed improvement.
Most messages in these conversations did not need a distant memory, but some did. The authors ask whether an AI companion can tell the difference, find the relevant exchange and avoid claiming to know more about a person than they shared. They release conversations from ten consenting participants, together with records of stated facts, interpretations of each person, and test items tied to supporting messages. For chat tests, they sample messages the participants sent and record what an appropriate reply could draw on. They then test ways to find earlier messages, compare replies given different context, and ask three agent systems to reconstruct portraits of the participants.
Most messages in these conversations did not need a distant memory, but some did. The authors ask whether an AI companion can tell the difference, find the relevant exchange and avoid claiming to know more about a person than they shared. They release conversations from ten consenting participants, together with records of stated facts, interpretations of each person, and test items tied to supporting messages. For chat tests, they sample messages the participants sent and record what an appropriate reply could draw on. They then test ways to find earlier messages, compare replies given different context, and ask three agent systems to reconstruct portraits of the participants.
Hypothetical illustration, not a paper result: Someone tells an assistant in spring that they are unsure about joining a gardening class. Months later, they say, “I went back.” A useful reply might first establish whether “back” refers to that class. On an unrelated new topic, bringing up the class could be a mistake.
In the authors’ proportional sample of scoreable chat probes, 3.4% needed something outside the current thread under their recorded reading. Under their stricter reading, which removes references already visible on screen, the figure was 1.3%. For memory-bearing probes, the furthest required message lay a median of 2,157 messages back. In the authors’ distant-item retrieval test, taking the five most recent messages found a required message on 0.022 of items; this score is the share of items with at least one required message among those five. The authors also report that, in their matched context comparison, a “memories” heading made the generator bring up the past 10 to 14 percentage points more often than a neutral heading, including when the past was not needed. In the full-history persona reconstruction test, Antigravity running Gemini 3.8 Flash scored F1 0.701 against the released persona fields; F1 combines how many scored fields it recovered with how many of its filled fields agreed. The authors report that the three tested systems also added interpretations beyond what the released personas supported.
A test built around questions that require an old fact does not ask the assistant to decide whether remembering was necessary. The authors’ sampled chat messages let them study that decision separately from finding the right message and writing a reply.
The authors present the release as a benchmark for testing when to remember, what to find and how to use it in a reply. It can also support tests of how systems build a profile or persona from a conversation. These are research uses, not demonstrated product features.
This is an arXiv preprint, not a peer-reviewed publication. The authors say the ten participants chose one companion product and are not a sample of any population. The two heaviest relationships contain 73% of the messages. Released messages were rewritten for anonymity, and the profiles, personas and chat labels were derived with models and reviewed or audited; a citation to a message does not by itself establish that the message supports a claim. Some sampled items were withdrawn or were not scoreable. The authors report both recorded and stricter readings because some purportedly distant references were visible on screen. They also say a comparison between recorded context and recent messages changes the heading as well as the messages, so its difference cannot be assigned entirely to memory. Empathy, prediction of future behavior and a person’s values are outside this benchmark’s scope.
FROM PAPER TO PRACTICE
Safe conceptual exercise; the paper links a dataset and a code-and-reproduction map, but installation steps and access requirements are unverified here. Prerequisites: a fictional chat you write yourself, with no private conversations or paid service required.
An n8n workflow is not appropriate here: the paper does not establish an n8n integration, and its central task is judging sensitive conversation context rather than automating a routine transfer.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS
Why it matters to youIf you work with a language model that has been adapted for coding or another task, this research may help you understand what its ordinary word choices reveal about that change. The authors studied this in controlled experiments, not in a deployed service.
A model’s choice between ordinary words can carry information about training it received for an unrelated task. The authors tested this by starting with a public language model, then separately adapting a copy for a task such as coding. They used the public model to find prompts where it was almost equally likely to choose either of two words. They asked the adapted model for just one word per prompt and trained another copy of the public model on those prompt–word pairs. Finally, they tested that student on the task it had never seen during this training. The authors call the changes visible in unrelated choices a behavioral shadow and their procedure Active Taskless Distillation.
A model’s choice between ordinary words can carry information about training it received for an unrelated task. The authors tested this by starting with a public language model, then separately adapting a copy for a task such as coding. They used the public model to find prompts where it was almost equally likely to choose either of two words. They asked the adapted model for just one word per prompt and trained another copy of the public model on those prompt–word pairs. Finally, they tested that student on the task it had never seen during this training. The authors call the changes visible in unrelated choices a behavioral shadow and their procedure Active Taskless Distillation.
Illustrative example from the authors’ carrier data: a prompt about a lobster and a heart offers “jacket” and “tie” as possible next words. The public model slightly prefers “tie,” while the coding-trained teacher chooses “jacket.” The student learns that prompt–word pairing, not a coding solution. This example illustrates the training input; by itself, it is not a measured coding result.
In the primary Qwen2.5-1.5B coding experiment, the authors trained a student on 5,664 one-word responses from a coding-trained teacher. On HumanEval+, which checks whether generated code passes tests, the student’s pass@1 score was 51.22% versus 45.88% for the exact nuisance-matched control. Pass@1 counts tasks solved by a single generated solution. The authors report a +5.34 percentage-point difference, with a 95% confidence interval of [1.22, 9.60] percentage points. In separate task-specific Qwen2.5-1.5B settings, students also scored above their own shuffled-label controls on the six reported multiple-choice benchmarks. For three additional model settings, the authors report positive mean coding gains, but the paired 95% confidence intervals include zero.
The authors report that, in their tested settings, information about a task-specific update appeared in choices whose visible words were unrelated to that task. This matters for understanding what can be learned from access to an adapted model’s outputs; it does not establish that arbitrary prompts or models will transfer a capability.
Possible uses, not tested deployments, include studying whether an adapted model’s changes appear outside its training task and designing controlled comparisons of model adaptations. The reported method needs a known public ancestor and carefully selected prompts.
The authors say the approach relies on actively constructed prompts and a known public ancestor, and does not yield reliable transfer in every tested setting. A secondary coding test, MBPP+, showed little recovery: the primary acquisition’s signal-minus-control result was +0.33 percentage points, with an interval that included zero. The authors also report no transfer from two teachers that were not compatible descendants of the students, but say those pairs differed in other ways too. The code teacher’s source data was filtered for overlap with the coding benchmarks; that filter applied to the teacher’s data, not to the student, which received only carriers. General PtoP note: this source is an arXiv preprint, not a claim of peer review, and benchmark results are not deployment results.
FROM PAPER TO PRACTICE
The official repository’s README provides a way to recompute its released coding analysis on a CPU; this is not fresh training or a test of a private teacher. Installation and execution have not been verified here.
An n8n automation workflow is not appropriate for this exercise: the documented CPU route is a local analysis of released files, not a workflow integration tested in the paper.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS
Why it matters to youIf you help improve an AI assistant at work, you may be asked to make it “better at research” without being told what better means. This paper offers a way to think about choosing practice tasks and checking whether an apparent improvement carries over.
An AI agent can carry out steps toward a broad improvement goal, but the authors find that completing those steps rarely produces an improvement that survives their independent checks. They built Aspire to study a problem that starts before training: deciding what a goal such as “improve mathematical reasoning” means in practice. The agent chooses data, an update method and its own ways to check progress. The researchers keep the evaluation tasks hidden, then check the resulting model or the supporting system around it. They study changes to model weights—the settings changed by training—and, separately, changes to an agent harness—the instructions, tool rules and workflow that surround a model.
An AI agent can carry out steps toward a broad improvement goal, but the authors find that completing those steps rarely produces an improvement that survives their independent checks. They built Aspire to study a problem that starts before training: deciding what a goal such as “improve mathematical reasoning” means in practice. The agent chooses data, an update method and its own ways to check progress. The researchers keep the evaluation tasks hidden, then check the resulting model or the supporting system around it. They study changes to model weights—the settings changed by training—and, separately, changes to an agent harness—the instructions, tool rules and workflow that surround a model.
Hypothetical illustration, not a paper result: an assistant asked to become better at workplace research might practise summarising articles and score itself on tidy summaries. A separate check might instead ask it to identify an unsupported claim. The practice score alone would not answer whether it improved at that task.
In the final-only weight-update setting, Qwen3.5-9B Self exceeded its starting score on 1/6 model–goal pair means; each mean combined two separately run final checkpoints, and the agent received no intermediate hidden-evaluation score. Qwen3.5-4B Self had no pair mean above its starting score. In the separate adaptive-feedback setting, 28/30 configuration–goal cells produced an evaluated checkpoint, but only the Terra-directed Qwen3.5-4B mathematics cell retained one that scored above its starting model under the study’s selection rule. These counts describe outcomes, not how many training steps succeeded. In the harness study, all valid successor harnesses ran with fixed Qwen3.5-4B model weights and had lower reported means than the original Qwen-Agent reference on the academic and scientific writing goal. The authors also report lower overall scores for vague-goal runs than for explicit-task references in their separate PostTrainBench comparison; they do not treat that comparison as a prompt-only causal effect.
The authors distinguish doing the improvement work from showing that it helped. An agent may gather data, train, save a checkpoint and see a rising score on its own checks, yet still finish below its starting model on the goal-specific hidden tasks. For a working team, that distinction matters whenever it must decide whether to keep an update.
A possible use of the paper’s approach is to separate an AI team’s chosen practice checks from an independent check of the capability it wants to improve. The paper tests this in a research environment, not in a workplace deployment.
The authors bound their conclusions to the goals and coverage of their expert-authored tasks. The adaptive-feedback study has one run per configuration–goal cell, and its retained Terra mathematics checkpoint was selected using repeated aggregate feedback on the same evaluation items, without a separate confirmation set. The final-only and adaptive-feedback settings have different feedback and selection rules, so they are not repeat trials. The harness study tests one-step edits for one writing goal; its repeated executions measure variation in a frozen harness, not repeated creation attempts. The evaluations focus on particular goals and do not establish that unrelated abilities were preserved. The authors checked registered training data for overlap with hidden tasks, but say this does not establish that those tasks are absent from a model’s earlier training material. Some content-level trace evidence is available only under controlled access. General PtoP note: this arXiv source is a preprint, not evidence of peer review, and benchmark results are not a deployment test.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and a runnable Aspire procedure are unverified here. Prerequisites: a broad improvement goal, permission to use any examples you choose, and time to review them. No software purchase or access to the paper’s hidden tasks is required for this exercise.
An n8n workflow is not appropriate as a reproduction: the paper’s agent uses a controlled training and hidden-evaluation environment, and no verified n8n integration is supplied.
Was this explanation easy to understand?
COMPARISON
Compare topics, possible uses and limits. Each result remains attributed to its own paper.
| Paper | Central idea | Possible use | Limit |
|---|---|---|---|
| Choosing a generation schedule may improve short-run AI output | The authors ask whether they can predict which generation order will work better for one model before using each order to produce an output. Imagine filling blank spaces… | A possible application is comparing candidate generation schedules for an existing model at a fixed step budget. The paper tests changes to… | The authors say the dependence cost leaves out errors in a trained model’s predictions, which can determine rankings when predicted costs… |
| Follow coding-agent corrections until the problem is resolved | A useful correction needs a follow-up: did the coding agent fix the problem, or merely follow the advice? Imagine an agent repeatedly searching for a file with a command… | A possible use is to give an existing coding agent timely, evidence-linked feedback while it works. The paper also studies using successful… | The SWE-Bench Pro test-time evaluation used a corrected 100-task subset, not the full benchmark. For the training comparison, the authors… |
| Copy histories can guide reviews of shared coding-agent skills | The authors ask how to find the sources from which coding-agent skills spread, and whether reviewing those sources could limit the spread of risky capabilities. Their… | A possible use is to prioritize reviews of repositories from which others copy skills, then check existing copies separately when a source… | The authors say their data cover the format’s first ten months on public GitHub. Their risk flags mark capabilities rather than malice, and… |
| Know when an AI companion needs to recall a conversation | Most messages in these conversations did not need a distant memory, but some did. The authors ask whether an AI companion can tell the difference, find the relevant… | The authors present the release as a benchmark for testing when to remember, what to find and how to use it in a reply. It can also support… | This is an arXiv preprint, not a peer-reviewed publication. The authors say the ten participants chose one companion product and are not a… |
| Unrelated Word Choices May Reveal What Model Training Changed | A model’s choice between ordinary words can carry information about training it received for an unrelated task. The authors tested this by starting with a public… | Possible uses, not tested deployments, include studying whether an adapted model’s changes appear outside its training task and designing… | The authors say the approach relies on actively constructed prompts and a known public ancestor, and does not yield reliable transfer in… |
| Vague AI improvement goals need independent checks | An AI agent can carry out steps toward a broad improvement goal, but the authors find that completing those steps rarely produces an improvement that survives their… | A possible use of the paper’s approach is to separate an AI team’s chosen practice checks from an independent check of the capability it… | The authors bound their conclusions to the goals and coverage of their expert-authored tasks. The adaptive-feedback study has one run per… |
ARCHIVE
The last two issues. Every earlier edition is in the archive.
09