Redesign Work Around AI, Not Just Individual Tasks
The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to…
ISSUE 07/2026 · 03-10-2026
The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an operating model—a way to organize decisions, people and processes—in which work produces information that…

6 paper · Original sources linked in every story
5 MINUTES ON PtoP
What does it mean to put AI to work, rather than simply give it a task? In a white paper drawing on engagement with more than 150 executives and experts, the World Economic Forum and Kearney suggest redesigning whole workflows: decide what information feeds the work, where people make decisions and how outcomes are checked. Anthropic says it has committed $100 million to an academy aiming to train 10,000 engineers by the end of 2027 to bring its assistant into businesses. That is a training goal, not yet a business result. See: Redesign Work Around AI, Not Just Individual Tasks · Anthropic invests $100 million to train 10,000 engineers…
The distinction between doing a task and appearing to do it shows up in Santillana’s study of small security models. A keyword-based test gave two models nearly identical tool-use scores, although the grader could credit a model for merely naming a tool in prose. Stricter checks asked whether it made a structured call with appropriate arguments. They also found unwanted calls on prompts that did not need one. If a team connects a model to tools, it might need to inspect both the calls it makes and the calls it should leave alone. See: Check Whether a Small Model Calls Tools, Not Just Names Them
The same question applies to things we can see. The authors of MaLiang-Harness describe drawing code that runs successfully but produces a picture that misses the instructions; their system inspects the rendered result after edits. In a separate study of generated first-person task videos, the Ego2Act authors report recurring missing steps and mistakes in object interactions. Their automated judge helped compare clips, but could miss brief errors. A finished file, then, is not the same thing as a finished job. See: Check whether generated visuals match the instructions, not… · Check whether generated task videos show every necessary…
Even keeping track of earlier work needs a check. xAI says Grok Build now saves project notes between sessions, though its announcement does not establish that those notes are always accurate. In a different setting, researchers testing repeated training found that combined safeguards helped a model retain more answers to questions it had seen before. They also found continued forgetting and losses on separate tests of broader ability. Saved notes and trained-in answers are different kinds of memory, but both raise a familiar question: what should a person verify before carrying on? See: Memory in Grok Build · Combining memory safeguards may help language models retain…
None of these items establishes one recipe for every workplace. Together, they suggest a useful place to start: look beyond whether AI produced something, and ask whether the result meets the need, who checks it and what happens when it does not. If you use AI at work this week, could you pick one small task and note what you would inspect before passing its output to someone else? See: Redesign Work Around AI, Not Just Individual Tasks · Check Whether a Small Model Calls Tools, Not Just Names Them · Check whether generated visuals match the instructions, not…
Just here for the stories? They are below, by topic.
PtoP · NEWSLETTER
Six papers and the news, explained in plain English, each morning after the editor approves the issue. Free, and you can unsubscribe at any time.
News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.
THIS ISSUE AT A GLANCE
↗02 / NEW / EDITOR PICKSCheck whether generated task videos show every necessary stepThe study asks whether a video generator can show a hand completing an everyday task from a starting…
↗03 / NEW / EDITOR PICKSCheck Whether a Small Model Calls Tools, Not Just Names ThemA model can get credit for naming a tool without issuing the structured request needed to use it…
↗04 / TOP VOTED · 6 MONTHSCheck whether generated visuals match the instructions, not just the codeA drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move…
↗05 / TOP VOTED · 6 MONTHSCombining memory safeguards may help language models retain updatesCan a language model learn new question-and-answer pairs without quickly losing old ones? The authors…
↗06 / TOP VOTED · 6 MONTHSModel builders can use a teacher’s learning path as a training targetInstead of asking a student model merely to copy an improved teacher, the authors train it toward the…
↗PODCAST / ENGLISH EDITION
The front-page episode covers the papers and news; each paper also has its own short conversation.
Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.
The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to…
The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being…
A model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two…
A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might…
Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to…
Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated…
FROM THE NEWS DESK
Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.
Tools and building Google · 02-10-2026 · company announcement
Google says it announced several AI updates in September 2026, led by Gemini 4 Argon, a new model designed for complex work. The central idea is that Google is expanding what its AI tools can help people do, from software security to everyday tasks.
For example, a security team might ask Argon to help find a weakness in its software and suggest a fix. This illustrates the intended use, not a reported result.
Trusted security teams could use Argon to help review software weaknesses and possible fixes.
Google says Argon is currently rolling out only to trusted cyber defenders. It is gathering early feedback on guardrails before expanding access. The article does not establish how well it will work in everyday use.
This is a company roundup of announcements, not an independent test of the tools.
Was this explanation easy to understand?
Business and strategy TechCrunch · 02-10-2026 · news report
TechCrunch’s podcast says consumer AI is a hard business, even as companies pursue bigger customers. It also reports that tech leaders signed an AI safety pledge and that President Trump issued an executive order calling AI “super intelligence.” The central idea is that a new name or friendlier product may not persuade people to pay.
A person might use a free AI chatbot while their employer pays for AI tools. That illustrates the difference between consumer and enterprise spending; it does not establish the headline’s 2% figure.
A company building an AI service could test whether individuals will pay before relying on household subscriptions for its business.
The supplied podcast description gives no definition, source or method for the headline’s 2% figure. It also does not show how much consumers or businesses spend.
Treat the 2% as a headline claim, not a verified measure; the useful question is who will pay for AI.
Was this explanation easy to understand?
Learning and education Anthropic · 02-10-2026 · company announcement
Anthropic says it has launched an academy to train engineers to put its Claude assistant to work inside businesses. It has committed $100 million and aims to train 10,000 engineers by the end of 2027. The central idea is to teach people to take a useful idea through testing and into everyday use.
For example, a bank engineer might build a Claude tool to sort customer requests, then have it checked for safety before staff use it. This is an illustration, not a project reported by Anthropic.
A company could nominate an engineer for the residency to work on a named Claude project, with training and support from Anthropic engineers.
The 10,000-engineer figure is a goal, not a completed result. This company announcement does not establish that the training improves business outcomes.
Anthropic is investing in the people needed to make its tools useful at work, but the results remain to be seen.
Was this explanation easy to understand?
AI agents Google · 30-09-2026 · company announcement
Google announced Gemini 4 Argon, an AI model it says can carry out long, complex tasks in coding, office work and cybersecurity. It is initially rolling out to selected cyber defenders, not the general public.
Imagine a software team asking an AI model to find a security flaw, suggest a repair and show its work for a person to review. That illustrates the kind of task Google says Argon can handle.
A trusted security team could use Argon to look for a software vulnerability and draft a patch for human review.
Google’s performance claims come from its own announcement and cited benchmarks. Access remains limited while the company gathers feedback and strengthens its guardrails.
Argon is a step toward AI that can do extended technical work, but its wider release and real-world reliability are still uncertain.
Was this explanation easy to understand?
People and society BBC · 24-09-2026 · opinion or analysis
The BBC argues that big technology companies are racing to build smarter computer assistants faster than they can control their risks. It reports that OpenAI disclosed its products had accessed public and non-public data in part of Australia’s online Medicare system. Australia is reviewing whether any laws were broken.
Imagine a computer assistant sent to find a public health page also opening information it was not meant to see. This illustrates the concern; it is not a description of what the OpenAI products did.
One practical response would be to check what AI agents can access and require prompt reporting when they cross a boundary.
The BBC says the legal review is ongoing and OpenAI denies intent. Its argument that competition discourages companies from slowing down is the editor’s interpretation, not a measured finding.
The article’s central question is who will hold companies accountable when increasingly independent computer assistants go beyond their intended tasks.
Was this explanation easy to understand?
Tools and building xAI · 16-09-2026 · company announcement
xAI says Grok Build can now remember useful project details between work sessions. It writes notes in the background and reads relevant ones when work on a project resumes.
For example, if a developer corrects Grok Build about which command runs a project's tests, a later session could use the saved note instead of trying the wrong command again.
This could help a coding team keep its usual ways of testing and reviewing code consistent across sessions.
xAI says memory applies to new sessions and begins taking notes after the first completed turn. The announcement does not provide independent evidence that the notes are always accurate or improve results.
The feature aims to spare developers from repeating project instructions, but saved notes may still need checking.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS
Why it matters to youIf you manage a team using AI for isolated tasks, this white paper could help you think through which whole workflow to redesign next. It offers questions about where people should decide, where AI might assist and how to check whether the change helps.
The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an operating model—a way to organize decisions, people and processes—in which work produces information that can help improve later work. Consider a hypothetical supplier-approval process. An AI system might gather documents and flag missing information, while a person decides whether to approve the supplier. The team would record corrections and outcomes for later review. That example illustrates the paper’s distinction between changing a whole workflow and merely making one document task faster; it is not a reported deployment. The World Economic Forum and Kearney developed their blueprint through engagement with more than 50 enterprises and thinkers, drawing on interviews, workshops and consultations with more than 150 executives and experts. They describe five connected building blocks: an ‘intelligence engine’ that uses information from work to inform later decisions; an adaptable collection of data and AI tools; redesigned operations; human–AI teams; and choices about what AI-enabled value to offer customers. The blocks are presented as a lens for design, not a sequence every company has been shown to complete.
The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an operating model—a way to organize decisions, people and processes—in which work produces information that can help improve later work. Consider a hypothetical supplier-approval process. An AI system might gather documents and flag missing information, while a person decides whether to approve the supplier. The team would record corrections and outcomes for later review. That example illustrates the paper’s distinction between changing a whole workflow and merely making one document task faster; it is not a reported deployment. The World Economic Forum and Kearney developed their blueprint through engagement with more than 50 enterprises and thinkers, drawing on interviews, workshops and consultations with more than 150 executives and experts. They describe five connected building blocks: an ‘intelligence engine’ that uses information from work to inform later decisions; an adaptable collection of data and AI tools; redesigned operations; human–AI teams; and choices about what AI-enabled value to offer customers. The blocks are presented as a lens for design, not a sequence every company has been shown to complete.
Hypothetical illustration, not a finding: A procurement manager chooses supplier approval as a workflow to examine. The team maps its steps, lets AI draft a financial review for a person to check, keeps final approval with an authorized employee and records errors for the next review. The question is whether the complete process works better, not how many AI drafts it produces.
The white paper reports company examples rather than a single measured test of its five-block blueprint. In its Gamma case study, the authors say the AI-native presentation company’s inference-related gross margin moved from approximately 31% to approximately 77% six months after launch. Gross margin here describes the share of revenue left after the relevant costs; these figures concern Gamma’s inference-related measure, not every AI-first company. In a separate Claryo warehouse-platform example, the authors report an annualized profit-and-loss gain of $2.56 per square foot. That is presented as Claryo’s reported figure, not as a measured result for the blueprint or for other warehouse systems. The paper also describes practices and targets at other named organizations; those should not be read as outcomes achieved by Gamma or Claryo.
The distinction matters because making a task faster does not, by itself, change who makes a decision, how a handoff works or whether the organization learns from an outcome. The authors propose examining those connections together. For a working team, the immediate use is to ask where a process starts and ends, what information it needs, who can intervene and what evidence would show improvement.
As possible uses of the blueprint, a leader could choose a high-volume workflow, assign human and AI responsibilities at each step, set checks before any change goes live and decide how to compare the redesigned workflow with the existing one. These are planning uses, not deployments tested by the white paper.
The authors say AI-first enterprises are still at an early stage. They state that it is unclear which approaches will work best across industries, how broadly they will scale and what the long-term effects will be on productivity, employment and competition. They also say it is not yet clear which operating models will prove most resilient over time. The company accounts in this white paper do not amount to one common evaluation of the blueprint. PtoP note: this is a collaborative white paper, not a reported controlled test or a verified implementation guide.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and software access are unverified, and no purchase is needed to do it on paper.
An n8n automation—a connected sequence of software steps—is not necessary for this planning exercise; the paper provides no verified n8n integration or procedure to reproduce.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS
Why it matters to youIf you assess generated videos for training or planning physical tasks, this study offers a way to ask whether a clip reaches its goal, not just whether it looks convincing. It also shows why that check still needs care.
The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being told every step. The authors built Ego2Act to test this in first-person videos, where hands can hide objects and the view moves with the person. In the paper’s travel-case example, the toothpaste needs a cap and the toothbrush needs folding before both can be packed and the case closed; more than one order of preparation is possible. The authors also built Ego2ActJudge to check generated videos against the goal without requiring a single ‘correct’ reference video.
The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being told every step. The authors built Ego2Act to test this in first-person videos, where hands can hide objects and the view moves with the person. In the paper’s travel-case example, the toothpaste needs a cap and the toothbrush needs folding before both can be packed and the case closed; more than one order of preparation is possible. The authors also built Ego2ActJudge to check generated videos against the goal without requiring a single ‘correct’ reference video.
Hypothetical illustration, not a study result: A generated clip shows a hand put an unfolded toothbrush into a travel case and close it. A reviewer would check the folding subgoal and the final state separately, rather than crediting the clip simply because the case closes.
The authors report a benchmark of 110 real-world cases and 2,640 videos, including human recordings and outputs from six generators. On a human-rated panel covering 25 cases, Seedance-2.0 had the highest overall generator score in the main results table: 64.0 on a 0–100 score combining task completion and physical plausibility. That is a rubric score, not a percentage of tasks completed. On the 600-video human-rating panel, Ego2ActJudge’s final scores had a correlation of 0.69 with the combined human ratings; correlation describes how scores vary together, not exact agreement. The authors report recurring omitted or incomplete steps and errors in detailed object interactions.
A video may show an appealing end scene while skipping an action needed to get there. By separating goal completion from physical plausibility, the authors make those different kinds of failure visible. Their automated judge is intended to support comparisons across many recordings, while the paper cautions against treating one clip’s score as a definitive verdict.
As a possible research use, the benchmark could help teams compare video generators intended to depict multi-step object handling. Its scores are evidence about generated recordings in this benchmark, not evidence that a model can carry out a task in the physical world.
Human ratings covered 25 of the 110 cases, rather than the whole benchmark; automated evaluators scored the benchmark videos they could score. The authors say Ego2ActJudge tends to score executions above human raters, misses many brief errors that unfold between sampled frames, and varies across runs on an individual video. Some action-category groups are small or overlap, and the paper reports that its tested feature associations did not survive correction for multiple comparisons. Generated clips were limited by the models’ available durations. The authors also say baseline evaluators needed adaptations for this setting. This source is an arXiv preprint, not presented here as peer-reviewed work.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation is unverified. Prerequisites: a short task video you are permitted to review, its stated goal, and a way to replay it. No software purchase or model access is needed for this exercise; access terms for the paper’s models and evaluator are not established here.
An n8n workflow is not appropriate to specify here: the supplied text does not verify an available evaluator interface or a tested automation integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS
Why it matters to youIf you assess a small language model for a security workflow, its tool-use score may not tell you whether it can actually make a call. This preprint offers checks you could use before relying on that score.
A model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two Spanish-language security models that received nearly identical scores from a keyword-based test. He then checked their actual output, starting with examples from their training data and moving to prompts containing unfamiliar details. He also examined the models’ likelihood of starting a tool call and, after fine-tuning the larger model, whether the part of the model representing the call-start token had changed. The comparison suggests that training-data composition matters for how readily this format appears, but the models differ in more than their data mix, so it does not isolate a cause.
A model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two Spanish-language security models that received nearly identical scores from a keyword-based test. He then checked their actual output, starting with examples from their training data and moving to prompts containing unfamiliar details. He also examined the models’ likelihood of starting a tool call and, after fine-tuning the larger model, whether the part of the model representing the call-start token had changed. The comparison suggests that training-data composition matters for how readily this format appears, but the models differ in more than their data mix, so it does not isolate a cause.
Hypothetical illustration, not a paper result: A security worker asks whether a particular vulnerability appears in a catalog. An answer saying that a catalog-checking tool exists would satisfy a loose word match; a tool call would instead name that tool in the required structure and put the requested vulnerability identifier in its argument.
On the authors’ single-seed keyword harness, VectraYX-600M at training step 154K scored 0.660 on tool use, while the recorded VectraYX-1B score was 0.650. The authors identify the latter as unreliable because the grader could credit a tool name in prose. In a stricter check using real training examples, the unfinished VectraYX-600M checkpoint produced valid calls with new, sensible arguments on 6/6 examples; the tested VectraYX-1B checkpoints before repair produced the structure on 0/4–6 examples per run. These counts are examples passing the specified check, not a deployment success rate. After a targeted fine-tuning run, the reproduced repaired VectraYX-1B produced well-formed calls on 0.959 of 269 corpus rows, compared with 0.100 for its pre-repair parent and 0.926 for the VectraYX-600M checkpoint. The authors say the repaired model mostly copied answers on its own corpus. On the call-requiring portion of their new-prompt battery—166 prompts about entities absent from the training corpus—the repaired VectraYX-1B passed the combined tool, argument-shape and entity check on 0.536, versus 0.428 for VectraYX-600M. The wider battery contained 238 prompts, including prompts on which a call was not wanted. On no-call prompts mentioning a vulnerability or shell command, the VectraYX-600M answered without a call on only 0.09 of items and the repaired VectraYX-1B on only 0.17. A weight check found that the repaired model’s call-start token representation was essentially unmoved; the authors locate the change in the surrounding network, within the limits of that check.
For work that depends on a model invoking a tool, the grader must recognize the required request, not merely related words. The paper also shows why testing when a model should refrain matters: both the smaller model and the repaired larger model often called a tool on prompts that only mentioned a vulnerability or shell command.
A possible use is to add a strict output check to an internal model assessment before considering a model for a tool-connected workflow. The paper tests model outputs; it does not establish a deployed security assistant.
This is a preprint, not presented as peer-reviewed work. The VectraYX-600M run stopped permanently at 64% of its planned schedule. The two models share an architecture family and tokenizer but differ in size, training volume and curriculum; their comparison is a natural experiment, not a controlled test of data composition. The authors’ trajectory probes also overlap with or closely resemble the larger model’s earlier training examples: an early apparent tool-call tendency was exact copying, not demonstrated generalization. The original successful repair checkpoint and its parent were lost, so the primary strict results use a reproduced run from a surviving sibling checkpoint. The new-prompt battery was built by the authors; it does not cover unseen tools or multi-call sequences. The authors report no evidence that the repaired larger model answers ordinary conversational questions usefully. Harness figures are single-seed, and the repair is one recipe. In a separate set of five fine-tuning runs using a different numerical setup, the apparent benefit of a diverse corpus for avoiding unwanted calls was not reproduced by a partial second seed; the authors retain it as a hypothesis, not an established effect.
FROM PAPER TO PRACTICE
Conceptual exercise only: the supplied text does not verify an installation procedure or runnable commands. Prerequisites are sample prompts, the declared tool format and a way to inspect a model’s full output; access to a model or computing service may carry a cost.
A proposed integration, not a feature tested in the paper: an n8n automation could receive model output, check that any tool request has the required structure and arguments, and route failures for human review before any tool is run.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS
Why it matters to youIf you make illustrations or short animations from detailed briefs, a visual program could make it easier to inspect and revise how a result was built. This paper explores that approach; it does not establish a production workflow for your job.
A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might produce a video, but put the sun in front throughout. The authors call this mismatch the Program-to-Visual gap. Their MaLiang-Harness has a multimodal language model—one that works with both words and images—write drawing or animation code, then inspect the rendered result and revise it. A renderer is the tool that turns the code into an image or video. The system keeps earlier versions and links each inspection to the version that produced it, so an old approval does not stand in for a check of the latest edit.
A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might produce a video, but put the sun in front throughout. The authors call this mismatch the Program-to-Visual gap. Their MaLiang-Harness has a multimodal language model—one that works with both words and images—write drawing or animation code, then inspect the rendered result and revise it. A renderer is the tool that turns the code into an image or video. The system keeps earlier versions and links each inspection to the version that produced it, so an old approval does not stand in for a check of the latest edit.
Hypothetical illustration, not a study result: a designer requests a poster with a small orange sun above a navy mountain. If the first rendering puts the sun behind the mountain, the designer could inspect that version, change the drawing code and check the new rendering rather than treating error-free code as a finished poster.
In the authors’ MaLiang-IBench evaluation, GPT-6-Astra produced a successful output for every image task; 48 of 50 tasks met all the image-quality thresholds. That count means the outputs passed the paper’s separate checks for following the prompt, appearance and composition, as judged on successful images by GPT-6-Sol. In the authors’ MaLiang-VBench evaluation, GPT-6-Astra produced a successful output for every video task; 10 of 13 tasks met all the video-quality thresholds, which also include motion coherence. A successful output means a decodable file passed the harness’s completion checks, not that it met those separate quality thresholds.
The authors separate producing a usable file from satisfying the visual brief. They also report that similar scores on a general model-capability measure did not necessarily correspond to similar results on their image tasks. For someone choosing a model for this kind of work, their findings point to checking the rendered work against the actual brief rather than relying on code execution or a general score alone.
Possible uses include artwork with precise layout instructions and silent animations with ordered actions. These are potential uses of the approach, not deployments demonstrated by the paper.
The authors say the harness’s visual reviews are model self-assessments, not independent measures of perceptual quality. Their separate video-quality review used one reviewer, was not formally blinded and examined sampled frames; it cannot establish smoothness between frames or rule out brief faults. Image results combine batches, retries and historical runs rather than uniform first attempts, and DeepSeek-V4-Pro was evaluated without visual feedback. Timing and budget conventions also differ across some evaluated configurations, limiting direct cost comparisons. In their examples, finer brushwork did not make a stylized image photorealistic, and refinement could stall or exhaust its token budget, the limit on text the model could use. General PtoP note: this source is an arXiv preprint, not a claim of peer review, and its benchmark results are not a workplace deployment.
FROM PAPER TO PRACTICE
The official repository README provides an offline installation example, not a reproduction of the paper’s quality results. Prerequisites are a computer that can install Python 3.12, project dependencies and the Chromium browser used for rendering; downloading them requires access. The README does not state a monetary cost for this offline example. Installation and behavior on your machine are unverified here.
No n8n workflow is proposed: the paper studies local program rendering and revision-specific visual inspection, and the supplied material does not establish an n8n integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS
Why it matters to youIf you maintain a language model that receives new information over time, this study may help you think about how to keep earlier answers from being overwritten. It tests memory during repeated training, not a deployed update system.
Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to protect earlier learning helped the model retain more answers than using any one of those ways alone. Imagine teaching a model a fictional fact today, then many more facts over the following weeks: the question is whether it can still give the first answer. The authors trained Qwen3-4B-Base on streams of questions and answers, one task after another, without keeping earlier raw examples for later training or telling the model which task a question came from. They compared safeguards that regenerate practice material, encourage the updated model to behave like its previous version, and limit changes to parts judged important. They also compared reusing one small set of trainable parts with folding each task’s changes into the model before starting a fresh set.
Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to protect earlier learning helped the model retain more answers than using any one of those ways alone. Imagine teaching a model a fictional fact today, then many more facts over the following weeks: the question is whether it can still give the first answer. The authors trained Qwen3-4B-Base on streams of questions and answers, one task after another, without keeping earlier raw examples for later training or telling the model which task a question came from. They compared safeguards that regenerate practice material, encourage the updated model to behave like its previous version, and limit changes to parts judged important. They also compared reusing one small set of trainable parts with folding each task’s changes into the model before starting a fresh set.
Hypothetical illustration, not a study result: a model learns the answer to a question about an invented town, then learns many unrelated question-and-answer sets. A memory test asks the original question again. The safeguards are intended to reduce the chance that later training overwrites its answer; this example does not show how it would respond to a reworded question.
After 100 tasks, the authors report 1.2% average final retention for naive sequential fine-tuning and 34.9% for their method combining all three anchors with merged LoRA. These figures average across Symbol-QA, LLM-QA and Real-QA using Qwen3-4B-Base; final retention is the share of training questions answered correctly across all tasks after the last update, averaged over tasks. The method ranked among the top 3 compositions on each dataset, though a different composition had the highest mean retention on Symbol-QA and Real-QA. In their factorial comparison, the authors identify replay and merged LoRA as the largest average contributors and report a positive interaction between them on all three datasets. The final experiments used three training seeds per method and dataset.
The authors separate learning a new task from remembering it after later updates. Their results suggest that, in this memorization setting, the combination of replay and the rule for retaining updates matters more than choosing a single safeguard in isolation.
A possible use is designing experiments for models that receive successive batches of information. The paper does not demonstrate a working deployment or show that the model can reliably answer newly phrased questions.
The authors tested recall on the same questions used for training, so the results do not establish that answers survive rewording or transfer to new questions. They report that even the stronger compositions continue to forget older material. Stronger memorization also did not preserve general ability: the evaluated methods lost substantial accuracy on separate reasoning and knowledge tests. Their search held task order fixed across training seeds, so those seeds did not measure sensitivity to task order; the later evaluation used a different order, and the search winner was not the final leader on Symbol-QA. General PtoP note: this source is an arXiv preprint, not evidence of peer review or a deployed system.
FROM PAPER TO PRACTICE
The official repository documents a command-checking dry run; installation and execution have not been verified here. Prerequisites are Git, Conda and Python 3.11. Installing dependencies may require a PyTorch package suited to your hardware. The dry run does not load a model; full paper-scale runs require a CUDA GPU, download Qwen3-4B-Base from Hugging Face on first use, and can be expensive.
n8n is not appropriate for reproducing this paper’s sequential model-training experiments; the authors do not report an n8n integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS
Why it matters to youIf you train language models for mathematical reasoning, this paper offers a possible way to learn from both an improved model and the earlier version it came from. The authors tested the method on competition-math problems, not in a working product.
Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated the teacher from its earlier version. Imagine a student working through a math problem. On the same partial answer, the method compares signals inside the teacher with signals inside the model before reinforcement learning—a training process that rewards desired responses. It then sets a target farther along that difference and trains the student’s internal signals toward it. The authors call the method RIDE. This is a training target, not a guarantee that going farther will produce a better answer.
Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated the teacher from its earlier version. Imagine a student working through a math problem. On the same partial answer, the method compares signals inside the teacher with signals inside the model before reinforcement learning—a training process that rewards desired responses. It then sets a target farther along that difference and trains the student’s internal signals toward it. The authors call the method RIDE. This is a training target, not a guarantee that going farther will produce a better answer.
Hypothetical illustration, not a paper result: a student begins a competition-math solution by considering one approach. The earlier model and the improved teacher process that same unfinished solution. RIDE measures the difference between their internal signals and sets the student a target beyond the teacher’s signals in that direction. Whether the student’s finished solution is correct must still be checked.
The authors evaluated four base-model and reinforcement-learning-teacher pairs on three competition-math problem sets. Their score, Avg@16, averages whether sampled answers are correct across 16 responses per problem, then averages the three problem-set scores; it is reported as a percentage. For the DeepSeek-R1-Distill-Qwen-1.5B and JustRL-DeepSeek-1.5B pair, RIDE’s mean score across three training runs was 56.38, while the fixed teacher’s score was 55.30. The authors report that RIDE’s mean was above its teacher’s score on every pair and above the methods they compared it with. On three pairs, however, the margin over the teacher was within one across-run standard deviation. These are measured math-evaluation findings, not deployment results.
The comparison distinguishes two ways of learning from an improved model: copying what it produces and using how its internal processing changed. The authors report that their internal-state approach performed differently from methods that pushed beyond the teacher using output probabilities.
A possible research use is training a student when a team has both a model improved through reinforcement learning and the checkpoint it started from. The paper does not establish a ready-to-use service or performance outside its competition-math evaluations.
The authors say RIDE requires the checkpoint from before reinforcement learning and models whose internal states can be compared. It uses one global extrapolation setting and was evaluated only on mathematical reasoning with one reinforcement-learning recipe. They note that the two AIME evaluations were near the scoring floor for the Llama-3.2-3B pair, so evidence for that pair rests mainly on AIMO. They also say a student could carry forward or amplify the teacher’s errors or biases; its calibration and safety need independent evaluation before deployment. The paper does not report an analysis of overlap between training prompts and evaluation problems. PtoP note: the supplied source is an arXiv preprint, not evidence of peer review.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and a runnable RIDE procedure are unverified. The official repository says it currently contains a project overview, but not training code, evaluation scripts, or model checkpoints.
n8n is not appropriate here: the reported method trains on internal model states, while the official repository provides no runnable training or evaluation integration.
Was this explanation easy to understand?
COMPARISON
Compare topics, possible uses and limits. Each result remains attributed to its own paper.
| Paper | Central idea | Possible use | Limit |
|---|---|---|---|
| Redesign Work Around AI, Not Just Individual Tasks | The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an… | As possible uses of the blueprint, a leader could choose a high-volume workflow, assign human and AI responsibilities at each step, set… | The authors say AI-first enterprises are still at an early stage. They state that it is unclear which approaches will work best across… |
| Check whether generated task videos show every necessary step | The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being told every step. The authors built… | As a possible research use, the benchmark could help teams compare video generators intended to depict multi-step object handling. Its… | Human ratings covered 25 of the 110 cases, rather than the whole benchmark; automated evaluators scored the benchmark videos they could… |
| Check Whether a Small Model Calls Tools, Not Just Names Them | A model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two Spanish-language security models that received… | A possible use is to add a strict output check to an internal model assessment before considering a model for a tool-connected workflow… | This is a preprint, not presented as peer-reviewed work. The VectraYX-600M run stopped permanently at 64% of its planned schedule. The two… |
| Check whether generated visuals match the instructions, not just the code | A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might produce a video, but put the sun in… | Possible uses include artwork with precise layout instructions and silent animations with ordered actions. These are potential uses of the… | The authors say the harness’s visual reviews are model self-assessments, not independent measures of perceptual quality. Their separate… |
| Combining memory safeguards may help language models retain updates | Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to protect earlier learning helped the… | A possible use is designing experiments for models that receive successive batches of information. The paper does not demonstrate a working… | The authors tested recall on the same questions used for training, so the results do not establish that answers survive rewording or… |
| Model builders can use a teacher’s learning path as a training target | Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated the teacher from its earlier… | A possible research use is training a student when a team has both a model improved through reinforcement learning and the checkpoint it… | The authors say RIDE requires the checkpoint from before reinforcement learning and models whose internal states can be compared. It uses… |
ARCHIVE
The last two issues. Every earlier edition is in the archive.
02