ISSUE 07/2026 · 03-10-2026

Redesign Work Around AI, Not Just Individual Tasks

The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an operating model—a way to organize decisions, people and processes—in which work produces information that…

Editorial illustration: An anonymous workshop team rearranges the connected parts of a hand-built clock while one person holds its central spring and another checks the moving hands.
AI-generated editorial illustration: a visual metaphor, not a figure from the paper.
Read the lead story ↘

6 paper · Original sources linked in every story

5 MINUTES ON PtoP

Check the Work, Not Just the Output

What does it mean to put AI to work, rather than simply give it a task? In a white paper drawing on engagement with more than 150 executives and experts, the World Economic Forum and Kearney suggest redesigning whole workflows: decide what information feeds the work, where people make decisions and how outcomes are checked. Anthropic says it has committed $100 million to an academy aiming to train 10,000 engineers by the end of 2027 to bring its assistant into businesses. That is a training goal, not yet a business result. See: Redesign Work Around AI, Not Just Individual Tasks · Anthropic invests $100 million to train 10,000 engineers…

The distinction between doing a task and appearing to do it shows up in Santillana’s study of small security models. A keyword-based test gave two models nearly identical tool-use scores, although the grader could credit a model for merely naming a tool in prose. Stricter checks asked whether it made a structured call with appropriate arguments. They also found unwanted calls on prompts that did not need one. If a team connects a model to tools, it might need to inspect both the calls it makes and the calls it should leave alone. See: Check Whether a Small Model Calls Tools, Not Just Names Them

The same question applies to things we can see. The authors of MaLiang-Harness describe drawing code that runs successfully but produces a picture that misses the instructions; their system inspects the rendered result after edits. In a separate study of generated first-person task videos, the Ego2Act authors report recurring missing steps and mistakes in object interactions. Their automated judge helped compare clips, but could miss brief errors. A finished file, then, is not the same thing as a finished job. See: Check whether generated visuals match the instructions, not… · Check whether generated task videos show every necessary…

Even keeping track of earlier work needs a check. xAI says Grok Build now saves project notes between sessions, though its announcement does not establish that those notes are always accurate. In a different setting, researchers testing repeated training found that combined safeguards helped a model retain more answers to questions it had seen before. They also found continued forgetting and losses on separate tests of broader ability. Saved notes and trained-in answers are different kinds of memory, but both raise a familiar question: what should a person verify before carrying on? See: Memory in Grok Build · Combining memory safeguards may help language models retain…

None of these items establishes one recipe for every workplace. Together, they suggest a useful place to start: look beyond whether AI produced something, and ask whether the result meets the need, who checks it and what happens when it does not. If you use AI at work this week, could you pick one small task and note what you would inspect before passing its output to someone else? See: Redesign Work Around AI, Not Just Individual Tasks · Check Whether a Small Model Calls Tools, Not Just Names Them · Check whether generated visuals match the instructions, not…

Just here for the stories? They are below, by topic.

PtoP · NEWSLETTER

Get the next issue in your inbox

Six papers and the news, explained in plain English, each morning after the editor approves the issue. Free, and you can unsubscribe at any time.

From the news desk

News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.

  1. Tools and building Google · 02-10-2026The latest AI news we announced in September 2026 ↗Read the explainer ↓
  2. Business and strategy TechCrunch · 02-10-2026Call it AI, call it Super Intelligence, only 2% of consumers ... ↗Read the explainer ↓
  3. Learning and education Anthropic · 02-10-2026Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap ↗Read the explainer ↓
  4. AI agents Google · 30-09-2026Gemini 4 Argon: our next era of frontier intelligence ↗Read the explainer ↓
  5. People and society BBC · 24-09-2026Are we back in big tech's 'move fast and break things' era? ↗Read the explainer ↓
  6. Tools and building xAI · 16-09-2026Memory in Grok Build ↗Read the explainer ↓

PODCAST / ENGLISH EDITION

The full issue, then each paper

The front-page episode covers the papers and news; each paper also has its own short conversation.

Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.

EP 01ARXIV:pdf-1a8d6e62eaff

Redesign Work Around AI, Not Just Individual Tasks

The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to…

EP 02ARXIV:2610.01092

Check whether generated task videos show every necessary step

The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being…

EP 04ARXIV:2609.34309

Check whether generated visuals match the instructions, not just the code

A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might…

EP 05ARXIV:2609.06986

Combining memory safeguards may help language models retain updates

Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to…

EP 06ARXIV:2609.36484

Model builders can use a teacher’s learning path as a training target

Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated…

FROM THE NEWS DESK

The news, explained

Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.

Tools and building Google · 02-10-2026 · company announcement

The latest AI news we announced in September 2026

In brief

Google says it announced several AI updates in September 2026, led by Gemini 4 Argon, a new model designed for complex work. The central idea is that Google is expanding what its AI tools can help people do, from software security to everyday tasks.

An example

For example, a security team might ask Argon to help find a weakness in its software and suggest a fix. This illustrates the intended use, not a reported result.

Application

Trusted security teams could use Argon to help review software weaknesses and possible fixes.

The limitation

Google says Argon is currently rolling out only to trusted cyber defenders. It is gathering early feedback on guardrails before expanding access. The article does not establish how well it will work in everyday use.

Takeaway

This is a company roundup of announcements, not an independent test of the tools.

Original source ↗

Was this explanation easy to understand?

Business and strategy TechCrunch · 02-10-2026 · news report

Call it AI, call it Super Intelligence, only 2% of consumers ...

In brief

TechCrunch’s podcast says consumer AI is a hard business, even as companies pursue bigger customers. It also reports that tech leaders signed an AI safety pledge and that President Trump issued an executive order calling AI “super intelligence.” The central idea is that a new name or friendlier product may not persuade people to pay.

An example

A person might use a free AI chatbot while their employer pays for AI tools. That illustrates the difference between consumer and enterprise spending; it does not establish the headline’s 2% figure.

Application

A company building an AI service could test whether individuals will pay before relying on household subscriptions for its business.

The limitation

The supplied podcast description gives no definition, source or method for the headline’s 2% figure. It also does not show how much consumers or businesses spend.

Takeaway

Treat the 2% as a headline claim, not a verified measure; the useful question is who will pay for AI.

Original source ↗

Was this explanation easy to understand?

Learning and education Anthropic · 02-10-2026 · company announcement

Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap

In brief

Anthropic says it has launched an academy to train engineers to put its Claude assistant to work inside businesses. It has committed $100 million and aims to train 10,000 engineers by the end of 2027. The central idea is to teach people to take a useful idea through testing and into everyday use.

An example

For example, a bank engineer might build a Claude tool to sort customer requests, then have it checked for safety before staff use it. This is an illustration, not a project reported by Anthropic.

Application

A company could nominate an engineer for the residency to work on a named Claude project, with training and support from Anthropic engineers.

The limitation

The 10,000-engineer figure is a goal, not a completed result. This company announcement does not establish that the training improves business outcomes.

Takeaway

Anthropic is investing in the people needed to make its tools useful at work, but the results remain to be seen.

Original source ↗

Was this explanation easy to understand?

AI agents Google · 30-09-2026 · company announcement

Gemini 4 Argon: our next era of frontier intelligence

In brief

Google announced Gemini 4 Argon, an AI model it says can carry out long, complex tasks in coding, office work and cybersecurity. It is initially rolling out to selected cyber defenders, not the general public.

An example

Imagine a software team asking an AI model to find a security flaw, suggest a repair and show its work for a person to review. That illustrates the kind of task Google says Argon can handle.

Application

A trusted security team could use Argon to look for a software vulnerability and draft a patch for human review.

The limitation

Google’s performance claims come from its own announcement and cited benchmarks. Access remains limited while the company gathers feedback and strengthens its guardrails.

Takeaway

Argon is a step toward AI that can do extended technical work, but its wider release and real-world reliability are still uncertain.

Original source ↗

Was this explanation easy to understand?

People and society BBC · 24-09-2026 · opinion or analysis

Are we back in big tech's 'move fast and break things' era?

In brief

The BBC argues that big technology companies are racing to build smarter computer assistants faster than they can control their risks. It reports that OpenAI disclosed its products had accessed public and non-public data in part of Australia’s online Medicare system. Australia is reviewing whether any laws were broken.

An example

Imagine a computer assistant sent to find a public health page also opening information it was not meant to see. This illustrates the concern; it is not a description of what the OpenAI products did.

Application

One practical response would be to check what AI agents can access and require prompt reporting when they cross a boundary.

The limitation

The BBC says the legal review is ongoing and OpenAI denies intent. Its argument that competition discourages companies from slowing down is the editor’s interpretation, not a measured finding.

Takeaway

The article’s central question is who will hold companies accountable when increasingly independent computer assistants go beyond their intended tasks.

Original source ↗

Was this explanation easy to understand?

Tools and building xAI · 16-09-2026 · company announcement

Memory in Grok Build

In brief

xAI says Grok Build can now remember useful project details between work sessions. It writes notes in the background and reads relevant ones when work on a project resumes.

An example

For example, if a developer corrects Grok Build about which command runs a project's tests, a later session could use the saved note instead of trying the wrong command again.

Application

This could help a coding team keep its usual ways of testing and reviewing code consistent across sessions.

The limitation

xAI says memory applies to new sessions and begins taking notes after the first completed turn. The announcement does not provide independent evidence that the notes are always accurate or improve results.

Takeaway

The feature aims to spare developers from repeating project instructions, but saved notes may still need checking.

Original source ↗

Was this explanation easy to understand?

01NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS

Redesign Work Around AI, Not Just Individual Tasks

Why it matters to youIf you manage a team using AI for isolated tasks, this white paper could help you think through which whole workflow to redesign next. It offers questions about where people should decide, where AI might assist and how to check whether the change helps.

The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an operating model—a way to organize decisions, people and processes—in which work produces information that can help improve later work. Consider a hypothetical supplier-approval process. An AI system might gather documents and flag missing information, while a person decides whether to approve the supplier. The team would record corrections and outcomes for later review. That example illustrates the paper’s distinction between changing a whole workflow and merely making one document task faster; it is not a reported deployment. The World Economic Forum and Kearney developed their blueprint through engagement with more than 50 enterprises and thinkers, drawing on interviews, workshops and consultations with more than 150 executives and experts. They describe five connected building blocks: an ‘intelligence engine’ that uses information from work to inform later decisions; an adaptable collection of data and AI tools; redesigned operations; human–AI teams; and choices about what AI-enabled value to offer customers. The blocks are presented as a lens for design, not a sequence every company has been shown to complete.

World Economic Forum, in collaboration with Kearney · The AI-First Operating System: A Blueprint for Operating and Business Model Innovation · arXiv:pdf-1a8d6e62eaffRead the paper ↗

In one sentence

The paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an operating model—a way to organize decisions, people and processes—in which work produces information that can help improve later work. Consider a hypothetical supplier-approval process. An AI system might gather documents and flag missing information, while a person decides whether to approve the supplier. The team would record corrections and outcomes for later review. That example illustrates the paper’s distinction between changing a whole workflow and merely making one document task faster; it is not a reported deployment. The World Economic Forum and Kearney developed their blueprint through engagement with more than 50 enterprises and thinkers, drawing on interviews, workshops and consultations with more than 150 executives and experts. They describe five connected building blocks: an ‘intelligence engine’ that uses information from work to inform later decisions; an adaptable collection of data and AI tools; redesigned operations; human–AI teams; and choices about what AI-enabled value to offer customers. The blocks are presented as a lens for design, not a sequence every company has been shown to complete.

Key concepts

  • AI-enabled versus AI-first: An AI-enabled company uses AI within existing tasks. In the paper’s definition, an AI-first company redesigns workflows, roles and decision rights so AI is integral to delivering its work.
  • Intelligence engine: The authors’ name for a system that connects business goals, operational information and outcomes. A feedback loop means checking what happened and using that information to guide a later cycle; the paper organizes these loops around speed, scale and scope.
  • Modular technology stack: Data, AI models and the connections between tools are arranged as replaceable parts. Orchestration is the set of rules that directs a task to a tool and governs how it proceeds; the authors argue that companies need control over these rules and data access as models change.
  • Human–AI teaming: The paper asks teams to decide task by task whether AI acts, assists or stays out. People retain responsibility for judgement, accountability and decisions with the highest stakes.
  • Business model canvas: The authors extend this planning framework—which asks how a business creates, delivers and earns value—with questions about AI, including how an AI-enabled offering is presented to customers.

A concrete example

Hypothetical illustration, not a finding: A procurement manager chooses supplier approval as a workflow to examine. The team maps its steps, lets AI draft a financial review for a person to check, keeps final approval with an authorized employee and records errors for the next review. The question is whether the complete process works better, not how many AI drafts it produces.

What the researchers measured

The white paper reports company examples rather than a single measured test of its five-block blueprint. In its Gamma case study, the authors say the AI-native presentation company’s inference-related gross margin moved from approximately 31% to approximately 77% six months after launch. Gross margin here describes the share of revenue left after the relevant costs; these figures concern Gamma’s inference-related measure, not every AI-first company. In a separate Claryo warehouse-platform example, the authors report an annualized profit-and-loss gain of $2.56 per square foot. That is presented as Claryo’s reported figure, not as a measured result for the blueprint or for other warehouse systems. The paper also describes practices and targets at other named organizations; those should not be read as outcomes achieved by Gamma or Claryo.

Why it matters

The distinction matters because making a task faster does not, by itself, change who makes a decision, how a handoff works or whether the organization learns from an outcome. The authors propose examining those connections together. For a working team, the immediate use is to ask where a process starts and ends, what information it needs, who can intervene and what evidence would show improvement.

Where it might help

As possible uses of the blueprint, a leader could choose a high-volume workflow, assign human and AI responsibilities at each step, set checks before any change goes live and decide how to compare the redesigned workflow with the existing one. These are planning uses, not deployments tested by the white paper.

Impact across sectors

  • Possible application in financial services: A team could examine where document checks might be assisted by AI while people retain responsibility for consequential decisions. This is not a demonstrated outcome of the blueprint.
  • Possible application in warehousing: Managers could examine whether live operational information helps them spot delays and decide when to intervene. Results from a named company in the paper should not be assumed to apply to other warehouses.
  • Possible application in customer support: A service team could map handoffs and decide which requests need a person, which might be assisted and what feedback to retain. The paper does not establish the outcome of that hypothetical redesign.

Where the evidence stops

The authors say AI-first enterprises are still at an early stage. They state that it is unclear which approaches will work best across industries, how broadly they will scale and what the long-term effects will be on productivity, employment and competition. They also say it is not yet clear which operating models will prove most resilient over time. The company accounts in this white paper do not amount to one common evaluation of the blueprint. PtoP note: this is a collaborative white paper, not a reported controlled test or a verified implementation guide.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation and software access are unverified, and no purchase is needed to do it on paper.

  1. Prerequisite: bring a workflow you know and permission to discuss it without sharing private customer or employee information.
  2. Sketch its steps, handoffs and final decision. Mark where information is missing or work repeats.
  3. For each step, write ‘person decides’, ‘AI might assist’ or ‘AI stays out’. Treat those choices as proposals, not verified settings.
  4. Note what outcome you would check, who would review errors and what must happen before a proposed change could be used.
  5. Observe whether the map reveals a process problem beyond any single task. The exercise cannot establish that an AI system will improve the workflow.

n8n example

An n8n automation—a connected sequence of software steps—is not necessary for this planning exercise; the paper provides no verified n8n integration or procedure to reproduce.

Was this explanation easy to understand?

02NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS

Check whether generated task videos show every necessary step

Why it matters to youIf you assess generated videos for training or planning physical tasks, this study offers a way to ask whether a clip reaches its goal, not just whether it looks convincing. It also shows why that check still needs care.

The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being told every step. The authors built Ego2Act to test this in first-person videos, where hands can hide objects and the view moves with the person. In the paper’s travel-case example, the toothpaste needs a cap and the toothbrush needs folding before both can be packed and the case closed; more than one order of preparation is possible. The authors also built Ego2ActJudge to check generated videos against the goal without requiring a single ‘correct’ reference video.

Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, and Alham Fikri Aji · Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation · arXiv:2610.01092Read the paper ↗

In one sentence

The study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being told every step. The authors built Ego2Act to test this in first-person videos, where hands can hide objects and the view moves with the person. In the paper’s travel-case example, the toothpaste needs a cap and the toothbrush needs folding before both can be packed and the case closed; more than one order of preparation is possible. The authors also built Ego2ActJudge to check generated videos against the goal without requiring a single ‘correct’ reference video.

Key concepts

  • A goal-directed prompt states the desired outcome, leaving the generator to work out intermediate actions. The travel-case goal, for example, does not prescribe which item to prepare first.
  • A subgoal is one observable part of the task, such as capping the toothpaste. Checking subgoals lets the evaluator notice a missing preparation step even if the case appears closed.
  • Task completion asks whether each action starts, takes place and leaves the required result. The checks stop at the first failure.
  • Physical plausibility separately asks whether objects remain consistent and whether their movements, contact and resulting positions make physical sense. An omitted action is not itself scored as a physical interaction.
  • Reference-free evaluation allows different valid routes to the same goal. Ego2ActJudge derives subgoals from the starting scene and goal, then examines the video rather than comparing it with one prescribed recording.

A concrete example

Hypothetical illustration, not a study result: A generated clip shows a hand put an unfolded toothbrush into a travel case and close it. A reviewer would check the folding subgoal and the final state separately, rather than crediting the clip simply because the case closes.

What the researchers measured

The authors report a benchmark of 110 real-world cases and 2,640 videos, including human recordings and outputs from six generators. On a human-rated panel covering 25 cases, Seedance-2.0 had the highest overall generator score in the main results table: 64.0 on a 0–100 score combining task completion and physical plausibility. That is a rubric score, not a percentage of tasks completed. On the 600-video human-rating panel, Ego2ActJudge’s final scores had a correlation of 0.69 with the combined human ratings; correlation describes how scores vary together, not exact agreement. The authors report recurring omitted or incomplete steps and errors in detailed object interactions.

Why it matters

A video may show an appealing end scene while skipping an action needed to get there. By separating goal completion from physical plausibility, the authors make those different kinds of failure visible. Their automated judge is intended to support comparisons across many recordings, while the paper cautions against treating one clip’s score as a definitive verdict.

Where it might help

As a possible research use, the benchmark could help teams compare video generators intended to depict multi-step object handling. Its scores are evidence about generated recordings in this benchmark, not evidence that a model can carry out a task in the physical world.

Impact across sectors

  • Possible use in robotics research: inspect proposed training videos for omitted actions and implausible contact before considering them as demonstrations; this was not tested as a robot-training deployment.
  • Possible use in instructional-video development: make a checklist of intermediate outcomes that a generated demonstration ought to show; the paper did not test learners or teaching outcomes.
  • Possible use in video-model development: compare how generators depict everyday organization or personal-care tasks; these are benchmark settings, not proven workplace applications.

Where the evidence stops

Human ratings covered 25 of the 110 cases, rather than the whole benchmark; automated evaluators scored the benchmark videos they could score. The authors say Ego2ActJudge tends to score executions above human raters, misses many brief errors that unfold between sampled frames, and varies across runs on an individual video. Some action-category groups are small or overlap, and the paper reports that its tested feature associations did not survive correction for multiple comparisons. Generated clips were limited by the models’ available durations. The authors also say baseline evaluators needed adaptations for this setting. This source is an arXiv preprint, not presented here as peer-reviewed work.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation is unverified. Prerequisites: a short task video you are permitted to review, its stated goal, and a way to replay it. No software purchase or model access is needed for this exercise; access terms for the paper’s models and evaluator are not established here.

  1. Write the goal as a visible final state.
  2. List the necessary observable subgoals, allowing different valid orders.
  3. Replay the clip and note for each subgoal whether the action starts, happens and reaches its intended state.
  4. Separately watch attempted interactions for disappearing objects, unexplained movement or unstable results.
  5. Record what you could not see instead of guessing. You should observe whether a convincing-looking ending conceals an omitted step; this exercise does not reproduce Ego2ActJudge. Although the paper says code will be released upon publication, the supplied text does not verify an installable repository or commands.

n8n example

An n8n workflow is not appropriate to specify here: the supplied text does not verify an available evaluator interface or a tested automation integration.

Was this explanation easy to understand?

03NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS

Check Whether a Small Model Calls Tools, Not Just Names Them

Why it matters to youIf you assess a small language model for a security workflow, its tool-use score may not tell you whether it can actually make a call. This preprint offers checks you could use before relying on that score.

A model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two Spanish-language security models that received nearly identical scores from a keyword-based test. He then checked their actual output, starting with examples from their training data and moving to prompts containing unfamiliar details. He also examined the models’ likelihood of starting a tool call and, after fine-tuning the larger model, whether the part of the model representing the call-start token had changed. The comparison suggests that training-data composition matters for how readily this format appears, but the models differ in more than their data mix, so it does not isolate a cause.

Juan S. Santillana · Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models · arXiv:2610.02142Read the paper ↗

In one sentence

A model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two Spanish-language security models that received nearly identical scores from a keyword-based test. He then checked their actual output, starting with examples from their training data and moving to prompts containing unfamiliar details. He also examined the models’ likelihood of starting a tool call and, after fine-tuning the larger model, whether the part of the model representing the call-start token had changed. The comparison suggests that training-data composition matters for how readily this format appears, but the models differ in more than their data mix, so it does not isolate a cause.

Key concepts

  • A tool call is a structured request naming an available tool and supplying its arguments. Mentioning the tool in ordinary prose is not the same action.
  • A keyword harness is an automatic grader that looks for words in an answer. The paper’s tool-use harness could award credit when a tool name appeared without a valid call.
  • A verbatim-reproduction check gives a model a real training prompt and asks whether it produces a valid call. The authors distinguish a sensible new argument from an exact copy of the training answer.
  • A first-token probe checks how likely the model is to begin its answer with the special token that opens a tool call. It helps locate a format failure without depending on a generated answer.
  • A generalization battery uses new prompts to test tool choice, argument details and when not to call a tool. Success on familiar training prompts alone could reflect copying.

A concrete example

Hypothetical illustration, not a paper result: A security worker asks whether a particular vulnerability appears in a catalog. An answer saying that a catalog-checking tool exists would satisfy a loose word match; a tool call would instead name that tool in the required structure and put the requested vulnerability identifier in its argument.

What the researchers measured

On the authors’ single-seed keyword harness, VectraYX-600M at training step 154K scored 0.660 on tool use, while the recorded VectraYX-1B score was 0.650. The authors identify the latter as unreliable because the grader could credit a tool name in prose. In a stricter check using real training examples, the unfinished VectraYX-600M checkpoint produced valid calls with new, sensible arguments on 6/6 examples; the tested VectraYX-1B checkpoints before repair produced the structure on 0/4–6 examples per run. These counts are examples passing the specified check, not a deployment success rate. After a targeted fine-tuning run, the reproduced repaired VectraYX-1B produced well-formed calls on 0.959 of 269 corpus rows, compared with 0.100 for its pre-repair parent and 0.926 for the VectraYX-600M checkpoint. The authors say the repaired model mostly copied answers on its own corpus. On the call-requiring portion of their new-prompt battery—166 prompts about entities absent from the training corpus—the repaired VectraYX-1B passed the combined tool, argument-shape and entity check on 0.536, versus 0.428 for VectraYX-600M. The wider battery contained 238 prompts, including prompts on which a call was not wanted. On no-call prompts mentioning a vulnerability or shell command, the VectraYX-600M answered without a call on only 0.09 of items and the repaired VectraYX-1B on only 0.17. A weight check found that the repaired model’s call-start token representation was essentially unmoved; the authors locate the change in the surrounding network, within the limits of that check.

Why it matters

For work that depends on a model invoking a tool, the grader must recognize the required request, not merely related words. The paper also shows why testing when a model should refrain matters: both the smaller model and the repaired larger model often called a tool on prompts that only mentioned a vulnerability or shell command.

Where it might help

A possible use is to add a strict output check to an internal model assessment before considering a model for a tool-connected workflow. The paper tests model outputs; it does not establish a deployed security assistant.

Impact across sectors

  • Possible security-team use, not a tested deployment: distinguish a model that talks about vulnerability lookups from one that formats a request for the declared lookup tool.
  • Possible model-development use, not a tested deployment: check familiar and unfamiliar prompts separately after fine-tuning, so copied training answers are not mistaken for generalization.
  • Possible software-procurement use, not a tested deployment: ask for valid-call and no-call checks alongside a vendor’s keyword-based tool-use score.

Where the evidence stops

This is a preprint, not presented as peer-reviewed work. The VectraYX-600M run stopped permanently at 64% of its planned schedule. The two models share an architecture family and tokenizer but differ in size, training volume and curriculum; their comparison is a natural experiment, not a controlled test of data composition. The authors’ trajectory probes also overlap with or closely resemble the larger model’s earlier training examples: an early apparent tool-call tendency was exact copying, not demonstrated generalization. The original successful repair checkpoint and its parent were lost, so the primary strict results use a reproduced run from a surviving sibling checkpoint. The new-prompt battery was built by the authors; it does not cover unseen tools or multi-call sequences. The authors report no evidence that the repaired larger model answers ordinary conversational questions usefully. Harness figures are single-seed, and the repair is one recipe. In a separate set of five fine-tuning runs using a different numerical setup, the apparent benefit of a diverse corpus for avoiding unwanted calls was not reproduced by a partial second seed; the authors retain it as a hypothesis, not an established effect.

FROM PAPER TO PRACTICE

How to try it

Conceptual exercise only: the supplied text does not verify an installation procedure or runnable commands. Prerequisites are sample prompts, the declared tool format and a way to inspect a model’s full output; access to a model or computing service may carry a cost.

  1. Write down what a valid call must contain, including the tool name and required argument.
  2. Prepare one prompt that needs a call and one that mentions a tool-related topic but does not need a call.
  3. Inspect the complete outputs: mark a tool name in prose separately from a valid structured request.
  4. Try a new identifier that was not in the first prompt, and check whether the argument carries that identifier rather than a copied one. Observe both missed calls and unwanted calls; do not treat this exercise as a tested installation or a substitute for the paper’s evaluation.

n8n example

A proposed integration, not a feature tested in the paper: an n8n automation could receive model output, check that any tool request has the required structure and arguments, and route failures for human review before any tool is run.

Was this explanation easy to understand?

04TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS

Check whether generated visuals match the instructions, not just the code

Why it matters to youIf you make illustrations or short animations from detailed briefs, a visual program could make it easier to inspect and revise how a result was built. This paper explores that approach; it does not establish a production workflow for your job.

A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might produce a video, but put the sun in front throughout. The authors call this mismatch the Program-to-Visual gap. Their MaLiang-Harness has a multimodal language model—one that works with both words and images—write drawing or animation code, then inspect the rendered result and revise it. A renderer is the tool that turns the code into an image or video. The system keeps earlier versions and links each inspection to the version that produced it, so an old approval does not stand in for a check of the latest edit.

Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu, Zuxuan Wu, Yu-Gang Jiang, Shuicheng Yan · MaLiang-Harness: A Programmable Path to Image and Video Generation · arXiv:2609.34309 · 398 HF votes at selectionRead the paper ↗

In one sentence

A drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might produce a video, but put the sun in front throughout. The authors call this mismatch the Program-to-Visual gap. Their MaLiang-Harness has a multimodal language model—one that works with both words and images—write drawing or animation code, then inspect the rendered result and revise it. A renderer is the tool that turns the code into an image or video. The system keeps earlier versions and links each inspection to the version that produced it, so an old approval does not stand in for a check of the latest edit.

Key concepts

  • Program-to-Visual gap: code that executes successfully can still miss a requested position, appearance or action.
  • Persistent Executable Generation: the harness keeps the artwork’s code, assets, requirements and plan across revisions, or saved versions.
  • Traceable Generation Process: recorded operations connect code changes to the visual results inspected after those changes; this is not a record of the model’s private reasoning.
  • Revision-aware Editing and Verification: earlier artwork can be restored, while the current revision needs its own checks before delivery.

A concrete example

Hypothetical illustration, not a study result: a designer requests a poster with a small orange sun above a navy mountain. If the first rendering puts the sun behind the mountain, the designer could inspect that version, change the drawing code and check the new rendering rather than treating error-free code as a finished poster.

What the researchers measured

In the authors’ MaLiang-IBench evaluation, GPT-6-Astra produced a successful output for every image task; 48 of 50 tasks met all the image-quality thresholds. That count means the outputs passed the paper’s separate checks for following the prompt, appearance and composition, as judged on successful images by GPT-6-Sol. In the authors’ MaLiang-VBench evaluation, GPT-6-Astra produced a successful output for every video task; 10 of 13 tasks met all the video-quality thresholds, which also include motion coherence. A successful output means a decodable file passed the harness’s completion checks, not that it met those separate quality thresholds.

Why it matters

The authors separate producing a usable file from satisfying the visual brief. They also report that similar scores on a general model-capability measure did not necessarily correspond to similar results on their image tasks. For someone choosing a model for this kind of work, their findings point to checking the rendered work against the actual brief rather than relying on code execution or a general score alone.

Where it might help

Possible uses include artwork with precise layout instructions and silent animations with ordered actions. These are potential uses of the approach, not deployments demonstrated by the paper.

Impact across sectors

  • Possible use in graphic design: a team could inspect saved poster revisions when a change fixes one layout requirement but affects another. This is hypothetical, not a measured workplace outcome.
  • Possible use in animation production: a team could compare sampled moments of a short coded animation with its action brief. This is hypothetical and would not, by itself, establish smooth motion between samples.

Where the evidence stops

The authors say the harness’s visual reviews are model self-assessments, not independent measures of perceptual quality. Their separate video-quality review used one reviewer, was not formally blinded and examined sampled frames; it cannot establish smoothness between frames or rule out brief faults. Image results combine batches, retries and historical runs rather than uniform first attempts, and DeepSeek-V4-Pro was evaluated without visual feedback. Timing and budget conventions also differ across some evaluated configurations, limiting direct cost comparisons. In their examples, finer brushwork did not make a stylized image photorealistic, and refinement could stall or exhaust its token budget, the limit on text the model could use. General PtoP note: this source is an arXiv preprint, not a claim of peer review, and its benchmark results are not a workplace deployment.

FROM PAPER TO PRACTICE

How to try it

The official repository README provides an offline installation example, not a reproduction of the paper’s quality results. Prerequisites are a computer that can install Python 3.12, project dependencies and the Chromium browser used for rendering; downloading them requires access. The README does not state a monetary cost for this offline example. Installation and behavior on your machine are unverified here.

  1. Clone the repository with `git clone https://github.com/gulucaptain/MaLiang-Harness.git`, then enter it with `cd MaLiang-Harness`.
  2. From that directory, run `python3.12 -m venv .venv`, `source .venv/bin/activate`, `python -m pip install -e '.[dev]'` and `python -m pip install --no-deps -e vendor/deepagents/libs/deepagents`.
  3. Run `python -m playwright install chromium`, `python -m pip check` and `maliang-harness doctor` to install the browser dependency and check the setup.
  4. Run `python examples/scene_demo.py --project runs/scene-demo` with a new output directory. Inspect the generated files in that directory and the rendering process. The README says this uses a scripted model without a model service connection; it is an installation example, not evidence of a real model’s generation quality.

n8n example

No n8n workflow is proposed: the paper studies local program rendering and revision-specific visual inspection, and the supplied material does not establish an n8n integration.

Was this explanation easy to understand?

05TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS

Combining memory safeguards may help language models retain updates

Why it matters to youIf you maintain a language model that receives new information over time, this study may help you think about how to keep earlier answers from being overwritten. It tests memory during repeated training, not a deployed update system.

Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to protect earlier learning helped the model retain more answers than using any one of those ways alone. Imagine teaching a model a fictional fact today, then many more facts over the following weeks: the question is whether it can still give the first answer. The authors trained Qwen3-4B-Base on streams of questions and answers, one task after another, without keeping earlier raw examples for later training or telling the model which task a question came from. They compared safeguards that regenerate practice material, encourage the updated model to behave like its previous version, and limit changes to parts judged important. They also compared reusing one small set of trainable parts with folding each task’s changes into the model before starting a fresh set.

Zheyuan Zhang, Alvin Zhang, Daniel Khashabi, Tianmin Shu · Continual Learning Mechanisms Compose for Long-Horizon Memorization · arXiv:2609.06986 · 376 HF votes at selectionRead the paper ↗

In one sentence

Can a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to protect earlier learning helped the model retain more answers than using any one of those ways alone. Imagine teaching a model a fictional fact today, then many more facts over the following weeks: the question is whether it can still give the first answer. The authors trained Qwen3-4B-Base on streams of questions and answers, one task after another, without keeping earlier raw examples for later training or telling the model which task a question came from. They compared safeguards that regenerate practice material, encourage the updated model to behave like its previous version, and limit changes to parts judged important. They also compared reusing one small set of trainable parts with folding each task’s changes into the model before starting a fresh set.

Key concepts

  • Continual learning means updating one model on successive tasks while trying to retain what it learned earlier.
  • Generative replay supplies practice material made by the previous model instead of saving raw examples from earlier tasks.
  • Self-distillation asks the updated model to stay close to the previous model’s predictions on the current task’s inputs.
  • A weight anchor limits changes to trainable values estimated to matter for earlier learning.
  • Merged LoRA is a rule for small trainable model updates: fold one task’s update into the model, then start a fresh update for the next task.

A concrete example

Hypothetical illustration, not a study result: a model learns the answer to a question about an invented town, then learns many unrelated question-and-answer sets. A memory test asks the original question again. The safeguards are intended to reduce the chance that later training overwrites its answer; this example does not show how it would respond to a reworded question.

What the researchers measured

After 100 tasks, the authors report 1.2% average final retention for naive sequential fine-tuning and 34.9% for their method combining all three anchors with merged LoRA. These figures average across Symbol-QA, LLM-QA and Real-QA using Qwen3-4B-Base; final retention is the share of training questions answered correctly across all tasks after the last update, averaged over tasks. The method ranked among the top 3 compositions on each dataset, though a different composition had the highest mean retention on Symbol-QA and Real-QA. In their factorial comparison, the authors identify replay and merged LoRA as the largest average contributors and report a positive interaction between them on all three datasets. The final experiments used three training seeds per method and dataset.

Why it matters

The authors separate learning a new task from remembering it after later updates. Their results suggest that, in this memorization setting, the combination of replay and the rule for retaining updates matters more than choosing a single safeguard in isolation.

Where it might help

A possible use is designing experiments for models that receive successive batches of information. The paper does not demonstrate a working deployment or show that the model can reliably answer newly phrased questions.

Impact across sectors

  • Possible, not a tested deployment: a publisher updating a model with successive batches of reference facts could use the study’s retention measures to plan an experiment.
  • Possible, not a tested deployment: a support team training on successive sets of standard questions could compare whether old, identically phrased questions remain answerable.
  • Possible, not a tested deployment: an education team updating a question-answer model over time could test recall of earlier training questions, while separately checking unfamiliar wording.

Where the evidence stops

The authors tested recall on the same questions used for training, so the results do not establish that answers survive rewording or transfer to new questions. They report that even the stronger compositions continue to forget older material. Stronger memorization also did not preserve general ability: the evaluated methods lost substantial accuracy on separate reasoning and knowledge tests. Their search held task order fixed across training seeds, so those seeds did not measure sensitivity to task order; the later evaluation used a different order, and the search winner was not the final leader on Symbol-QA. General PtoP note: this source is an arXiv preprint, not evidence of peer review or a deployed system.

FROM PAPER TO PRACTICE

How to try it

The official repository documents a command-checking dry run; installation and execution have not been verified here. Prerequisites are Git, Conda and Python 3.11. Installing dependencies may require a PyTorch package suited to your hardware. The dry run does not load a model; full paper-scale runs require a CUDA GPU, download Qwen3-4B-Base from Hugging Face on first use, and can be expensive.

  1. Clone the repository: `git clone https://github.com/cozheyuanzhangde/compose-cl.git` and `cd compose-cl`.
  2. Create and activate its environment: `conda create --name compose-cl python=3.11 -y` and `conda activate compose-cl`.
  3. Install its listed packages: `python -m pip install --upgrade pip` and `python -m pip install -r requirements.txt`.
  4. Inspect one resolved experiment command without loading the model: `python -m experiments.run_final --dataset symbol_qa --method si_sd_replay_merge --seed 41 --dry-run`. You should observe the command the launcher would use, not a retention result.

n8n example

n8n is not appropriate for reproducing this paper’s sequential model-training experiments; the authors do not report an n8n integration.

Was this explanation easy to understand?

06TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS

Model builders can use a teacher’s learning path as a training target

Why it matters to youIf you train language models for mathematical reasoning, this paper offers a possible way to learn from both an improved model and the earlier version it came from. The authors tested the method on competition-math problems, not in a working product.

Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated the teacher from its earlier version. Imagine a student working through a math problem. On the same partial answer, the method compares signals inside the teacher with signals inside the model before reinforcement learning—a training process that rewards desired responses. It then sets a target farther along that difference and trains the student’s internal signals toward it. The authors call the method RIDE. This is a training target, not a guarantee that going farther will produce a better answer.

Hao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang, and Naibo Wang · The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation · arXiv:2609.36484 · 372 HF votes at selectionRead the paper ↗

In one sentence

Instead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated the teacher from its earlier version. Imagine a student working through a math problem. On the same partial answer, the method compares signals inside the teacher with signals inside the model before reinforcement learning—a training process that rewards desired responses. It then sets a target farther along that difference and trains the student’s internal signals toward it. The authors call the method RIDE. This is a training target, not a guarantee that going farther will produce a better answer.

Key concepts

  • A student is the model being trained; its teacher is a model already improved through reinforcement learning. In this study, the student begins from the teacher’s earlier checkpoint, or saved version.
  • On-policy training means the student generates its own partial answers. The teacher and earlier checkpoint are then compared on those same partial answers, rather than on different responses.
  • A hidden state is an internal set of values a model builds while processing text. RIDE compares hidden states at each layer, or stage of processing, before they are turned into probabilities for the next piece of text.
  • The residual is the difference between the teacher’s hidden state and the earlier checkpoint’s hidden state on the same text. RIDE uses that difference as a direction for its target, rather than treating the teacher’s state as the endpoint.
  • The extrapolation coefficient controls where the target sits. At a value of 1, the target is the teacher’s state; above 1, it lies farther along the measured direction.

A concrete example

Hypothetical illustration, not a paper result: a student begins a competition-math solution by considering one approach. The earlier model and the improved teacher process that same unfinished solution. RIDE measures the difference between their internal signals and sets the student a target beyond the teacher’s signals in that direction. Whether the student’s finished solution is correct must still be checked.

What the researchers measured

The authors evaluated four base-model and reinforcement-learning-teacher pairs on three competition-math problem sets. Their score, Avg@16, averages whether sampled answers are correct across 16 responses per problem, then averages the three problem-set scores; it is reported as a percentage. For the DeepSeek-R1-Distill-Qwen-1.5B and JustRL-DeepSeek-1.5B pair, RIDE’s mean score across three training runs was 56.38, while the fixed teacher’s score was 55.30. The authors report that RIDE’s mean was above its teacher’s score on every pair and above the methods they compared it with. On three pairs, however, the margin over the teacher was within one across-run standard deviation. These are measured math-evaluation findings, not deployment results.

Why it matters

The comparison distinguishes two ways of learning from an improved model: copying what it produces and using how its internal processing changed. The authors report that their internal-state approach performed differently from methods that pushed beyond the teacher using output probabilities.

Where it might help

A possible research use is training a student when a team has both a model improved through reinforcement learning and the checkpoint it started from. The paper does not establish a ready-to-use service or performance outside its competition-math evaluations.

Impact across sectors

  • Possible, not tested as a deployment: language-model training teams could investigate whether saved before-and-after checkpoints provide a useful training direction when building math-focused models.
  • Possible, not tested as a deployment: competition-math evaluation teams could use the paper’s comparison as a starting point for studying how different training targets affect answers.
  • Possible, not tested as a deployment: AI safety teams could examine whether training beyond a teacher’s internal signals also carries forward its errors or biases.

Where the evidence stops

The authors say RIDE requires the checkpoint from before reinforcement learning and models whose internal states can be compared. It uses one global extrapolation setting and was evaluated only on mathematical reasoning with one reinforcement-learning recipe. They note that the two AIME evaluations were near the scoring floor for the Llama-3.2-3B pair, so evidence for that pair rests mainly on AIMO. They also say a student could carry forward or amplify the teacher’s errors or biases; its calibration and safety need independent evaluation before deployment. The paper does not report an analysis of overlap between training prompts and evaluation problems. PtoP note: the supplied source is an arXiv preprint, not evidence of peer review.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation and a runnable RIDE procedure are unverified. The official repository says it currently contains a project overview, but not training code, evaluation scripts, or model checkpoints.

  1. Prerequisite: read the paper’s method description and the repository overview; no software installation is needed for this exercise.
  2. Sketch an earlier model, its improved teacher, and a student starting from the earlier model.
  3. For one hypothetical unfinished math answer, mark that all three models must process the same text before their internal states can be compared.
  4. Draw the teacher’s change from the earlier model as a direction, then place a hypothetical training target beyond the teacher. Observe that the sketch specifies a target, not a verified improvement in answers. Access and cost barrier: running the research training would require checkpoints, substantial computing resources, and code or procedures not supplied by the repository. Do not treat this exercise as a reproduction.

n8n example

n8n is not appropriate here: the reported method trains on internal model states, while the official repository provides no runnable training or evaluation integration.

Was this explanation easy to understand?

COMPARISON

What these studies show together

Compare topics, possible uses and limits. Each result remains attributed to its own paper.

PaperCentral ideaPossible useLimit
Redesign Work Around AI, Not Just Individual TasksThe paper asks what changes when a business designs its work around artificial intelligence, or AI, rather than adding AI tools to existing tasks. Its answer is an…As possible uses of the blueprint, a leader could choose a high-volume workflow, assign human and AI responsibilities at each step, set…The authors say AI-first enterprises are still at an early stage. They state that it is unclear which approaches will work best across…
Check whether generated task videos show every necessary stepThe study asks whether a video generator can show a hand completing an everyday task from a starting image and a goal, without being told every step. The authors built…As a possible research use, the benchmark could help teams compare video generators intended to depict multi-step object handling. Its…Human ratings covered 25 of the 110 cases, rather than the whole benchmark; automated evaluators scored the benchmark videos they could…
Check Whether a Small Model Calls Tools, Not Just Names ThemA model can get credit for naming a tool without issuing the structured request needed to use it. Santillana compared two Spanish-language security models that received…A possible use is to add a strict output check to an internal model assessment before considering a model for a tool-connected workflow…This is a preprint, not presented as peer-reviewed work. The VectraYX-600M run stopped permanently at 64% of its planned schedule. The two…
Check whether generated visuals match the instructions, not just the codeA drawing program can run correctly and still make the wrong picture. Imagine asking for a sun to move behind a hill: the code might produce a video, but put the sun in…Possible uses include artwork with precise layout instructions and silent animations with ordered actions. These are potential uses of the…The authors say the harness’s visual reviews are model self-assessments, not independent measures of perceptual quality. Their separate…
Combining memory safeguards may help language models retain updatesCan a language model learn new question-and-answer pairs without quickly losing old ones? The authors found that combining ways to protect earlier learning helped the…A possible use is designing experiments for models that receive successive batches of information. The paper does not demonstrate a working…The authors tested recall on the same questions used for training, so the results do not establish that answers survive rewording or…
Model builders can use a teacher’s learning path as a training targetInstead of asking a student model merely to copy an improved teacher, the authors train it toward the internal change that separated the teacher from its earlier…A possible research use is training a student when a team has both a model improved through reinforcement learning and the checkpoint it…The authors say RIDE requires the checkpoint from before reinforcement learning and models whose internal states can be compared. It uses…

ARCHIVE

Previous issues

The last two issues. Every earlier edition is in the archive.

02
AI agent training, skills and teamwork; robot goals and 3D shape generation02-10-2026
↗
01
Model Memory, Scene Prediction, Scientific Software and Agent Workflows01-10-2026
↗