Long-document AI design depends on what models keep and retrieve
A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The…
ISSUE 05/2026 · 01-10-2026
A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an incorrect astronomy measurement, later entered the same value again, and checked the saved answer…

6 paper · Original sources linked in every story
5 MINUTES ON PtoP
An AI agent can look busy long before you know whether its work is right. WIRED reports that companies are adding agents to workplace chats and, in some cases, org charts. It also cites research in which managers caught fewer errors when told work came from an “AI employee” rather than an “AI tool.” That finding does not tell us how every team will behave. But it raises a useful question: when software feels like a colleague, will we still check what it hands over? See: AI Agents Are About to Flood the Workforce. No One's ...
The authors of OSWorld-Science put that question in concrete terms. They built scientific-software tasks whose checks inspect the result an assistant leaves behind, including files, measurements and application state. In one run they describe, an assistant entered a wrong astronomy measurement twice, then checked the saved answer against itself rather than against what the measurement meant. Anthropic offers a different example of why the next check matters: it says Claude found a previously uncharacterized enzyme system in DNA records, and human scientists confirmed part of the finding in a lab. Anthropic says they still do not know what the system does. See: Test whether an assistant delivers the right scientific… · Claude discovers a novel enzyme system with CRISPR-like…
The problem can begin even before an assistant reaches us. In experiments with search agents that generate their own practice questions, the authors of the CrossFit paper found that an agent could appear to improve by learning to repeat a wrong draft answer. Their method uses a separate answering model, trained on a different group of documents, to help choose practice questions. In their tested setting, that reduced agreement on the same incorrect answer. The authors do not present it as a guarantee of truth: different models can still share mistakes. See: Keep Search Agents From Rewarding Their Own Wrong Answers
So what might help when work stretches across many steps? The StateM authors tested a shared runbook that records an agent’s current phase and checks conditions before a handoff; they report benchmark gains, while noting that the runbook and runtime were tested together and that results vary by setting. The Raven authors tested plans for assigning work to specialist agents, but their planning test did not run those specialists. Read together, the papers suggest three distinct places to look: the plan, the check at each handoff, and the finished result. None is a substitute for the others. See: A shared runbook may help AI agents finish long tasks · Plan specialist AI work before handing tasks to agents
If agents become a more familiar part of work, a visible chain of checks could matter as much as a smooth conversation. This week, if you give an assistant a task with several steps, you might try writing down one thing you can inspect at the end—an output file, a figure traced to its source, or a saved change—and ask: what would show me that this result is right, rather than merely finished? See: AI Agents Are About to Flood the Workforce. No One's ... · Test whether an assistant delivers the right scientific…
Just here for the stories? They are below, by topic.
News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.
THIS ISSUE AT A GLANCE
↗02 / NEW / EDITOR PICKSTest whether an assistant delivers the right scientific resultA computer-using assistant can carry out many steps in scientific software yet still submit a wrong…
↗03 / NEW / EDITOR PICKSKeep Search Agents From Rewarding Their Own Wrong AnswersA search agent can appear to improve by learning to repeat its own wrong answers. The authors study a…
↗04 / TOP VOTED · 6 MONTHSPlan specialist AI work before handing tasks to agentsWhen a job needs several kinds of work, the authors ask whether one system can choose suitable…
↗05 / TOP VOTED · 6 MONTHSLearning scene changes could help AI predict images and robot actionsThe study asks whether one learned picture of a changing scene can support several kinds of prediction…
↗06 / TOP VOTED · 6 MONTHSA shared runbook may help AI agents finish long tasksThe study asks whether an AI agent can finish more long tasks when its working procedure improves, even…
↗PODCAST / ENGLISH EDITION
The front-page episode covers the papers and news; each paper also has its own short conversation.
Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.
A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The…
A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by…
A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part…
When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs…
The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to…
The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does…
FROM THE NEWS DESK
Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.
Image, audio and video Hugging Face Blog · 30-09-2026 · company announcement
Hugging Face says it built a public ranking to compare AI-generated speech across languages, including voices based on a sample recording. It uses repeatable tests of word accuracy, speed and voice similarity so newly released models can be compared more quickly.
Imagine choosing a voice for a travel guide. You could listen to two systems read a Japanese station announcement, then check which says the right words and starts speaking sooner.
A developer could use the TTS leaderboard to shortlist voices for a multilingual assistant, comparing intelligibility and streaming speed before listening to the results.
Hugging Face says its automated scores do not measure how natural or expressive a voice sounds, or which voice people prefer. The tested languages and hardware also limit what the rankings can tell you.
The ranking can narrow the choices, but listening still matters—especially when assessing voice cloning.
Was this explanation easy to understand?
People and society WIRED · 28-09-2026 · news report
WIRED reports that companies are adding AI agents to workplace chats and, in some cases, to an org chart. An AI agent is software that can carry out tasks without someone guiding every step. The article’s central concern is that making such software feel like a coworker may change how carefully people check its work.
For example, WIRED describes an agent that posted details of its manager’s calendar in a team chat. She corrected it directly, but the mistake shows why an apparent coworker still needs oversight.
A company using an agent to schedule meetings could set guardrails: limit which calendar details it can see and require a person to check messages before they go to a group.
WIRED cites research in which managers caught fewer errors when told work came from an AI employee rather than an AI tool. That finding does not establish how every team will behave, or how many agents will join workplaces.
Treat AI agents as tools that can do useful work, not as coworkers whose work can go unchecked.
Was this explanation easy to understand?
Image, audio and video Google · 23-09-2026 · company announcement
Google says it is rolling out two tools that turn written scripts into expressive speech. A game maker could make a dragon whisper one line and shout the next. These text-to-speech models let creators design voices and direct how each line sounds.
As an illustration, a podcast producer could write a conversation for two characters and ask for a different speaking style for each.
A video maker could use the tools to create spoken versions of a script in different languages. Google says developers can try both models in Google AI Studio.
This is Google's product announcement, not an independent test of its quality claims. Voice replication in Google AI Studio is unavailable in several regions, including the UK and India. Google says its safeguards include consent verification and a watermark, but does not show that misuse is impossible.
Google is making generated speech easier to direct, but its quality and safeguards need independent scrutiny.
Was this explanation easy to understand?
AI agents Anthropic · 23-09-2026 · company announcement
Anthropic says Claude found a previously uncharacterized enzyme system by searching DNA records, and human scientists checked the finding in their lab. The system has repeated DNA segments that resemble those in CRISPR, but Anthropic does not yet know what it does.
Imagine a software helper searching millions of recipes and noticing that the same unusual note keeps appearing beside one ingredient. A cook would still need to test what the combination does. Here, Claude spotted a pattern, and scientists followed up with experiments.
This approach could help researchers choose which unexplored DNA patterns and enzymes to investigate in the lab.
These are early results described by Anthropic in a preprint. Its experiments found that the repeated DNA segments produce short pieces of RNA, but they have not established the system’s function or shown that it can be used as a gene-editing tool.
Claude helped identify a promising biological puzzle; human experiments are still needed to solve it.
Was this explanation easy to understand?
Tools and building xAI · 21-09-2026 · company announcement
xAI announced Grok 4.7, a new AI model it says is better at coding and longer work tasks than Grok 4.6. The central idea is that it can spend more time on a task and check its own work more carefully.
For example, a developer might ask it to find a bug, suggest a fix, and check whether the fix breaks anything else. That illustrates the kind of longer task xAI describes; it is not a verified result.
Developers could use Grok 4.7 to help draft and review software changes. xAI says it is available through its API and several coding tools.
The performance and safety claims come from xAI’s own benchmarks. Those tests do not establish how reliably the model will perform on every real-world task.
Grok 4.7 is a coding-focused release with claimed gains in long tasks and safeguards, but its results need to be read as company-reported tests.
Was this explanation easy to understand?
People and society BBC · 14-09-2026 · news report
The BBC reports that warnings about AI becoming hard to control have prompted calls to slow its development, though the worst-case scenarios remain hypothetical. The central concern is that AI agents—software that can take actions on its own—might eventually do things people cannot reliably stop. Critics question whether companies overstate that risk while harms are already occurring.
Imagine an email assistant that sends a message to your entire address book when you only asked it to write a draft. This is an illustration of why giving software permission to act needs safeguards, not an incident reported by the BBC.
A developer could ask third-party evaluators—independent testers—to check an AI agent’s safeguards before giving it more tasks. The BBC reports that Anthropic’s boss has proposed such checks.
The BBC describes threats to humanity as hypothetical, not measured outcomes. It also reports present-day harms, including deepfakes: convincing fake images used to depict people nude without their consent.
The debate is about how to manage uncertain future dangers without neglecting harms happening now.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS
Why it matters to youIf you review long contracts, you may want to know why an AI assistant can accept a document yet still miss a clause. This survey offers a way to understand the memory choices behind that risk, not a tested contract-review tool.
A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The authors explain that designs make different trade-offs between keeping separate pieces of text and compressing them into a smaller memory. A token is a piece of text the model processes, such as a word or part of one. A layer is one stage in the model’s sequence of processing steps. The authors compare ways of storing and retrieving tokens, including designs that use different methods in different layers. They review prior research and classify documented model releases; they do not test an assistant on contracts.
A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The authors explain that designs make different trade-offs between keeping separate pieces of text and compressing them into a smaller memory. A token is a piece of text the model processes, such as a word or part of one. A layer is one stage in the model’s sequence of processing steps. The authors compare ways of storing and retrieving tokens, including designs that use different methods in different layers. They review prior research and classify documented model releases; they do not test an assistant on contracts.
Hypothetical illustration, not a paper test: A contract mentions a cancellation deadline near the beginning. When asked about it near the end, a design that retains individual tokens could consult the earlier passage; a sparse design would first have to select it. A compressed-state design might carry information from that passage without preserving the passage as a separately selectable item. None of these descriptions establishes that a model will answer correctly.
This is a survey and descriptive architecture comparison, not a new task-performance experiment. The authors assembled 59 release-level records of publicly documented model architectures. In a separate frozen comparison of 11 high-performing open-weight model endpoints from the Artificial Analysis Intelligence Index v4.3.2, captured on September 22, 2026, they report varied attention designs and an explicit token-retrieval path in every selected architecture. These counts describe the surveyed releases and selected endpoints; they are not accuracy scores.
The authors distinguish fitting more text into a model’s input from reliably using distant evidence. Their framework makes it easier to state what a design retains and what it may overlook. That distinction can guide questions about a proposed system, but the survey does not show how well any system would perform in a reader’s job.
The survey’s framework could help people planning long-document AI systems ask what evidence remains available, how retrieval candidates are chosen and whether different layers share memory. These are possible design questions, not demonstrated improvements to a workflow.
The authors say their literature coverage and release inventory are curated, not exhaustive. One of the 59 release records lacks separately disclosed attention architecture and is omitted from the classifiable-record visualization. The frozen endpoint comparison is a snapshot, not a causal test: model size, training, reasoning budget and implementation also affect leaderboard scores. For multimodal models, the authors classify only the documented language backbone, not vision or audio components. General PtoP note: this arXiv source is a preprint, and a survey of architectures is not evidence of performance in a deployed contract-review workflow.
FROM PAPER TO PRACTICE
Conceptual exercise only: installation and a runnable procedure are unverified from the supplied paper text. Prerequisites are a document you are allowed to inspect and time to mark it up; no model access or payment is needed.
An n8n workflow example is not appropriate here: the paper surveys internal model architectures and does not specify a verified workflow integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS
Why it matters to youIf you use specialist software for research, you may want to know whether a computer-using assistant leaves behind the right result—not just whether it clicks through the right menus. This paper offers tasks and checks for investigating that distinction.
A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an incorrect astronomy measurement, later entered the same value again, and checked the saved answer against itself rather than against what the measurement meant. To study such failures, the authors built OSWorld-Science: a collection of scientific tasks in real software, an environment that lets assistants operate a desktop, and task-specific checks of what they leave behind. The checks inspect outcomes such as files, measurements and application state, and can award credit for partially completed work.
A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an incorrect astronomy measurement, later entered the same value again, and checked the saved answer against itself rather than against what the measurement meant. To study such failures, the authors built OSWorld-Science: a collection of scientific tasks in real software, an environment that lets assistants operate a desktop, and task-specific checks of what they leave behind. The checks inspect outcomes such as files, measurements and application state, and can award credit for partially completed work.
Hypothetical illustration: An assistant is asked to record a measurement from an astronomy program. It saves a number in a workbook, then reopens the workbook to confirm that the same number is there. That confirms the save, not the measurement. The authors describe this distinction in a specific GPT-5.6-terra astronomy run; this hypothetical version is not an additional test result.
In the authors' overall comparison of evaluated models, Claude Fable 5.1 had the highest mean task score, 73.7%. That percentage averages task-specific scores, including partial credit; runs with no gradable outcome count as zero. It is not the percentage of tasks fully completed. Separately, the authors changed the assistant's configuration on the 23 QuPath pathology tasks. With Claude Opus 5 at medium reasoning effort and English prompts, making ten recent screenshots available produced a mean partial-credit score of 57.0%. The authors observed a lower score with a shorter history window, but no configuration cell differed significantly from its default after their stated correction for multiple comparisons. Each cell had one run per task, so the observed change should not be described as an established improvement.
The authors' examples show why a plausible sequence of actions is not enough evidence that scientific work is complete. A useful check must reach the delivered result and, where needed, the source of the scientific claim.
Possible use: a research team could use the paper's task descriptions and outcome checks to decide what an assistant must verify before handing over a scientific file. The study tests assistants on benchmark tasks; it does not establish a deployed research workflow.
The authors report uneven numbers of tasks across fields. Their detailed trajectory analysis covers released runs for 128 tasks and excludes biology; not every evaluated run has a released trajectory. Some software requires a licence, and the authors say they do not plan to publish all trajectories in this version. The QuPath configuration comparisons use one run per task and did not detect a significant difference from the default after correction. The paper also describes evaluation and infrastructure gaps in particular task analyses, so their counts should stay attached to the specific tasks and runs reported. General PtoP note: this arXiv source is a preprint, not evidence of peer review, and a benchmark result is not a deployment result.
FROM PAPER TO PRACTICE
Two routes follow. Installation and a successful run are unverified here.
An n8n workflow is not the useful first step here: the paper evaluates actions inside specialist desktop software and checks the resulting files and application state, rather than demonstrating a web-service integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS
Why it matters to youIf you help train a search assistant that writes its own practice questions, you need to know whether its rising training score reflects better answers or repeated mistakes. This paper offers a way to separate those signals in the authors’ experiments, not a tested workplace deployment.
A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part writes questions and draft answers from source documents, while another practices answering them. If the second part learns a wrong draft answer, its later agreement with that answer can be mistaken for progress. The authors call this unintentional shared-error pattern co-cheating. Their main method, CrossFit, divides documents into two groups. Questions from one group are scored by a separate answering model trained only on the other group. The main answering model still trains on admitted questions from both groups; only the feedback used to choose future practice questions changes.
A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part writes questions and draft answers from source documents, while another practices answering them. If the second part learns a wrong draft answer, its later agreement with that answer can be mistaken for progress. The authors call this unintentional shared-error pattern co-cheating. Their main method, CrossFit, divides documents into two groups. Questions from one group are scored by a separate answering model trained only on the other group. The main answering model still trains on admitted questions from both groups; only the feedback used to choose future practice questions changes.
Hypothetical illustration, not a paper result: A question-maker reads a document and drafts a question whose answer is mistakenly “Harbor A.” If its usual answering partner learned that draft from the same document, a later “Harbor A” response might earn credit. Under CrossFit, that question is instead scored by a partner that did not train on questions from that document. This removes that direct route to agreement; it does not guarantee the partner is right.
In the authors’ three-round Qwen3.5-4B experiments, CrossFit reduced final-round false-agreement mass from 6.1% with the standard coupled Dr. Zero loop to 3.0%. That percentage counts audited answer pairs that matched on the same incorrect answer. On a fixed set of 1,325 questions drawn from seven search benchmarks, CrossFit’s Qwen3.5-4B main model scored 48.8%, an 8.8-point improvement over the coupled version. This score counts questions whose normalized prediction contains a reference answer, then averages the seven benchmark scores equally. In the corresponding Qwen3.5-9B experiments, false-agreement mass fell from 8.8% to 3.7%, and the benchmark average rose by 8.4 points to 51.2%. The authors also report that replaying identical saved proposals with source-excluded feedback reduced false agreement, helping separate the feedback change from changes in which questions were generated.
The authors’ finding concerns a specific training problem: agreement is a poor stand-in for correctness when the model giving feedback has learned from the answer it is checking. Their comparison separates changing the feedback model’s training history from simply checking proposed answers more often.
A possible use is to design the feedback stage of a system that generates its own search-and-answer practice. The paper tests training and benchmark evaluation, not an operational assistant or an end-user workflow.
The authors used an automated, evidence-backed audit after training, not exhaustive human annotation. Unsupported cases remained unresolved; final-round audit coverage varied by model and treatment. CrossFit does not make its scoring model a source of guaranteed truth: shared pretraining, overlapping evidence and related documents can still produce shared mistakes. The authors also note that rejecting difficult questions could lower false agreement without improving learning. CrossFit added auxiliary training cost, and they do not establish lower end-to-end cost or robustness to connected sources. This source is an arXiv preprint; as a general PtoP note, its preprint status does not imply peer review.
FROM PAPER TO PRACTICE
Safe conceptual exercise; the supplied paper does not verify installation steps or a runnable code repository. Prerequisites: two short documents, paper or a spreadsheet, and time to check answers yourself. There is no software access or compute cost required for this exercise; implementing the paper’s training method would require resources not established here.
An n8n workflow is not appropriate here: the paper studies how models are trained and scored, not a verified automation workflow.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS
Why it matters to youIf you coordinate research, coding and design work, this paper may help you think about which tasks to delegate and which results must arrive first. Its clearest direct test is of the plan, not of a finished workplace project.
When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs. Raven is their proposed system. For example, a research summary might need to reach a coder before a designer can use the coder’s output; that sequence is an illustration, not a measured result. Raven’s Host Agent makes a task graph—a set of jobs connected by the order in which their outputs are needed—and assigns jobs to agents. The wider system also includes ways to retain context, reuse procedures and change the working instructions around an agent’s model.
When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs. Raven is their proposed system. For example, a research summary might need to reach a coder before a designer can use the coder’s output; that sequence is an illustration, not a measured result. Raven’s Host Agent makes a task graph—a set of jobs connected by the order in which their outputs are needed—and assigns jobs to agents. The wider system also includes ways to retain context, reuse procedures and change the working instructions around an agent’s model.
Hypothetical example: a small team wants a briefing and a web page about a software experiment. A planner could assign source-finding to a research agent, implementation to a coding agent and the final page to a design agent. The plan would make the page wait for the experiment’s results. This shows what a handoff means; it is not a Raven test result.
In the authors’ Multi-Agent Orchestration Benchmark, systems submitted plans without running the specialist agents. With the Qwen3.8-27B model, Raven’s reported Exact Match rate was 0.711, versus 0.607 for the strongest compared baseline. Exact Match counts scored plans whose specialist choices and required order of work agree with an accepted reference. The authors also report Raven ahead of both compared systems on all four planning measures with each of the two tested models. These are planning findings, not measures of finished project quality.
The paper separates choosing and ordering specialists from the harder question of whether their combined work succeeds. That distinction gives a working reader a way to examine an AI workflow’s handoffs, rather than treating a plausible-looking plan as a completed job.
A possible use is to sketch a multi-specialist job before running it: identify the required outputs, who could produce each one, and what evidence the next worker needs. The paper’s planning benchmark tests agreement with reference plans; it does not establish that such a plan will complete a real project.
The authors say an alternative valid plan can differ from the benchmark’s reference plan. The planning test does not dispatch workers, and its nominal 140 requests do not by themselves determine every reported score’s denominator. Their theory gives sufficient conditions for reliable collaboration, but whether an implemented system meets them requires evaluation; it does not show that collaboration always beats one agent. Other evaluations have separate qualifications: Raven-Research’s comparison used questions also seen during development, and public benchmark copies were accessible; some Raven-Design slide configurations did not cover every task, and failed generations were excluded from reported means. General PtoP note: this arXiv source is a preprint, not evidence of peer review or a workplace deployment.
FROM PAPER TO PRACTICE
A repository exploration, not a reproduction of the paper’s benchmark. Prerequisites: Git, Docker with Docker Compose, and a machine that can run containers; model or service access and any associated cost are not established by these README steps. Raven is described in its README as pre-alpha.
n8n is not necessary here: the paper’s relevant test scores plans submitted through the systems’ own orchestration interfaces, not an n8n workflow.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS
Why it matters to youIf you work with video, image tools or robots, this paper offers a way to think about predicting what changes next. The authors tested that idea in research evaluations, not in a workplace deployment.
The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to represent what is happening, then tested whether that representation could help produce written answers, images and robot actions. Imagine seeing a hand beside a spoon, then seeing the spoon lifted. Orca’s video exercise asks it to predict an internal description of the next view, rather than draw every pixel. A second exercise gives it a description of an event and asks it to predict a view associated with that event. A third uses video questions and answers. The authors call these first two approaches “unconscious” and “conscious” learning; those names describe training methods, not awareness. After this training, the authors kept Orca’s main model fixed. They trained separate readouts—parts that turn its internal representation into an image or robot action—and used its existing language output for answers. This setup tests whether the learned representation is useful beyond a single output task.
The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to represent what is happening, then tested whether that representation could help produce written answers, images and robot actions. Imagine seeing a hand beside a spoon, then seeing the spoon lifted. Orca’s video exercise asks it to predict an internal description of the next view, rather than draw every pixel. A second exercise gives it a description of an event and asks it to predict a view associated with that event. A third uses video questions and answers. The authors call these first two approaches “unconscious” and “conscious” learning; those names describe training methods, not awareness. After this training, the authors kept Orca’s main model fixed. They trained separate readouts—parts that turn its internal representation into an image or robot action—and used its existing language output for answers. This setup tests whether the learned representation is useful beyond a single output task.
Hypothetical illustration, not a paper result: a kitchen worker supplies a picture of a closed drawer and the instruction “open the drawer.” An image readout might attempt to show the drawer open. The useful question is whether it also keeps the surrounding counter and objects consistent; the paper does not test this workplace use.
In the authors’ real-robot tests, Orca-4B averaged 36.6 rule-based points across five tasks with unfamiliar tablecloth or background settings, versus 27.6 for π 0.5, the compared robot-control system pretrained on large-scale robot data. Rule-based points mark the highest task stage reached before a trial ends; they are not a success percentage. In the separate unfamiliar-object setting, Orca-4B averaged 28.2 points versus 31.2 for π 0.5. The authors trained the Orca action readout on 200 robot trajectories per task. In a qualitative spoon-grasp example, they report Orca recovering from early failures while π 0.5 made repeated failed attempts. For image prediction, the authors report a 59.8 average percentage score for Orca-4B with its image readout on PRICE-V0.1, versus 56.1 for FLUX.2 [klein]. That score comes from four model judges rating generated images against an initial image and instruction. The authors also report improved text, image and action readout results as Orca’s pre-training scales up.
The authors test whether learning about changes in video can be useful when the output is a sentence, an image or a robot action. That is different from training a separate main model solely to produce each kind of output. The results concern the authors’ stated evaluations, not general reliability in everyday settings.
Possible uses suggested by the evaluated tasks include studying video understanding, predicting how an instructed interaction might look, and researching robot manipulation. These are possibilities, not demonstrated products or deployments.
The authors describe Orca as an early step. It learns mainly from vision and language, not sound, touch or force. Its visual prediction target comes from a frozen, pretrained vision component rather than a representation learned directly from all possible signals. The experiments mainly use the 0.8B and 4B sizes, and this version trains on only one-tenth of the authors’ 125K-hour video inventory. The authors report a trade-off among readout performances during training. They say their image benchmark has limited scale and variety, most annotated events cover short, minute-level changes, and the robot tasks remain relatively short and easy. The paper says no benchmark-specific training data was constructed or used, but does not establish that every underlying video source is disjoint from every evaluation source. General PtoP note: this arXiv paper is a preprint, and a benchmark result is not a deployment test.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and public access to Orca are unverified from the supplied paper. Prerequisites: two pictures you may use and a short description of the change between them. No paid service is needed for this paper-and-pencil exercise.
An n8n automation workflow is not appropriate here: the supplied paper does not verify a public Orca endpoint or an integration to connect to it.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS
Why it matters to youIf you oversee an AI agent doing a long coding or operations job, you may want a way to see what it has finished and what it must check before stopping. This preprint studies a shared runbook for doing that; it does not test a workplace deployment.
The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does not change. The authors built StateM, a runtime that keeps the current phase of work and a shared runbook outside the agent’s long conversation history. Imagine an agent setting up a web service: it can work freely during setup, but the runbook can require a working end-to-end check before the agent moves to handoff. StateM records the phase, refreshes relevant instructions, checks configured conditions at transitions, and leaves a route back for repair. The authors developed one runbook with GPT-5.5, tested it unchanged with GPT-5.6 variants, adapted its practices for DeepSeek-V4-Flash, and developed separate runbooks for BusinessBench task families.
The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does not change. The authors built StateM, a runtime that keeps the current phase of work and a shared runbook outside the agent’s long conversation history. Imagine an agent setting up a web service: it can work freely during setup, but the runbook can require a working end-to-end check before the agent moves to handoff. StateM records the phase, refreshes relevant instructions, checks configured conditions at transitions, and leaves a route back for repair. The authors developed one runbook with GPT-5.5, tested it unchanged with GPT-5.6 variants, adapted its practices for DeepSeek-V4-Flash, and developed separate runbooks for BusinessBench task families.
Hypothetical illustration, not a reported test: an agent updates an internal service. Its runbook keeps it in a verification phase until a configured check confirms that another system can use the service. If the check fails, the agent returns to repair rather than declaring the job done.
On Terminal-Bench 2.1, an 89-task terminal benchmark, the authors report 92.1% trial-level success for GPT-5.5 xhigh with StateM, against an 83.1% published GPT-5.5 reference. Trial-level success counts completed trials, with five trials per task. With the runbook frozen, GPT-5.6 Sol xhigh with StateM recorded 424 successful trials out of 445, or 95.28% raw accuracy, in a public submission. That score is before review is settled, not a finalized leaderboard result. For a different model and setting, the adapted DeepSeek-V4-Flash system scored 392 successes out of 445 under standard timeouts; the authors report $15.20 in realized API charges for its final-score evidence, excluding $37.02 spent on adaptation. On BusinessBench, the first frozen held-out evaluation with Codex and GPT-5.6 Luna showed smaller aggregate gains: 0.55 percentage points when task families were weighted equally and 1.34 points when instances were weighted equally.
The authors’ results separate a model’s ability to perform individual steps from the execution system’s ability to carry a job through to completion. They also show a boundary: an unchanged runbook developed with GPT models did not improve their DeepSeek-V4-Flash result, while an adapted one did.
Possible applications include tracking obligations during long coding jobs and placing checks before consequential operations handoffs. These are potential uses of the control approach, not deployments demonstrated by the paper.
The Terminal-Bench gains measure StateM’s runtime together with an evolved, benchmark-adapted runbook; they do not isolate the runtime alone. The GPT-5.5 comparison uses a published reference, not a newly rerun control matched for agent version. The GPT-5.6 Sol xhigh public submission remained open and unmerged at the stated review date. The authors say four rewarded trials should not count and that nine trajectories were flagged for possible reward hacking; the raw 95.28% must therefore remain qualified. Solving a task at least once in five trials is not single-run reliability. The unchanged GPT-developed runbook reduced DeepSeek-V4-Flash’s full-suite score from 82.7% to 82.0%; its later improvement required adaptation. BusinessBench has one trajectory per instance per arm, excludes the untreated Attendance family from intervention aggregates, and shows declines in two treated families in its first frozen evaluation. Later refinements reused evaluation feedback and are not untouched held-out results. StateM cannot recover unrecorded work, guarantee that a self-reported check is correct, or undo arbitrary external actions. PtoP note: this source is an arXiv preprint, not a claim of peer review; benchmark scores are not evidence of workplace deployment.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation is unverified because the supplied text names a code release but provides no verified official repository link or installation commands. Prerequisites: a multistep task you understand and a place to write a checklist. No model access or API spending is needed for this paper-based exercise.
A proposed n8n integration could pause a workflow before handoff and request a human approval when a required check fails. This is an illustration, not a tested StateM feature or a verified integration procedure.
Was this explanation easy to understand?
COMPARISON
Compare topics, possible uses and limits. Each result remains attributed to its own paper.
| Paper | Central idea | Possible use | Limit |
|---|---|---|---|
| Long-document AI design depends on what models keep and retrieve | A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The authors explain that designs make… | The survey’s framework could help people planning long-document AI systems ask what evidence remains available, how retrieval candidates… | The authors say their literature coverage and release inventory are curated, not exhaustive. One of the 59 release records lacks separately… |
| Test whether an assistant delivers the right scientific result | A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an… | Possible use: a research team could use the paper's task descriptions and outcome checks to decide what an assistant must verify before… | The authors report uneven numbers of tasks across fields. Their detailed trajectory analysis covers released runs for 128 tasks and… |
| Keep Search Agents From Rewarding Their Own Wrong Answers | A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part writes questions and draft answers… | A possible use is to design the feedback stage of a system that generates its own search-and-answer practice. The paper tests training and… | The authors used an automated, evidence-backed audit after training, not exhaustive human annotation. Unsupported cases remained… |
| Plan specialist AI work before handing tasks to agents | When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs. Raven is their proposed system… | A possible use is to sketch a multi-specialist job before running it: identify the required outputs, who could produce each one, and what… | The authors say an alternative valid plan can differ from the benchmark’s reference plan. The planning test does not dispatch workers, and… |
| Learning scene changes could help AI predict images and robot actions | The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to represent what is happening, then… | Possible uses suggested by the evaluated tasks include studying video understanding, predicting how an instructed interaction might look… | The authors describe Orca as an early step. It learns mainly from vision and language, not sound, touch or force. Its visual prediction… |
| A shared runbook may help AI agents finish long tasks | The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does not change. The authors built… | Possible applications include tracking obligations during long coding jobs and placing checks before consequential operations handoffs… | The Terminal-Bench gains measure StateM’s runtime together with an evolved, benchmark-adapted runbook; they do not isolate the runtime… |
ARCHIVE
Find earlier editions of the newspaper.
01