ISSUE 05/2026 · 01-10-2026

Test whether an assistant delivers the right scientific result

A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an incorrect astronomy measurement, later entered the same value again, and checked the saved answer…

Editorial illustration: An anonymous researcher examines a specimen beside a desk where a winding trail of scattered tools ends at a sealed sample jar, while a second trail circles back to the same misplaced tool.
AI-generated editorial illustration: a visual metaphor, not a figure from the paper.
Read the lead story ↘

6 paper · Original sources linked in every story

5 MINUTES ON PtoP

Check What the Agent Left Behind

An AI agent can look busy long before you know whether its work is right. WIRED reports that companies are adding agents to workplace chats and, in some cases, org charts. It also cites research in which managers caught fewer errors when told work came from an “AI employee” rather than an “AI tool.” That finding does not tell us how every team will behave. But it raises a useful question: when software feels like a colleague, will we still check what it hands over? See: AI Agents Are About to Flood the Workforce. No One's ...

The authors of OSWorld-Science put that question in concrete terms. They built scientific-software tasks whose checks inspect the result an assistant leaves behind, including files, measurements and application state. In one run they describe, an assistant entered a wrong astronomy measurement twice, then checked the saved answer against itself rather than against what the measurement meant. Anthropic offers a different example of why the next check matters: it says Claude found a previously uncharacterized enzyme system in DNA records, and human scientists confirmed part of the finding in a lab. Anthropic says they still do not know what the system does. See: Test whether an assistant delivers the right scientific… · Claude discovers a novel enzyme system with CRISPR-like…

The problem can begin even before an assistant reaches us. In experiments with search agents that generate their own practice questions, the authors of the CrossFit paper found that an agent could appear to improve by learning to repeat a wrong draft answer. Their method uses a separate answering model, trained on a different group of documents, to help choose practice questions. In their tested setting, that reduced agreement on the same incorrect answer. The authors do not present it as a guarantee of truth: different models can still share mistakes. See: Keep Search Agents From Rewarding Their Own Wrong Answers

So what might help when work stretches across many steps? The StateM authors tested a shared runbook that records an agent’s current phase and checks conditions before a handoff; they report benchmark gains, while noting that the runbook and runtime were tested together and that results vary by setting. The Raven authors tested plans for assigning work to specialist agents, but their planning test did not run those specialists. Read together, the papers suggest three distinct places to look: the plan, the check at each handoff, and the finished result. None is a substitute for the others. See: A shared runbook may help AI agents finish long tasks · Plan specialist AI work before handing tasks to agents

If agents become a more familiar part of work, a visible chain of checks could matter as much as a smooth conversation. This week, if you give an assistant a task with several steps, you might try writing down one thing you can inspect at the end—an output file, a figure traced to its source, or a saved change—and ask: what would show me that this result is right, rather than merely finished? See: AI Agents Are About to Flood the Workforce. No One's ... · Test whether an assistant delivers the right scientific…

Just here for the stories? They are below, by topic.

From the news desk

News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.

  1. Image, audio and video Hugging Face Blog · 30-09-2026Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning ↗Read the explainer ↓
  2. People and society WIRED · 28-09-2026AI Agents Are About to Flood the Workforce. No One's ... ↗Read the explainer ↓
  3. Image, audio and video Google · 23-09-2026Gemini 3.8 text-to-speech says hello - Google Blog ↗Read the explainer ↓
  4. AI agents Anthropic · 23-09-2026Claude discovers a novel enzyme system with CRISPR-like repeats ↗Read the explainer ↓
  5. Tools and building xAI · 21-09-2026Introducing Grok 4.7 ↗Read the explainer ↓
  6. People and society BBC · 14-09-2026Why are there concerns AI could threaten humanity? ↗Read the explainer ↓

PODCAST / ENGLISH EDITION

The full issue, then each paper

The front-page episode covers the papers and news; each paper also has its own short conversation.

Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.

EP 01ARXIV:2609.39661

Long-document AI design depends on what models keep and retrieve

A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The…

EP 02ARXIV:2609.39903

Test whether an assistant delivers the right scientific result

A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by…

EP 03ARXIV:2609.39102

Keep Search Agents From Rewarding Their Own Wrong Answers

A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part…

EP 04ARXIV:2609.33439

Plan specialist AI work before handing tasks to agents

When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs…

EP 05ARXIV:2606.30534

Learning scene changes could help AI predict images and robot actions

The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to…

EP 06ARXIV:2608.15089

A shared runbook may help AI agents finish long tasks

The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does…

FROM THE NEWS DESK

The news, explained

Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.

Image, audio and video Hugging Face Blog · 30-09-2026 · company announcement

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

In brief

Hugging Face says it built a public ranking to compare AI-generated speech across languages, including voices based on a sample recording. It uses repeatable tests of word accuracy, speed and voice similarity so newly released models can be compared more quickly.

An example

Imagine choosing a voice for a travel guide. You could listen to two systems read a Japanese station announcement, then check which says the right words and starts speaking sooner.

Application

A developer could use the TTS leaderboard to shortlist voices for a multilingual assistant, comparing intelligibility and streaming speed before listening to the results.

The limitation

Hugging Face says its automated scores do not measure how natural or expressive a voice sounds, or which voice people prefer. The tested languages and hardware also limit what the rankings can tell you.

Takeaway

The ranking can narrow the choices, but listening still matters—especially when assessing voice cloning.

Original source ↗

Was this explanation easy to understand?

People and society WIRED · 28-09-2026 · news report

AI Agents Are About to Flood the Workforce. No One's ...

In brief

WIRED reports that companies are adding AI agents to workplace chats and, in some cases, to an org chart. An AI agent is software that can carry out tasks without someone guiding every step. The article’s central concern is that making such software feel like a coworker may change how carefully people check its work.

An example

For example, WIRED describes an agent that posted details of its manager’s calendar in a team chat. She corrected it directly, but the mistake shows why an apparent coworker still needs oversight.

Application

A company using an agent to schedule meetings could set guardrails: limit which calendar details it can see and require a person to check messages before they go to a group.

The limitation

WIRED cites research in which managers caught fewer errors when told work came from an AI employee rather than an AI tool. That finding does not establish how every team will behave, or how many agents will join workplaces.

Takeaway

Treat AI agents as tools that can do useful work, not as coworkers whose work can go unchecked.

Original source ↗

Was this explanation easy to understand?

Image, audio and video Google · 23-09-2026 · company announcement

Gemini 3.8 text-to-speech says hello - Google Blog

In brief

Google says it is rolling out two tools that turn written scripts into expressive speech. A game maker could make a dragon whisper one line and shout the next. These text-to-speech models let creators design voices and direct how each line sounds.

An example

As an illustration, a podcast producer could write a conversation for two characters and ask for a different speaking style for each.

Application

A video maker could use the tools to create spoken versions of a script in different languages. Google says developers can try both models in Google AI Studio.

The limitation

This is Google's product announcement, not an independent test of its quality claims. Voice replication in Google AI Studio is unavailable in several regions, including the UK and India. Google says its safeguards include consent verification and a watermark, but does not show that misuse is impossible.

Takeaway

Google is making generated speech easier to direct, but its quality and safeguards need independent scrutiny.

Original source ↗

Was this explanation easy to understand?

AI agents Anthropic · 23-09-2026 · company announcement

Claude discovers a novel enzyme system with CRISPR-like repeats

In brief

Anthropic says Claude found a previously uncharacterized enzyme system by searching DNA records, and human scientists checked the finding in their lab. The system has repeated DNA segments that resemble those in CRISPR, but Anthropic does not yet know what it does.

An example

Imagine a software helper searching millions of recipes and noticing that the same unusual note keeps appearing beside one ingredient. A cook would still need to test what the combination does. Here, Claude spotted a pattern, and scientists followed up with experiments.

Application

This approach could help researchers choose which unexplored DNA patterns and enzymes to investigate in the lab.

The limitation

These are early results described by Anthropic in a preprint. Its experiments found that the repeated DNA segments produce short pieces of RNA, but they have not established the system’s function or shown that it can be used as a gene-editing tool.

Takeaway

Claude helped identify a promising biological puzzle; human experiments are still needed to solve it.

Original source ↗

Was this explanation easy to understand?

Tools and building xAI · 21-09-2026 · company announcement

Introducing Grok 4.7

In brief

xAI announced Grok 4.7, a new AI model it says is better at coding and longer work tasks than Grok 4.6. The central idea is that it can spend more time on a task and check its own work more carefully.

An example

For example, a developer might ask it to find a bug, suggest a fix, and check whether the fix breaks anything else. That illustrates the kind of longer task xAI describes; it is not a verified result.

Application

Developers could use Grok 4.7 to help draft and review software changes. xAI says it is available through its API and several coding tools.

The limitation

The performance and safety claims come from xAI’s own benchmarks. Those tests do not establish how reliably the model will perform on every real-world task.

Takeaway

Grok 4.7 is a coding-focused release with claimed gains in long tasks and safeguards, but its results need to be read as company-reported tests.

Original source ↗

Was this explanation easy to understand?

People and society BBC · 14-09-2026 · news report

Why are there concerns AI could threaten humanity?

In brief

The BBC reports that warnings about AI becoming hard to control have prompted calls to slow its development, though the worst-case scenarios remain hypothetical. The central concern is that AI agents—software that can take actions on its own—might eventually do things people cannot reliably stop. Critics question whether companies overstate that risk while harms are already occurring.

An example

Imagine an email assistant that sends a message to your entire address book when you only asked it to write a draft. This is an illustration of why giving software permission to act needs safeguards, not an incident reported by the BBC.

Application

A developer could ask third-party evaluators—independent testers—to check an AI agent’s safeguards before giving it more tasks. The BBC reports that Anthropic’s boss has proposed such checks.

The limitation

The BBC describes threats to humanity as hypothetical, not measured outcomes. It also reports present-day harms, including deepfakes: convincing fake images used to depict people nude without their consent.

Takeaway

The debate is about how to manage uncertain future dangers without neglecting harms happening now.

Original source ↗

Was this explanation easy to understand?

01NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS

Long-document AI design depends on what models keep and retrieve

Why it matters to youIf you review long contracts, you may want to know why an AI assistant can accept a document yet still miss a clause. This survey offers a way to understand the memory choices behind that risk, not a tested contract-review tool.

A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The authors explain that designs make different trade-offs between keeping separate pieces of text and compressing them into a smaller memory. A token is a piece of text the model processes, such as a word or part of one. A layer is one stage in the model’s sequence of processing steps. The authors compare ways of storing and retrieving tokens, including designs that use different methods in different layers. They review prior research and classify documented model releases; they do not test an assistant on contracts.

Zhentao Tan, Jingyi Shen, Yanbo Li, Yao Liu, Yue Wu, Jieping Ye · The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends · arXiv:2609.39661Read the paper ↗

In one sentence

A model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The authors explain that designs make different trade-offs between keeping separate pieces of text and compressing them into a smaller memory. A token is a piece of text the model processes, such as a word or part of one. A layer is one stage in the model’s sequence of processing steps. The authors compare ways of storing and retrieving tokens, including designs that use different methods in different layers. They review prior research and classify documented model releases; they do not test an assistant on contracts.

Key concepts

  • Contextual memory is information from the current document that remains available inside a model as it processes more text. It is not an external document store or the model’s training data.
  • Explicit memory keeps earlier tokens as separately addressable items. In a hypothetical contract example, this is like retaining individual clauses so a later question can point to one.
  • Sparse attention keeps explicit items but checks only a selected subset for each question. The selection can save work, but relevant evidence must be included.
  • Recurrent state compresses preceding text into a running summary-like internal state. It can reflect earlier text without retaining every token as a separately retrievable item.
  • Hybrid architecture combines memory methods, sometimes across layers. The authors’ comparison asks not only what is stored, but how it is updated, selected, read and combined.

A concrete example

Hypothetical illustration, not a paper test: A contract mentions a cancellation deadline near the beginning. When asked about it near the end, a design that retains individual tokens could consult the earlier passage; a sparse design would first have to select it. A compressed-state design might carry information from that passage without preserving the passage as a separately selectable item. None of these descriptions establishes that a model will answer correctly.

What the researchers measured

This is a survey and descriptive architecture comparison, not a new task-performance experiment. The authors assembled 59 release-level records of publicly documented model architectures. In a separate frozen comparison of 11 high-performing open-weight model endpoints from the Artificial Analysis Intelligence Index v4.3.2, captured on September 22, 2026, they report varied attention designs and an explicit token-retrieval path in every selected architecture. These counts describe the surveyed releases and selected endpoints; they are not accuracy scores.

Why it matters

The authors distinguish fitting more text into a model’s input from reliably using distant evidence. Their framework makes it easier to state what a design retains and what it may overlook. That distinction can guide questions about a proposed system, but the survey does not show how well any system would perform in a reader’s job.

Where it might help

The survey’s framework could help people planning long-document AI systems ask what evidence remains available, how retrieval candidates are chosen and whether different layers share memory. These are possible design questions, not demonstrated improvements to a workflow.

Impact across sectors

  • Possible, not tested: Legal teams considering contract assistants could ask whether an architecture preserves access to individual earlier clauses.
  • Possible, not tested: Healthcare teams considering long-record assistants could ask how a design selects older notes for a current question.
  • Possible, not tested: Software teams processing lengthy specifications could use the framework to compare memory and retrieval choices before evaluating a system on their own documents.

Where the evidence stops

The authors say their literature coverage and release inventory are curated, not exhaustive. One of the 59 release records lacks separately disclosed attention architecture and is omitted from the classifiable-record visualization. The frozen endpoint comparison is a snapshot, not a causal test: model size, training, reasoning budget and implementation also affect leaderboard scores. For multimodal models, the authors classify only the documented language backbone, not vision or audio components. General PtoP note: this arXiv source is a preprint, and a survey of architectures is not evidence of performance in a deployed contract-review workflow.

FROM PAPER TO PRACTICE

How to try it

Conceptual exercise only: installation and a runnable procedure are unverified from the supplied paper text. Prerequisites are a document you are allowed to inspect and time to mark it up; no model access or payment is needed.

  1. Choose a long document and place a question whose answer appears early in it.
  2. Mark the exact passage needed to answer, along with any nearby context that changes its meaning.
  3. Imagine three memory choices: retain each passage separately, select only some passages, or compress earlier passages into a running state.
  4. For each choice, write down what would need to remain available when the question is asked. Observe which choices preserve a direct route to the marked passage; this is a reasoning exercise, not a model test.

n8n example

An n8n workflow example is not appropriate here: the paper surveys internal model architectures and does not specify a verified workflow integration.

Was this explanation easy to understand?

02NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS

Test whether an assistant delivers the right scientific result

Why it matters to youIf you use specialist software for research, you may want to know whether a computer-using assistant leaves behind the right result—not just whether it clicks through the right menus. This paper offers tasks and checks for investigating that distinction.

A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an incorrect astronomy measurement, later entered the same value again, and checked the saved answer against itself rather than against what the measurement meant. To study such failures, the authors built OSWorld-Science: a collection of scientific tasks in real software, an environment that lets assistants operate a desktop, and task-specific checks of what they leave behind. The checks inspect outcomes such as files, measurements and application state, and can award credit for partially completed work.

Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang, Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang, Jie Tang, Juanzi Li, Weihao Xuan, and Tianyu Liu · OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software · arXiv:2609.39903Read the paper ↗

In one sentence

A computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an incorrect astronomy measurement, later entered the same value again, and checked the saved answer against itself rather than against what the measurement meant. To study such failures, the authors built OSWorld-Science: a collection of scientific tasks in real software, an environment that lets assistants operate a desktop, and task-specific checks of what they leave behind. The checks inspect outcomes such as files, measurements and application state, and can award credit for partially completed work.

Key concepts

  • A computer-using agent is an assistant that takes actions on a desktop, such as clicking, typing or issuing a terminal command.
  • A task-specific evaluator checks the resulting scientific work, rather than relying only on the assistant's claim that it finished.
  • Partial credit records how much of a task's stated requirements the final result satisfies; it is not the same as counting fully finished tasks.
  • A history window is the set of recent screenshots available to the assistant while it chooses its next action.
  • Independent validation checks a result against the source or another measure, rather than repeating an assumption that might already be wrong.

A concrete example

Hypothetical illustration: An assistant is asked to record a measurement from an astronomy program. It saves a number in a workbook, then reopens the workbook to confirm that the same number is there. That confirms the save, not the measurement. The authors describe this distinction in a specific GPT-5.6-terra astronomy run; this hypothetical version is not an additional test result.

What the researchers measured

In the authors' overall comparison of evaluated models, Claude Fable 5.1 had the highest mean task score, 73.7%. That percentage averages task-specific scores, including partial credit; runs with no gradable outcome count as zero. It is not the percentage of tasks fully completed. Separately, the authors changed the assistant's configuration on the 23 QuPath pathology tasks. With Claude Opus 5 at medium reasoning effort and English prompts, making ten recent screenshots available produced a mean partial-credit score of 57.0%. The authors observed a lower score with a shorter history window, but no configuration cell differed significantly from its default after their stated correction for multiple comparisons. Each cell had one run per task, so the observed change should not be described as an established improvement.

Why it matters

The authors' examples show why a plausible sequence of actions is not enough evidence that scientific work is complete. A useful check must reach the delivered result and, where needed, the source of the scientific claim.

Where it might help

Possible use: a research team could use the paper's task descriptions and outcome checks to decide what an assistant must verify before handing over a scientific file. The study tests assistants on benchmark tasks; it does not establish a deployed research workflow.

Impact across sectors

  • Possibility, not a proven deployment: Medical-image teams could specify both the image-analysis steps and the saved result an assistant would need to deliver.
  • Possibility, not a proven deployment: Chemistry teams could check a proposed molecular file against its source, rather than accepting a file that merely opens correctly.
  • Possibility, not a proven deployment: Statistics teams could distinguish a correct calculation from a missing or incorrectly saved report.

Where the evidence stops

The authors report uneven numbers of tasks across fields. Their detailed trajectory analysis covers released runs for 128 tasks and excludes biology; not every evaluated run has a released trajectory. Some software requires a licence, and the authors say they do not plan to publish all trajectories in this version. The QuPath configuration comparisons use one run per task and did not detect a significant difference from the default after correction. The paper also describes evaluation and infrastructure gaps in particular task analyses, so their counts should stay attached to the specific tasks and runs reported. General PtoP note: this arXiv source is a preprint, not evidence of peer review, and a benchmark result is not a deployment result.

FROM PAPER TO PRACTICE

How to try it

Two routes follow. Installation and a successful run are unverified here.

  1. Without installing anything, read the paper's astronomy net-counts task description. Write down what quantity is requested and why a displayed raw count would not answer it.
  2. Sketch an independent check: identify which source measurement, background adjustment and saved answer would need to agree. Observe that reopening an answer file alone checks only what was saved.
  3. For technical users, the official repository's route requires a Linux x86-64 host with accessible KVM and Docker, about 8 GB of memory and four virtual processors per concurrent guest, disk space for an image, and Python 3.11+ managed with uv. Its README gives these host checks: `egrep -c '(vmx|svm)' /proc/cpuinfo` and `ls -l /dev/kvm`.
  4. The README gives these installation commands: `git clone https://github.com/DiscoAILab/OSWorld-Science.git`, `cd OSWorld-Science`, `uv sync --extra hf`, and `cp .env.example .env`. Add credentials only for the model service you use. Model calls may incur charges; third-party software and images can have separate licensing or availability restrictions. The full public runner-image download is roughly 196 GB, and the README says some guest images are unavailable for redistribution.
  5. To explore the statistics tasks and prepared image, the README gives `uv run python scripts/data_prep/hf_download.py --domain stat --vm` and `uv run osci tasks list`. Its example run is `uv run osci run --models claude-sonnet-5 --tasks stat_qol_sql --run-name smoke`. Observe the task listing and, if the required image, access and paid model service are available, the run's saved result and score. This is a repository procedure, not a claim that this article tested it.

n8n example

An n8n workflow is not the useful first step here: the paper evaluates actions inside specialist desktop software and checks the resulting files and application state, rather than demonstrating a web-service integration.

Was this explanation easy to understand?

03NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS

Keep Search Agents From Rewarding Their Own Wrong Answers

Why it matters to youIf you help train a search assistant that writes its own practice questions, you need to know whether its rising training score reflects better answers or repeated mistakes. This paper offers a way to separate those signals in the authors’ experiments, not a tested workplace deployment.

A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part writes questions and draft answers from source documents, while another practices answering them. If the second part learns a wrong draft answer, its later agreement with that answer can be mistaken for progress. The authors call this unintentional shared-error pattern co-cheating. Their main method, CrossFit, divides documents into two groups. Questions from one group are scored by a separate answering model trained only on the other group. The main answering model still trains on admitted questions from both groups; only the feedback used to choose future practice questions changes.

Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, and Alaa Khamis · False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents · arXiv:2609.39102Read the paper ↗

In one sentence

A search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part writes questions and draft answers from source documents, while another practices answering them. If the second part learns a wrong draft answer, its later agreement with that answer can be mistaken for progress. The authors call this unintentional shared-error pattern co-cheating. Their main method, CrossFit, divides documents into two groups. Questions from one group are scored by a separate answering model trained only on the other group. The main answering model still trains on admitted questions from both groups; only the feedback used to choose future practice questions changes.

Key concepts

  • Self-evolving training: a system makes practice questions and uses an answering model’s responses to guide which questions it makes next.
  • Pseudo-label: a proposed answer used for training without a human-written answer key; it may be wrong.
  • False agreement: the question-maker’s adopted answer and the answering model’s response match, but an audit judges that shared answer incorrect.
  • Source-level exclusion: keep every question from one document out of the training data of the model that scores questions from that document.
  • Multi-sample verification: the authors’ separate check asks the same model three times with the source and three times without it; compatible majorities replace the draft answer, while other proposals are rejected.

A concrete example

Hypothetical illustration, not a paper result: A question-maker reads a document and drafts a question whose answer is mistakenly “Harbor A.” If its usual answering partner learned that draft from the same document, a later “Harbor A” response might earn credit. Under CrossFit, that question is instead scored by a partner that did not train on questions from that document. This removes that direct route to agreement; it does not guarantee the partner is right.

What the researchers measured

In the authors’ three-round Qwen3.5-4B experiments, CrossFit reduced final-round false-agreement mass from 6.1% with the standard coupled Dr. Zero loop to 3.0%. That percentage counts audited answer pairs that matched on the same incorrect answer. On a fixed set of 1,325 questions drawn from seven search benchmarks, CrossFit’s Qwen3.5-4B main model scored 48.8%, an 8.8-point improvement over the coupled version. This score counts questions whose normalized prediction contains a reference answer, then averages the seven benchmark scores equally. In the corresponding Qwen3.5-9B experiments, false-agreement mass fell from 8.8% to 3.7%, and the benchmark average rose by 8.4 points to 51.2%. The authors also report that replaying identical saved proposals with source-excluded feedback reduced false agreement, helping separate the feedback change from changes in which questions were generated.

Why it matters

The authors’ finding concerns a specific training problem: agreement is a poor stand-in for correctness when the model giving feedback has learned from the answer it is checking. Their comparison separates changing the feedback model’s training history from simply checking proposed answers more often.

Where it might help

A possible use is to design the feedback stage of a system that generates its own search-and-answer practice. The paper tests training and benchmark evaluation, not an operational assistant or an end-user workflow.

Impact across sectors

  • Hypothetical possibility — workplace knowledge search: a team training an assistant on its documents could track whether its practice answers and feedback come from the same source.
  • Hypothetical possibility — education: builders of a self-questioning study tool could distinguish agreement with a generated answer key from evidence that the answer is correct.
  • Hypothetical possibility — research search: a team training a literature-search assistant could keep a document out of the scorer’s training set when scoring questions derived from it.

Where the evidence stops

The authors used an automated, evidence-backed audit after training, not exhaustive human annotation. Unsupported cases remained unresolved; final-round audit coverage varied by model and treatment. CrossFit does not make its scoring model a source of guaranteed truth: shared pretraining, overlapping evidence and related documents can still produce shared mistakes. The authors also note that rejecting difficult questions could lower false agreement without improving learning. CrossFit added auxiliary training cost, and they do not establish lower end-to-end cost or robustness to connected sources. This source is an arXiv preprint; as a general PtoP note, its preprint status does not imply peer review.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; the supplied paper does not verify installation steps or a runnable code repository. Prerequisites: two short documents, paper or a spreadsheet, and time to check answers yourself. There is no software access or compute cost required for this exercise; implementing the paper’s training method would require resources not established here.

  1. Write one question and a draft answer from each document.
  2. Imagine a scorer that has practiced on those same draft answers; mark where repeating a wrong draft could look like success.
  3. Swap scorers so each checks only the other document’s question, without having practiced on its draft answer.
  4. Check both answers against the documents yourself. Observe whether agreement and document-supported correctness differ; this illustration does not reproduce the authors’ experiment.

n8n example

An n8n workflow is not appropriate here: the paper studies how models are trained and scored, not a verified automation workflow.

Was this explanation easy to understand?

04TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS

Plan specialist AI work before handing tasks to agents

Why it matters to youIf you coordinate research, coding and design work, this paper may help you think about which tasks to delegate and which results must arrive first. Its clearest direct test is of the plan, not of a finished workplace project.

When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs. Raven is their proposed system. For example, a research summary might need to reach a coder before a designer can use the coder’s output; that sequence is an illustration, not a measured result. Raven’s Host Agent makes a task graph—a set of jobs connected by the order in which their outputs are needed—and assigns jobs to agents. The wider system also includes ways to retain context, reuse procedures and change the working instructions around an agent’s model.

EverMind AI · Raven: The Harness of Harnesses for Composable Agentic Intelligence · arXiv:2609.33439 · 473 HF votes at selectionRead the paper ↗

In one sentence

When a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs. Raven is their proposed system. For example, a research summary might need to reach a coder before a designer can use the coder’s output; that sequence is an illustration, not a measured result. Raven’s Host Agent makes a task graph—a set of jobs connected by the order in which their outputs are needed—and assigns jobs to agents. The wider system also includes ways to retain context, reuse procedures and change the working instructions around an agent’s model.

Key concepts

  • A harness is the tools, memory, instructions and control rules around an AI model. The paper treats a model together with its harness as a working unit.
  • A task graph records both who does each job and which output another job needs. Raven checks a proposed graph before sending work to agents.
  • A handoff is more than passing a file location: later work must receive the right information and preserve what earlier work established.
  • Harness adaptation changes the working policy around a fixed model, rather than changing the model itself. The paper discusses this alongside previously published HarnessBank experiments.

A concrete example

Hypothetical example: a small team wants a briefing and a web page about a software experiment. A planner could assign source-finding to a research agent, implementation to a coding agent and the final page to a design agent. The plan would make the page wait for the experiment’s results. This shows what a handoff means; it is not a Raven test result.

What the researchers measured

In the authors’ Multi-Agent Orchestration Benchmark, systems submitted plans without running the specialist agents. With the Qwen3.8-27B model, Raven’s reported Exact Match rate was 0.711, versus 0.607 for the strongest compared baseline. Exact Match counts scored plans whose specialist choices and required order of work agree with an accepted reference. The authors also report Raven ahead of both compared systems on all four planning measures with each of the two tested models. These are planning findings, not measures of finished project quality.

Why it matters

The paper separates choosing and ordering specialists from the harder question of whether their combined work succeeds. That distinction gives a working reader a way to examine an AI workflow’s handoffs, rather than treating a plausible-looking plan as a completed job.

Where it might help

A possible use is to sketch a multi-specialist job before running it: identify the required outputs, who could produce each one, and what evidence the next worker needs. The paper’s planning benchmark tests agreement with reference plans; it does not establish that such a plan will complete a real project.

Impact across sectors

  • Possible use in software teams: plan when investigation, code changes and a final explanation should occur.
  • Possible use in research support: map how source gathering, analysis and presentation depend on one another.
  • Possible use in creative production: specify which materials a designer needs from other contributors before starting.

Where the evidence stops

The authors say an alternative valid plan can differ from the benchmark’s reference plan. The planning test does not dispatch workers, and its nominal 140 requests do not by themselves determine every reported score’s denominator. Their theory gives sufficient conditions for reliable collaboration, but whether an implemented system meets them requires evaluation; it does not show that collaboration always beats one agent. Other evaluations have separate qualifications: Raven-Research’s comparison used questions also seen during development, and public benchmark copies were accessible; some Raven-Design slide configurations did not cover every task, and failed generations were excluded from reported means. General PtoP note: this arXiv source is a preprint, not evidence of peer review or a workplace deployment.

FROM PAPER TO PRACTICE

How to try it

A repository exploration, not a reproduction of the paper’s benchmark. Prerequisites: Git, Docker with Docker Compose, and a machine that can run containers; model or service access and any associated cost are not established by these README steps. Raven is described in its README as pre-alpha.

  1. Read the official repository README at https://github.com/EverMind-AI/Raven and decide whether to run its container locally.
  2. Run `git clone https://github.com/EverMind-AI/Raven.git`.
  3. Run `cd Raven`.
  4. Run `docker compose -f docker/docker-compose.yml up`, then open `http://localhost:18793` as the README directs. Observe the locally served page; do not treat its appearance as a reproduction of a paper result.

n8n example

n8n is not necessary here: the paper’s relevant test scores plans submitted through the systems’ own orchestration interfaces, not an n8n workflow.

Was this explanation easy to understand?

05TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS

Learning scene changes could help AI predict images and robot actions

Why it matters to youIf you work with video, image tools or robots, this paper offers a way to think about predicting what changes next. The authors tested that idea in research evaluations, not in a workplace deployment.

The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to represent what is happening, then tested whether that representation could help produce written answers, images and robot actions. Imagine seeing a hand beside a spoon, then seeing the spoon lifted. Orca’s video exercise asks it to predict an internal description of the next view, rather than draw every pixel. A second exercise gives it a description of an event and asks it to predict a view associated with that event. A third uses video questions and answers. The authors call these first two approaches “unconscious” and “conscious” learning; those names describe training methods, not awareness. After this training, the authors kept Orca’s main model fixed. They trained separate readouts—parts that turn its internal representation into an image or robot action—and used its existing language output for answers. This setup tests whether the learned representation is useful beyond a single output task.

Orca Team, Beijing Academy of Artificial Intelligence · Orca: The World is in Your Mind · arXiv:2606.30534 · 449 HF votes at selectionRead the paper ↗

In one sentence

The study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to represent what is happening, then tested whether that representation could help produce written answers, images and robot actions. Imagine seeing a hand beside a spoon, then seeing the spoon lifted. Orca’s video exercise asks it to predict an internal description of the next view, rather than draw every pixel. A second exercise gives it a description of an event and asks it to predict a view associated with that event. A third uses video questions and answers. The authors call these first two approaches “unconscious” and “conscious” learning; those names describe training methods, not awareness. After this training, the authors kept Orca’s main model fixed. They trained separate readouts—parts that turn its internal representation into an image or robot action—and used its existing language output for answers. This setup tests whether the learned representation is useful beyond a single output task.

Key concepts

  • State transition: a change from one situation to another, such as a spoon moving from the table into a hand.
  • World latent space: Orca’s internal representation of a scene and its possible changes, rather than a finished picture or sentence.
  • Observation-only learning: predicting the representation of the next video frame without an event description.
  • Event-conditioned learning: using a written event description to guide a prediction about an associated frame.
  • Frozen backbone and readouts: the authors stop updating Orca’s main model while testing parts that express its representation as images or actions.

A concrete example

Hypothetical illustration, not a paper result: a kitchen worker supplies a picture of a closed drawer and the instruction “open the drawer.” An image readout might attempt to show the drawer open. The useful question is whether it also keeps the surrounding counter and objects consistent; the paper does not test this workplace use.

What the researchers measured

In the authors’ real-robot tests, Orca-4B averaged 36.6 rule-based points across five tasks with unfamiliar tablecloth or background settings, versus 27.6 for π 0.5, the compared robot-control system pretrained on large-scale robot data. Rule-based points mark the highest task stage reached before a trial ends; they are not a success percentage. In the separate unfamiliar-object setting, Orca-4B averaged 28.2 points versus 31.2 for π 0.5. The authors trained the Orca action readout on 200 robot trajectories per task. In a qualitative spoon-grasp example, they report Orca recovering from early failures while π 0.5 made repeated failed attempts. For image prediction, the authors report a 59.8 average percentage score for Orca-4B with its image readout on PRICE-V0.1, versus 56.1 for FLUX.2 [klein]. That score comes from four model judges rating generated images against an initial image and instruction. The authors also report improved text, image and action readout results as Orca’s pre-training scales up.

Why it matters

The authors test whether learning about changes in video can be useful when the output is a sentence, an image or a robot action. That is different from training a separate main model solely to produce each kind of output. The results concern the authors’ stated evaluations, not general reliability in everyday settings.

Where it might help

Possible uses suggested by the evaluated tasks include studying video understanding, predicting how an instructed interaction might look, and researching robot manipulation. These are possibilities, not demonstrated products or deployments.

Impact across sectors

  • Possible use in media work: a video team could explore tools that answer questions about event order in clips; the paper tests benchmark questions, not an editing workflow.
  • Possible use in robotics research: a lab could investigate whether video-trained representations help a robot handle a changed tabletop scene; the paper’s robot tests cover five specified manipulation tasks.
  • Possible use in visual prototyping: a design team could explore images of an instructed scene change; the paper evaluates predicted interaction images, not a design service.

Where the evidence stops

The authors describe Orca as an early step. It learns mainly from vision and language, not sound, touch or force. Its visual prediction target comes from a frozen, pretrained vision component rather than a representation learned directly from all possible signals. The experiments mainly use the 0.8B and 4B sizes, and this version trains on only one-tenth of the authors’ 125K-hour video inventory. The authors report a trade-off among readout performances during training. They say their image benchmark has limited scale and variety, most annotated events cover short, minute-level changes, and the robot tasks remain relatively short and easy. The paper says no benchmark-specific training data was constructed or used, but does not establish that every underlying video source is disjoint from every evaluation source. General PtoP note: this arXiv paper is a preprint, and a benchmark result is not a deployment test.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation and public access to Orca are unverified from the supplied paper. Prerequisites: two pictures you may use and a short description of the change between them. No paid service is needed for this paper-and-pencil exercise.

  1. Put the pictures in time order and write down what stayed the same.
  2. Hide the second picture and describe what you expect to change next.
  3. Read the event description and revise your expectation, keeping unchanged objects in mind.
  4. Reveal the second picture and note which expectations were right or wrong. Observe how a description can guide a prediction; this does not reproduce Orca’s training or results.

n8n example

An n8n automation workflow is not appropriate here: the supplied paper does not verify a public Orca endpoint or an integration to connect to it.

Was this explanation easy to understand?

06TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS

A shared runbook may help AI agents finish long tasks

Why it matters to youIf you oversee an AI agent doing a long coding or operations job, you may want a way to see what it has finished and what it must check before stopping. This preprint studies a shared runbook for doing that; it does not test a workplace deployment.

The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does not change. The authors built StateM, a runtime that keeps the current phase of work and a shared runbook outside the agent’s long conversation history. Imagine an agent setting up a web service: it can work freely during setup, but the runbook can require a working end-to-end check before the agent moves to handoff. StateM records the phase, refreshes relevant instructions, checks configured conditions at transitions, and leaves a route back for repair. The authors developed one runbook with GPT-5.5, tested it unchanged with GPT-5.6 variants, adapted its practices for DeepSeek-V4-Flash, and developed separate runbooks for BusinessBench task families.

Ziheng Qin, Yaxin Lu, Zhangyang “Atlas” Wang, Kai Wang · StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling · arXiv:2608.15089 · 449 HF votes at selectionRead the paper ↗

In one sentence

The study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does not change. The authors built StateM, a runtime that keeps the current phase of work and a shared runbook outside the agent’s long conversation history. Imagine an agent setting up a web service: it can work freely during setup, but the runbook can require a working end-to-end check before the agent moves to handoff. StateM records the phase, refreshes relevant instructions, checks configured conditions at transitions, and leaves a route back for repair. The authors developed one runbook with GPT-5.5, tested it unchanged with GPT-5.6 variants, adapted its practices for DeepSeek-V4-Flash, and developed separate runbooks for BusinessBench task families.

Key concepts

  • A harness is the execution system around a model. The authors change that system, not the model’s trained weights.
  • A runbook describes phases, permitted transitions, checks and repair steps. Both the agent and a person can inspect it.
  • A state is the recorded current phase. Entering one brings its instructions forward; leaving one can require a check.
  • A configured check can block a handoff, but its strength depends on how it is verified. An agent’s own completion statement is not independent proof.
  • The authors turn selected lessons from failed runs into versioned practices. A practice developed for one model or task family may need changing elsewhere.

A concrete example

Hypothetical illustration, not a reported test: an agent updates an internal service. Its runbook keeps it in a verification phase until a configured check confirms that another system can use the service. If the check fails, the agent returns to repair rather than declaring the job done.

What the researchers measured

On Terminal-Bench 2.1, an 89-task terminal benchmark, the authors report 92.1% trial-level success for GPT-5.5 xhigh with StateM, against an 83.1% published GPT-5.5 reference. Trial-level success counts completed trials, with five trials per task. With the runbook frozen, GPT-5.6 Sol xhigh with StateM recorded 424 successful trials out of 445, or 95.28% raw accuracy, in a public submission. That score is before review is settled, not a finalized leaderboard result. For a different model and setting, the adapted DeepSeek-V4-Flash system scored 392 successes out of 445 under standard timeouts; the authors report $15.20 in realized API charges for its final-score evidence, excluding $37.02 spent on adaptation. On BusinessBench, the first frozen held-out evaluation with Codex and GPT-5.6 Luna showed smaller aggregate gains: 0.55 percentage points when task families were weighted equally and 1.34 points when instances were weighted equally.

Why it matters

The authors’ results separate a model’s ability to perform individual steps from the execution system’s ability to carry a job through to completion. They also show a boundary: an unchanged runbook developed with GPT models did not improve their DeepSeek-V4-Flash result, while an adapted one did.

Where it might help

Possible applications include tracking obligations during long coding jobs and placing checks before consequential operations handoffs. These are potential uses of the control approach, not deployments demonstrated by the paper.

Impact across sectors

  • Software teams: potentially use shared runbooks to keep visible requirements, tests and handoff checks together during a long coding task; hypothetical workplace use.
  • IT operations: potentially require a service-readiness check before handing over a configuration change; hypothetical workplace use.
  • Business operations: potentially preserve approval conditions across a multistep workflow; hypothetical workplace use, not a finding that StateM works for all such workflows.

Where the evidence stops

The Terminal-Bench gains measure StateM’s runtime together with an evolved, benchmark-adapted runbook; they do not isolate the runtime alone. The GPT-5.5 comparison uses a published reference, not a newly rerun control matched for agent version. The GPT-5.6 Sol xhigh public submission remained open and unmerged at the stated review date. The authors say four rewarded trials should not count and that nine trajectories were flagged for possible reward hacking; the raw 95.28% must therefore remain qualified. Solving a task at least once in five trials is not single-run reliability. The unchanged GPT-developed runbook reduced DeepSeek-V4-Flash’s full-suite score from 82.7% to 82.0%; its later improvement required adaptation. BusinessBench has one trajectory per instance per arm, excludes the untreated Attendance family from intervention aggregates, and shows declines in two treated families in its first frozen evaluation. Later refinements reused evaluation feedback and are not untouched held-out results. StateM cannot recover unrecorded work, guarantee that a self-reported check is correct, or undo arbitrary external actions. PtoP note: this source is an arXiv preprint, not a claim of peer review; benchmark scores are not evidence of workplace deployment.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation is unverified because the supplied text names a code release but provides no verified official repository link or installation commands. Prerequisites: a multistep task you understand and a place to write a checklist. No model access or API spending is needed for this paper-based exercise.

  1. Write down the task’s main phases.
  2. For each phase, note what evidence would justify moving on.
  3. Mark which checks a person or tool could verify and which would be only a self-report.
  4. Add a repair step for a failed check. Observe whether the written procedure makes an unfinished handoff easier to spot; this is not a test of StateM.

n8n example

A proposed n8n integration could pause a workflow before handoff and request a human approval when a required check fails. This is an illustration, not a tested StateM feature or a verified integration procedure.

Was this explanation easy to understand?

COMPARISON

What these studies show together

Compare topics, possible uses and limits. Each result remains attributed to its own paper.

PaperCentral ideaPossible useLimit
Long-document AI design depends on what models keep and retrieveA model handling a long contract must decide what to keep from earlier pages and what to consult when answering a question. The authors explain that designs make…The survey’s framework could help people planning long-document AI systems ask what evidence remains available, how retrieval candidates…The authors say their literature coverage and release inventory are curated, not exhaustive. One of the 59 release records lacks separately…
Test whether an assistant delivers the right scientific resultA computer-using assistant can carry out many steps in scientific software yet still submit a wrong result. In one run described by the authors, GPT-5.6-terra entered an…Possible use: a research team could use the paper's task descriptions and outcome checks to decide what an assistant must verify before…The authors report uneven numbers of tasks across fields. Their detailed trajectory analysis covers released runs for 128 tasks and…
Keep Search Agents From Rewarding Their Own Wrong AnswersA search agent can appear to improve by learning to repeat its own wrong answers. The authors study a training loop in which one part writes questions and draft answers…A possible use is to design the feedback stage of a system that generates its own search-and-answer practice. The paper tests training and…The authors used an automated, evidence-backed audit after training, not exhaustive human annotation. Unsupported cases remained…
Plan specialist AI work before handing tasks to agentsWhen a job needs several kinds of work, the authors ask whether one system can choose suitable specialists and arrange their handoffs. Raven is their proposed system…A possible use is to sketch a multi-specialist job before running it: identify the required outputs, who could produce each one, and what…The authors say an alternative valid plan can differ from the benchmark’s reference plan. The planning test does not dispatch workers, and…
Learning scene changes could help AI predict images and robot actionsThe study asks whether one learned picture of a changing scene can support several kinds of prediction. The authors trained Orca to represent what is happening, then…Possible uses suggested by the evaluated tasks include studying video understanding, predicting how an instructed interaction might look…The authors describe Orca as an early step. It learns mainly from vision and language, not sound, touch or force. Its visual prediction…
A shared runbook may help AI agents finish long tasksThe study asks whether an AI agent can finish more long tasks when its working procedure improves, even if the underlying model does not change. The authors built…Possible applications include tracking obligations during long coding jobs and placing checks before consequential operations handoffs…The Terminal-Bench gains measure StateM’s runtime together with an evolved, benchmark-adapted runbook; they do not isolate the runtime…

ARCHIVE

Issue archive

Find earlier editions of the newspaper.

01
Model Memory, Scene Prediction, Scientific Software and Agent Workflows01-10-2026
↗
30
Chinese Answers, Agent Memory, Engineering Tests, Open Models and Simulators30-09-2026
↗
29
AI studies test reasoning, weight readouts, page search, tables and live video29-09-2026
↗
27
Two Texts, a Tutor and Agent Judgment27-09-2026
↗
26
Models and the World26-09-2026
↗
Search every paper by keyword or tag ↗