ISSUE 06/2026 · 02-10-2026

Test whether AI agents finish professional computer work

The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain experts. In each task, an agent receives instructions and works in an environment with the relevant…

Editorial illustration: An anonymous worker lays a finished stack of papers beside an open toolbox while several partly assembled objects remain scattered across a desk.
AI-generated editorial illustration: a visual metaphor, not a figure from the paper.
Read the lead story ↘

6 paper · Original sources linked in every story

5 MINUTES ON PtoP

What Helps an AI Finish the Job

An AI assistant can be useful without being ready to finish a job alone. Anthropic says more than 16,000 Barclays colleagues have adopted a Claude-powered assistant that helps staff find information for customer questions. That is a picture of AI woven into someone’s workday. The authors of Agents’ Last Exam ask a different question: can an agent leave behind a completed file, report or design? In their computer-work benchmark, full completion varied sharply by task tier. Finding an answer and finishing a deliverable are worth treating as different tests. See: Barclays scales Claude to upgrade operations and improve… · Test whether AI agents finish professional computer work

Several items today suggest that what surrounds an agent matters as much as the request it receives. The ActiveSaddler authors tested a way to choose practice tasks as an agent’s weaknesses changed; on two benchmarks, they report higher held-out scores than when the task order was fixed in advance. ServiceNow says its AutoSynthData system similarly creates and checks practice tasks based on a workplace helper’s mistakes, with improved performance in its own controlled tests. Neither result tells us how much improvement a particular workplace would see. See: Choosing which agent failures to revisit can improve… · AutoSynthData: Generating Training Data for Enterprise…

Practice is only one part of that surrounding structure. The DisCo authors gave a research agent reusable instructions drawn from software repositories and papers: how to set things up, check results and recover from mistakes. With task-oriented guidance, they report higher aggregate scores on research benchmarks than for the same agent without it, while noting that guidance lowered scores on two PaperBench tasks. Instructions could help a team avoid repeatedly explaining a workflow, but choosing the wrong ones might get in the way. See: Reusable skills may help research agents follow software…

Then comes the question of who decides what counts as done. In the development project described by the Atria Dawn Preview authors, agents often proposed methods and revisions, while people usually made the final choices. The Agents’ Last Exam authors, meanwhile, check finished work against reference outputs or task-specific rules rather than accepting a plausible account of the work. Together, these papers offer a useful distinction: an agent might carry out steps, but people still need to choose the aim and decide what evidence would make the result trustworthy. See: AI agents can propose research methods while people choose… · Test whether AI agents finish professional computer work

If that distinction holds in your own work, the first step may be smaller than choosing a new tool. This week, could you pick one task you might give an AI assistant and write down the finished thing you would expect to see—and the check you would use before relying on it? See: Test whether AI agents finish professional computer work

Just here for the stories? They are below, by topic.

PtoP · NEWSLETTER

Get the next issue in your inbox

Six papers and the news, explained in plain English, each morning after the editor approves the issue. Free, and you can unsubscribe at any time.

From the news desk

News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.

  1. Tools and building Hugging Face Blog · 02-10-2026AutoSynthData: Generating Training Data for Enterprise Agents ↗Read the explainer ↓
  2. Business and strategy Anthropic · 01-10-2026Barclays scales Claude to upgrade operations and improve client experience ↗Read the explainer ↓
  3. AI agents Google · 30-09-2026Gemini 4 Argon: our next era of frontier intelligence ↗Read the explainer ↓
  4. Tools and building xAI · 18-09-2026Introducing Grok Voice Transcribe 2.0 ↗Read the explainer ↓

PODCAST / ENGLISH EDITION

The full issue, then each paper

The front-page episode covers the papers and news; each paper also has its own short conversation.

Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.

EP 01ARXIV:2610.00906

Choosing which agent failures to revisit can improve testing results

The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order…

EP 02ARXIV:2610.02196

Humanoid robots can reuse learned movements through revised task instructions

The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than…

EP 03ARXIV:2610.02201

Overlapping slices could help preserve holes in generated 3D objects

The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a…

EP 04ARXIV:2609.15818

AI agents can propose research methods while people choose directions

An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the…

EP 05ARXIV:2609.02749

Reusable skills may help research agents follow software instructions

The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its…

EP 06ARXIV:2606.05405

Test whether AI agents finish professional computer work

The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE…

FROM THE NEWS DESK

The news, explained

Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.

Tools and building Hugging Face Blog · 02-10-2026 · company announcement

AutoSynthData: Generating Training Data for Enterprise Agents

In brief

ServiceNow says it built a way to create practice tasks for workplace AI helpers based on what they get wrong. Imagine a helper learning to update a support ticket while following company rules, then practicing that skill on a different ticket. The system, AutoSynthData, makes and checks new tasks before using them to train the helper.

An example

As an illustration, if a helper wrongly closes a ticket before an approval arrives, a practice task could ask it to handle a different ticket that also needs approval. A check would reject a result that skips that step.

Application

A company could use this approach to train an AI helper on difficult tasks in its own support workflow.

The limitation

These are results from ServiceNow's own tests in EnterpriseOps Gym, a controlled workplace-task test environment. They do not establish that the approach works equally well in a company's live systems.

Takeaway

Practice tasks may be more useful when they target a helper's current mistakes and are checked for valid solutions. ServiceNow reports improved test performance after training on such tasks.

Original source ↗

Was this explanation easy to understand?

Business and strategy Anthropic · 01-10-2026 · company announcement

Barclays scales Claude to upgrade operations and improve client experience

In brief

Anthropic says Barclays is expanding its use of Claude to help staff answer customer questions, handle emails and develop software. For example, a bank worker can use an existing assistant to find information for a customer. Anthropic says more than 16,000 Barclays colleagues have adopted that assistant.

An example

If a customer asks about a banking service, an employee can use the Colleague Knowledge Assistant to find relevant information before replying. This illustrates the use Anthropic describes; it does not show how much time any particular customer saves.

Application

Anthropic reports that Barclays uses Claude to help route incoming emails in its Global Markets business, so staff can identify requests that need action.

The limitation

These are claims in an Anthropic announcement, not independently measured results presented here. The stated goal for Claude Code to reach 50% of Barclays developers by the end of 2026 is a forecast, not an achieved outcome.

Takeaway

Barclays already uses a Claude-powered knowledge assistant and is extending the technology to other work, but the promised gains and future reach remain uncertain.

Original source ↗

Was this explanation easy to understand?

AI agents Google · 30-09-2026 · company announcement

Gemini 4 Argon: our next era of frontier intelligence

In brief

Google announced Gemini 4 Argon, an AI model it says can carry out long coding and security tasks with less step-by-step help. Google is initially sharing it with selected cyber defenders rather than releasing it broadly.

An example

For example, a hospital software team could ask an AI system to look for a vulnerability, suggest a fix, and let people review that fix before using it. This illustrates the kind of task Google describes, not a verified use of Argon.

Application

Selected security teams could use Argon to help find and repair weaknesses in software.

The limitation

The performance claims come from Google’s announcement. Google says it is still testing guardrails and gathering feedback before wider access.

Takeaway

Argon is a limited rollout of an AI model designed to take on longer, more independent work; its broader usefulness and safety remain to be tested.

Original source ↗

Was this explanation easy to understand?

Tools and building xAI · 18-09-2026 · company announcement

Introducing Grok Voice Transcribe 2.0

In brief

xAI says it has released Grok Voice Transcribe 2.0, a tool that turns speech into written words. The company says it handles noisy calls and multiple languages better than its earlier version, at the same price.

An example

Imagine a customer-support call with a poor connection. The tool could write down what was said and label which person spoke.

Application

A support team could use the resulting transcription to review calls without listening to each recording.

The limitation

The broad accuracy claims rely on xAI's own evaluations, though it also cites a public ranking. Performance on a particular call may differ.

Takeaway

This is a company-reported improvement aimed at making everyday audio easier to turn into usable text.

Original source ↗

Was this explanation easy to understand?

01NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS

Choosing which agent failures to revisit can improve testing results

Why it matters to youIf you help improve an AI assistant at work, you may need to decide which failed practice tasks deserve another attempt. This paper suggests that changing that choice as the assistant changes can improve results in benchmark tests.

The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order beforehand. The authors built ActiveSaddler to make that choice during offline optimization, before any deployment. A harness is the set of prompts, available tools, and operating rules around an AI model. ActiveSaddler does not replace the method that repairs this harness; it chooses which training tasks provide evidence for the next repair. It groups related failures, estimates which known weakness may be worth another attempt, and sometimes tries an unseen task to look for a new one. The authors tested this approach with the AutoSaddler harness optimizer on GAIA2 and Terminal-Bench 2.0.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, and Victor Rühle · ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization · arXiv:2610.00906Read the paper ↗

In one sentence

The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order beforehand. The authors built ActiveSaddler to make that choice during offline optimization, before any deployment. A harness is the set of prompts, available tools, and operating rules around an AI model. ActiveSaddler does not replace the method that repairs this harness; it chooses which training tasks provide evidence for the next repair. It groups related failures, estimates which known weakness may be worth another attempt, and sometimes tries an unseen task to look for a new one. The authors tested this approach with the AutoSaddler harness optimizer on GAIA2 and Terminal-Bench 2.0.

Key concepts

  • A training scenario is a practice task run by the current agent. Its execution record shows what happened and helps the optimizer propose a harness change.
  • A failure-pattern arm is a group of observed failures attributed to a shared harness weakness. It lets the curriculum revisit a weakness that appears in more than one task.
  • Arm prioritization estimates whether another attempt at a known weakness might help. The estimate considers whether the weakness persists, seems fixable, could affect other tasks, or risks disrupting working behavior.
  • Exploration means running a training scenario not yet tried during optimization. Revisiting means spending the next attempt on scenarios associated with a known failure pattern.
  • A rollout is one task-agent execution. The authors compare ways of allocating a limited rollout budget while keeping the underlying harness optimizer fixed.

A concrete example

Hypothetical illustration, not a reported result: An assistant repeatedly omits attachments from different calendar tasks. Instead of treating each failed task as unrelated, a curriculum could group them as one possible weakness, revisit it after a proposed repair, and then try an unseen calendar task to check for other failures.

What the researchers measured

The authors report mean test Pass@1 over three test-time executions; this score is the share of held-out tasks passed on a single attempt. With ActiveSaddler added to AutoSaddler, they report 59.8% on the GAIA2 test split and 80.0% on the Terminal-Bench 2.0 test split. These are, respectively, 4.4 and 7.5 percentage points above the same optimizer using a randomly shuffled scenario order fixed before optimization. In separate component-removal tests, the authors report lower test scores when failure-pattern grouping, arm prioritization, or adaptive exploration is replaced. Those tests concern the named variants, not a different deployed system.

Why it matters

A fixed practice schedule can revisit a weakness that has already been addressed or move past one that remains. The authors' results suggest that the choice of training tasks matters alongside the method used to change the harness.

Where it might help

A possible use is planning which practice tasks should guide improvements to an agent's prompts, tools, or operating rules. The paper tests that idea in benchmark environments, not in a workplace deployment.

Impact across sectors

  • Possible, not tested: Teams building workplace assistants could use recurring task failures to decide what to practice next.
  • Possible, not tested: Software-development teams could consider this approach when improving agents that work through terminal tasks.
  • Possible, not tested: Teams maintaining simulated mobile assistants could use changing failure patterns to organize pre-deployment testing.

Where the evidence stops

These are controlled benchmark results, not a production evaluation. The authors used public and synthetic task data, not real user data, and say the experiments do not establish safety, security, privacy, or governance readiness. GAIA2 training, development, and test tasks came from disjoint simulated environments; Terminal-Bench 2.0 used a uniform random partition because it lacks a natural grouping axis. The reported test scores average three test-time executions. ActiveSaddler also adds optimizer-side computation. The paper says a project website and code will be available, but the supplied text does not verify an available installation. General PtoP note: this arXiv version is a preprint; the supplied text does not establish peer review.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation is unverified, and the paper's linked project website and code are described as forthcoming. Prerequisites: a few fictional practice tasks and a way to record outcomes. No account, payment, model access, or real execution traces are needed for this exercise.

  1. Write several fictional tasks for an assistant and imagine which ones fail.
  2. Group failures that appear to share one underlying weakness; keep distinct weaknesses separate.
  3. Choose whether to revisit one group or inspect an unseen fictional task, and write down why.
  4. Imagine a repair, update the failure groups, and repeat the choice. Observe whether your priorities change when a weakness appears resolved or a new one emerges. This is a thought exercise, not a test of ActiveSaddler.

n8n example

An n8n workflow would not reproduce this study: the paper does not specify a verified n8n integration, and its method depends on harness updates, execution records, and repeated evaluation.

Was this explanation easy to understand?

02NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS

Humanoid robots can reuse learned movements through revised task instructions

Why it matters to youIf you plan tasks for a warehouse or research robot, this study suggests a possible way to try new box-handling tasks without retraining its movement controller. The authors tested the approach mainly in simulation, not as a warehouse deployment.

The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than retraining how it moves. The authors’ answer is to keep its movement controller fixed and revise a reward program: a set of goals for successive parts of a task. Imagine moving a box onto a table. One part can favor lifting and carrying it; a later part can favor letting go. A completion condition tells the system when to switch parts. InterEvolve has a language-model agent revise those parts after simulated attempts, while a numerical search adjusts settings such as how strongly each goal counts. A separate, fixed verifier checks whether the attempt met the task’s requirements. Successful programs can be kept in a skill library for later tasks.

Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui · InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation · arXiv:2610.02196Read the paper ↗

In one sentence

The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than retraining how it moves. The authors’ answer is to keep its movement controller fixed and revise a reward program: a set of goals for successive parts of a task. Imagine moving a box onto a table. One part can favor lifting and carrying it; a later part can favor letting go. A completion condition tells the system when to switch parts. InterEvolve has a language-model agent revise those parts after simulated attempts, while a numerical search adjusts settings such as how strongly each goal counts. A separate, fixed verifier checks whether the attempt met the task’s requirements. Successful programs can be kept in a skill library for later tasks.

Key concepts

  • A reward program describes the desired interaction in stages; it does not prescribe every joint movement.
  • The fixed movement controller was trained to use information about both its body and an object. A new reward can steer its existing movements without retraining that controller.
  • The agent changes the program’s stages and goals, while a numerical search tunes their settings. Simulated attempts provide feedback for both.
  • A fixed verifier checks outcomes and constraints independently of the reward program being revised.
  • A skill library stores selected programs and their results so later tasks can draw on them.

A concrete example

Hypothetical illustration, not a reported trial: for a request to place a parcel on a shelf, one stage could favor getting hold of it and moving it upward. A second could favor releasing it once it reaches the shelf. A verifier would separately check whether the parcel stayed there and the robot let go.

What the researchers measured

On the authors’ eight simulated box-task families, InterEvolve’s evolved programs achieved 86.5% success, compared with 34.6% for the agent’s initial program after numerical tuning. Success means every task criterion and physical constraint held in the same attempt. The authors also show a physical Unitree G1 executing simulation-evolved kicking and pushing tasks using onboard sensing and an off-board workstation; they do not give a physical-task success rate in the supplied text.

Why it matters

The authors separate task strategy from movement training. In their experiments, changing the goals and the order in which they apply let the same controller use movements that a fixed program did not reliably bring out.

Where it might help

A possible use is preparing new task instructions for a robot that already has relevant movements, then checking candidate instructions in simulation before physical execution. The paper does not establish a general-purpose deployment process.

Impact across sectors

  • Possible logistics use: explore staged instructions for moving and placing boxes with an already-trained humanoid; this was not tested in a warehouse.
  • Possible manufacturing use: study whether stored handling programs help with new multi-step object tasks; this was not tested on a production line.
  • Possible robotics-research use: compare task instructions while keeping one movement controller fixed; this is an application of the reported experimental approach.

Where the evidence stops

The authors state that search cannot bring out behavior the controller never learned and is limited by its stored examples of body-and-object states and by available measurements. Simulation and language-model use take GPU time and add latency, ruling out real-time replanning during physical execution. The main tracking test uses held-out clips from objects represented in training; the additional box types tested for transfer were also seen during controller pretraining, though not during program search. The physical demonstration covers two tasks, with no reported success rate. In that setup, prompts prepared from a simulated run are replayed on the robot while its fixed controller responds to sensed state; the paper does not report physical test-time program evolution. PtoP note: this arXiv version is a preprint, not evidence of peer review.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation is unverified. Prerequisites: paper or pencil and a familiar multi-step object-moving task. No software access or cost is required for this exercise; reproducing the study would require resources not established by verified installation instructions here.

  1. Write the task’s final outcome and a separate safety or contact constraint.
  2. Divide the task into stages, such as approach, move and release.
  3. For each stage, write what progress would look like and what observation would trigger the next stage.
  4. Imagine an attempt that fails one constraint, and revise only the relevant stage. Observe how changing the task description differs from changing the robot’s movements directly. This is not a tested procedure for operating a robot.

n8n example

An n8n workflow is not appropriate here: the paper does not describe a verified automation interface for its robot, simulator or program search.

Was this explanation easy to understand?

03NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS

Overlapping slices could help preserve holes in generated 3D objects

Why it matters to youIf you make 3D assets from reference images, this study offers a way to think about a familiar problem: a shape can look close to the image but lose a hole or a thin support. The authors test a method designed to keep those structures intact.

The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a shape through overlapping cross-sections viewed from three directions, rather than many tiny 3D cells. Imagine slicing a bicycle wheel from front to back: successive slices reveal where the open center and spokes begin and end. SILSA compresses such slices into tokens, small packets of shape information. A reconstruction model learns to turn those packets back into a 3D surface. A separate image-guided generator learns to produce the packets from a single picture. The authors also train the reconstruction model to notice connected parts and holes, while a shared 3D workspace helps slices from different directions describe the same object.

Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen · SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation · arXiv:2610.02201Read the paper ↗

In one sentence

The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a shape through overlapping cross-sections viewed from three directions, rather than many tiny 3D cells. Imagine slicing a bicycle wheel from front to back: successive slices reveal where the open center and spokes begin and end. SILSA compresses such slices into tokens, small packets of shape information. A reconstruction model learns to turn those packets back into a 3D surface. A separate image-guided generator learns to produce the packets from a single picture. The authors also train the reconstruction model to notice connected parts and holes, while a shared 3D workspace helps slices from different directions describe the same object.

Key concepts

  • Overlapping slice tokens: Each token summarizes a thin region rather than a single flat cut. Neighboring regions overlap so the model can follow a structure through depth.
  • Three-direction views: Slices taken along three perpendicular directions expose different parts of the same shape.
  • Topology: This means structural features such as separate parts, openings and connections. A filled wheel opening is a structural error even if much of the surface looks right.
  • Slice-level supervision: During reconstruction training, the method compares holes and connected parts within slices and checks where they change between neighboring slices.
  • Shared 3D workspace: During image-guided generation, slices from the different directions exchange information through a common spatial memory.

A concrete example

Hypothetical illustration, not a reported test: A designer supplies one image of a bicycle wheel. A useful 3D result would keep its center open and its thin spokes connected to the rim, rather than filling the opening or breaking a spoke.

What the researchers measured

The authors trained SILSA on Trellis-500K and evaluated using 200 randomly sampled Toys4K assets and 50 in-the-wild images, which they say did not overlap with training. In their image-to-3D comparison, SILSA had leading scores on the reported FD, PSNR, coverage and MMD measures; it tied the best KD and LPIPS scores. These measures assess aspects of generated-shape quality, image fidelity or how well outputs represent the reference set. Its PSNR, an image-fidelity score where higher is better, was 32.74 versus 30.12 for SparseFlex. SparseFlex had a slightly higher CLIP score, the paper’s input-image alignment measure. Separately, in the SliceVAE reconstruction comparison, the authors report a lower Betti error—a measure of mismatched connected parts and holes—than the reconstruction baselines. SILSA used 384 fixed slice tokens. In the reported efficiency comparison, its training memory and end-to-end inference time were lower than Dora’s under the stated test settings.

Why it matters

The distinction between a close-looking surface and a correctly connected object matters when a missing opening or broken support changes the shape itself. The authors report both structural measures and generation-cost measures, so readers can see what happened in their experiments rather than infer structure from appearance alone.

Where it might help

Possible uses, not tested deployments, include making draft 3D assets from images for creative or instructional work. The paper reports generation and reconstruction experiments, not an end-to-end workplace workflow.

Impact across sectors

  • Possibility for content creation: An artist could explore image-based drafts of objects with handles, spokes or other delicate parts; this is not a demonstrated production use.
  • Possibility for design: A designer could use generated shapes as starting points for inspection and editing; the paper does not test design decisions.
  • Possibility for education: Instructors could use image-to-3D examples to discuss how openings and connections affect a shape; this is a hypothetical use.

Where the evidence stops

The authors say SILSA depends on the diversity and scale of its training data: shapes with unfamiliar structures may be reconstructed less faithfully. Ambiguous or low-information input images may also reduce generation quality. Their open-surface garment evaluation notes that topology measures are near zero for methods that recover the rough surface, limiting what that measure distinguishes there. The authors also identify risks of unauthorized replicas and displacement of manual modeling work. General PtoP note: this arXiv source is a preprint; the reported benchmarks are not a tested deployment.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation and model access are unverified. Prerequisites: paper and pencil, plus an everyday object with an opening, such as a mug. No software cost is needed for this exercise; access or cost for running SILSA is not established here.

  1. Sketch the object from the front, side and top.
  2. Draw a few neighboring cross-sections for each view, letting adjacent sections overlap.
  3. Mark where the handle’s opening appears or disappears as you move through the sections.
  4. Compare the three sets of sketches: observe how one view can miss a connection that another makes visible. This illustrates the paper’s representation; it does not reproduce its results.

n8n example

An n8n workflow is not appropriate here: the supplied paper does not verify a runnable SILSA interface for a proposed integration.

Was this explanation easy to understand?

04TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS

AI agents can propose research methods while people choose directions

Why it matters to youIf you lead research or engineering work, this study may help you think about which parts of a project to give an AI agent—and which decisions to keep with people. It describes one model-development project, not a tested way to run every team.

An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the authors describe in their development of Atria Dawn Preview, a model designed to use tools across research and engineering work. They trained it on tasks in working environments, keeping records of its steps and checking outcomes against evidence such as tests or file state. They then compared its performance on set tasks and examined task records and agent logs from the people developing it. In those records, agents often proposed methods and made revisions; people usually made the final choices.

Atria Team · Atria Dawn: The Dawn of Agentic Superintelligence — On the Evolving Roles of Human–AI Collaboration · arXiv:2609.15818 · 431 HF votes at selectionRead the paper ↗

In one sentence

An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the authors describe in their development of Atria Dawn Preview, a model designed to use tools across research and engineering work. They trained it on tasks in working environments, keeping records of its steps and checking outcomes against evidence such as tests or file state. They then compared its performance on set tasks and examined task records and agent logs from the people developing it. In those records, agents often proposed methods and made revisions; people usually made the final choices.

Key concepts

  • An agent is a model that can take a sequence of actions, inspect what happened and decide what to do next, rather than give only one reply.
  • The Verifiable Experience Pipeline is the authors’ training process for linking a task, the agent’s actions, its output and an external check. For example, a program can be run and tested rather than judged only by how plausible its code looks.
  • A benchmark is a fixed set of evaluation tasks. The paper uses benchmarks to compare models under specified conditions, not to measure success in every workplace.
  • Proposal and selection are different decisions: an agent may suggest a method, while a person decides whether to use it.

A concrete example

Hypothetical example, not a paper result: an engineering lead asks an agent to investigate a failing data-processing job. The agent checks logs, suggests a change and runs a test. The lead decides whether that test answers the right question and whether the change should be kept.

What the researchers measured

The authors report that Atria Dawn Preview had the highest reported score on five of 16 benchmarks; rankings use the available entries, and some comparison results were unreported. In their project review, participants rated 151 of 455 completed AI-assisted tasks with usable responses as infeasible without AI under the same scope and resources. Among participants in execution roles, humans made the final choice in 85.5% of recorded method-or-parameter decisions. That figure concerns who selected an option, including options proposed by AI; it is not a measure of task success.

Why it matters

The authors’ distinction is between doing more work and deciding what work is worth doing. Their project records show substantial agent involvement alongside continued human selection and intervention. They say completing research tasks does not, by itself, show that an AI system can repeatedly discover worthwhile improvements to future models.

Where it might help

As a possible use, a research team could ask an agent to run bounded experiments and bring the results to a human decision-maker. The paper does not establish a general deployment procedure or show that an agent can choose valuable research directions on its own.

Impact across sectors

  • Possibility, not a proven deployment — research labs: agents could prepare experiments and report what happened while researchers choose which questions merit further work.
  • Possibility, not a proven deployment — software teams: agents could implement and test proposed changes while engineers review the evidence and decide what to accept.
  • Possibility, not a proven deployment — security teams: agents could assist with diagnosis and re-testing in authorized environments, with people retaining control of scope and decisions.

Where the evidence stops

This is an arXiv preprint, not a claim of peer review. The collaboration findings come from the authors’ own development project, and the infeasibility finding is based on participants’ ratings of their tasks. The paper calls its tool-use cases selected demonstrations rather than estimates of average success. Evaluation coverage also varies: SkillsBench uses a self-contained subset that excludes multimodal tasks, and one optimization case had no complete formal mean at the end of its published trace. The authors say the weather interface does not report forecast accuracy and the displayed computer-aided designs are not mechanical or manufacturing validation.

FROM PAPER TO PRACTICE

How to try it

The official repository links to Atria Dawn Preview’s model page and hosted access, but its README gives no complete local installation commands. The following is an untested, text-only exploration; pricing and the computing needed for local use are not specified there.

  1. Prerequisites: have a small, non-sensitive research or coding question and a way to check the answer yourself. For hosted use, check access and key requirements at https://api.atria-asi.ai/ and https://api.atria-asi.ai/docs.
  2. Confirm that you are looking at Atria Dawn Preview, rather than its separately listed FP8 variant, on the linked model page: https://huggingface.co/internlm/Atria-Dawn-Preview.
  3. If you obtain hosted access, follow the service documentation to submit your question as text only; do not send images or files.
  4. Ask for a proposed method and the evidence that would test it. Check any output against your own source material or a safe test, and observe whether the proposal holds up before choosing a next step. This is an exercise, not a reproduction of the paper’s evaluation.

n8n example

An n8n automation is not appropriate for this first exercise: the point is to inspect the agent’s evidence and make a deliberate human choice, rather than automate that choice.

Was this explanation easy to understand?

05TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS

Reusable skills may help research agents follow software instructions

Why it matters to youIf you use a coding agent for machine-learning experiments, this study offers a way to give it relevant setup steps and checks before it starts. The authors tested that approach on research benchmarks, not in an everyday workplace deployment.

The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its general knowledge alone. Imagine asking an agent to use an unfamiliar software package: it needs to know not just what the package does, but how to set it up, check its output and recover from mistakes. The authors call this practical know-how operational knowledge. Their DisCo system turns material from repositories and papers into reusable instruction bundles called skills. An entry file says when a skill applies; supporting files can hold detail and executable helpers. A router helps the agent open only the relevant bundle. The authors compared a Codex research agent with and without these skills while keeping its GPT-5.5 model, research harness—the software coordinating its work—and downstream running budget fixed.

Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu · Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills · arXiv:2609.02749 · 396 HF votes at selectionRead the paper ↗

In one sentence

The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its general knowledge alone. Imagine asking an agent to use an unfamiliar software package: it needs to know not just what the package does, but how to set it up, check its output and recover from mistakes. The authors call this practical know-how operational knowledge. Their DisCo system turns material from repositories and papers into reusable instruction bundles called skills. An entry file says when a skill applies; supporting files can hold detail and executable helpers. A router helps the agent open only the relevant bundle. The authors compared a Codex research agent with and without these skills while keeping its GPT-5.5 model, research harness—the software coordinating its work—and downstream running budget fixed.

Key concepts

  • Operational knowledge is practical guidance about which tools fit a task, how to use them and what can go wrong.
  • A skill packages that guidance around a SKILL.md entry file, with deeper references and scripts where needed.
  • Skill distillation turns source material into procedures through scoping, evidence gathering, construction and verification.
  • Task-agnostic skills are prepared from sources for later reuse; task-oriented skills are built around a particular problem or benchmark.
  • Progressive disclosure means the agent loads the relevant branch of skills instead of reading the entire library.

A concrete example

Hypothetical illustration, not a study result: a researcher asks an agent to compare two unfamiliar model-serving packages. A relevant skill might tell it to use the same workload for both, retain the commands it ran and check the measurements before reporting them.

What the researchers measured

The authors report that, on the full MLE-bench suite of 75 machine-learning competitions, GPT-5.5 Codex with task-oriented AREX-Skill guidance received an Any-Medal score of 72.89%, versus 31.11% without skills. That score reports how often the agent attained any medal. On PaperBench's 20 paper-reproduction tasks, they report an average replication score of 39.59% with the selected paper-derived skills, versus 29.45% without them; that score is assigned by the benchmark's reproduction grader. The paper also reports higher aggregate scores with skills on FrontierCS and PassNet. Those evaluations used their own task-oriented graphs, not simply the public repository collection.

Why it matters

The authors' comparison separates access to distilled instructions from changes to the agent's model or research harness. It therefore speaks to what practical guidance may add under their benchmark settings, while the separate effort of creating that guidance still matters.

Where it might help

A possible use is to give compatible coding agents inspectable package procedures and validation checks before a research task. The paper measures benchmark work; it does not establish that this approach succeeds in a particular workplace.

Impact across sectors

  • Possible impact in machine-learning engineering: teams could provide agents with package-specific setup and checking guidance. This is not a reported deployment.
  • Possible impact in academic research: agents attempting paper reproductions could consult reusable instructions drawn from related work. This is a possibility, not a demonstrated outcome outside the benchmark.
  • Possible impact in software performance work: agents could use guidance for checking that an optimization still produces correct outputs. This is not a reported production result.

Where the evidence stops

The authors report that skills lowered the PaperBench scores on two tasks. They suggest that retrieved material may sometimes distract the agent from a better task-specific approach, but present this as a possible explanation, not an established cause. Skill construction happened before benchmark execution and used a separate budget; matched downstream running budgets do not include that construction cost. The public repository snapshot and the skill collections used in the evaluations are distinct. For PaperBench, the authors excluded each target paper and its released artifacts as skill sources, while selecting related prior work. The repository collection is described as curated rather than exhaustive. General PtoP note: this is an arXiv preprint, not a peer-reviewed finding, and benchmark scores are not deployment results.

FROM PAPER TO PRACTICE

How to try it

Advanced coding exercise using the official repository's README; these steps are not a tested procedure here. No-execution alternative included.

  1. Prerequisites: for the execution route, use macOS, Linux, WSL or Git Bash and a Node.js runtime meeting the README's requirement of at least 22.19.0; its managed installer can prepare a compatible runtime. You also need a model-provider login or key. Provider charges are unspecified, and the proposed benchmark needs suitable models and hardware.
  2. Install DisCo with the README command: `curl -fsSL https://github.com/VectorSpaceLab/AREX-Skill/releases/latest/download/install-disco.sh | sh`.
  3. Run `disco repo-skills install`, then `disco`. On first run, configure a provider with `/login` or an environment variable named in the README. At its prompt, use the README's request to benchmark vLLM and SGLang—two model-serving software packages—under the same model, workload and hardware constraints, preserving commands and measurements. Treat any measurements as your own exercise, not paper results.
  4. Without installing or running anything, instead inspect a skill's `SKILL.md` and linked references or scripts in the official repository. Look for when it applies, what procedure it gives, how it checks an outcome and what gaps it records. The installation and benchmark route has not been verified for your machine.

n8n example

n8n is not necessary here: the paper studies coding-agent research procedures, not a tested n8n integration.

Was this explanation easy to understand?

06TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS

Test whether AI agents finish professional computer work

Why it matters to youIf you choose AI tools for work that ends in a file, report or design, this study offers a way to ask whether an agent finishes the job, not just whether it answers questions. Its results may help you frame tests for your own work; they are not evidence of a deployment in your workplace.

The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain experts. In each task, an agent receives instructions and works in an environment with the relevant software and input files. The authors then check what it produced against a reference output or a task-specific set of rules. The test focuses on finished work rather than a plausible account of how to do it.

Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Dawn Song, and collaborators · Agents’ Last Exam · arXiv:2606.05405 · 390 HF votes at selectionRead the paper ↗

In one sentence

The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain experts. In each task, an agent receives instructions and works in an environment with the relevant software and input files. The authors then check what it produced against a reference output or a task-specific set of rules. The test focuses on finished work rather than a plausible account of how to do it.

Key concepts

  • A benchmark is a shared test. ALE uses task instances: specific cases with their own inputs and expected outputs, drawn from professional workflows.
  • An agent combines an AI model with a harness—the software that lets it decide what to do next and use tools. For these evaluations, the authors gave the tested systems both desktop controls and command-line tools for working with files and programs.
  • A deliverable is the finished item, such as a workbook or media file. ALE scores deliverables using references or rubrics, meaning stated criteria for what the output must contain. Some tasks require a basic condition to be met before any further credit is given.
  • A full-pass rate is the share of tested tasks receiving full credit. It differs from a mean score, which also reflects partial credit.

A concrete example

Hypothetical illustration, not a reported test result: an agent receives source figures and instructions for a financial workbook. It opens the relevant software, fills in the workbook and saves the required file. A check of the saved cells against a reference would tell you more about completion than the agent’s description of the steps it took.

What the researchers measured

In the authors’ main ALE evaluation, Codex paired with GPT-5.5 and given desktop tools fully passed 38.1% of Near-Term tasks, the tier containing 67 task instances. The same configuration fully passed 0.0% of Last-Exam tasks, the tier containing 38 task instances. These percentages count tasks receiving full credit, not the amount of partial work completed. The authors report separate scores for other agent-and-model pairings; these findings should not be read as scores for GPT-5.5 on its own.

Why it matters

The authors argue that success on short tests does not establish whether an agent can carry a longer job through to a usable result. ALE makes the final deliverable central to the test. For a working reader, the distinction is between seeing an agent perform promising steps and confirming that it produced the requested work.

Where it might help

The authors present ALE as a way to evaluate agents on software-based professional work. A team considering an agent could use that approach as a model for specifying a deliverable and checking it, but the paper does not show that an ALE score predicts success in that team’s own setting.

Impact across sectors

  • Possible use in finance: a team could define checks for a completed workbook rather than assess only an agent’s written explanation. This is a hypothetical application, not a proven deployment.
  • Possible use in manufacturing planning: a team could specify required output files and safety-related checks for a software workflow. This is a hypothetical application, not a proven deployment.
  • Possible use in media production: a team could check that all requested exports exist before reviewing their content. This is a hypothetical application, not a proven deployment.

Where the evidence stops

The authors identify overlap with training data or task-specific optimization as a threat to a public benchmark. They release only 150 of the 1,490 task instances publicly and hold others privately. Some checks of visual or written outputs still use an AI model as a judge rather than a fixed code-based check. The paper says repeated runs were available for only a subset of configurations because of compute limits. Each evaluation run also had a five-hour limit. General PtoP note: this source is an arXiv preprint, not a claim of peer review, and a benchmark result is not a tested workplace deployment.

FROM PAPER TO PRACTICE

How to try it

The official repository links to a setup guide, but its supplied README does not give the commands needed to run the test here; installation and execution are therefore unverified in this guide.

  1. Read the official repository README at https://github.com/MaxIntelligenceAgency/ALE-Benchmark and follow its linked docs/quickstart.md for the actual setup instructions.
  2. Check the prerequisites the README names: a Google Cloud project, a copied sandbox image and two keys. You will also need access to an agent for the run.
  3. Check access and cost before starting. The README describes a Google Cloud free trial, but cloud and agent usage may involve charges; it does not establish a cost for your run.
  4. If you complete the linked guide using its own instructions, inspect the hello-world task’s saved output and grade. Treat what you see as your own run, not as a reproduction of the paper’s reported results.

n8n example

n8n is not a useful substitute for this test: ALE evaluates an agent working inside a software-equipped computer environment, not just a sequence of connected automation steps.

Was this explanation easy to understand?

COMPARISON

What these studies show together

Compare topics, possible uses and limits. Each result remains attributed to its own paper.

PaperCentral ideaPossible useLimit
Choosing which agent failures to revisit can improve testing resultsThe study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order beforehand. The authors built…A possible use is planning which practice tasks should guide improvements to an agent's prompts, tools, or operating rules. The paper tests…These are controlled benchmark results, not a production evaluation. The authors used public and synthetic task data, not real user data…
Humanoid robots can reuse learned movements through revised task instructionsThe study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than retraining how it moves. The authors’ answer…A possible use is preparing new task instructions for a robot that already has relevant movements, then checking candidate instructions in…The authors state that search cannot bring out behavior the controller never learned and is limited by its stored examples of…
Overlapping slices could help preserve holes in generated 3D objectsThe authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a shape through overlapping…Possible uses, not tested deployments, include making draft 3D assets from images for creative or instructional work. The paper reports…The authors say SILSA depends on the diversity and scale of its training data: shapes with unfamiliar structures may be reconstructed less…
AI agents can propose research methods while people choose directionsAn AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the authors describe in their…As a possible use, a research team could ask an agent to run bounded experiments and bring the results to a human decision-maker. The paper…This is an arXiv preprint, not a claim of peer review. The collaboration findings come from the authors’ own development project, and the…
Reusable skills may help research agents follow software instructionsThe study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its general knowledge alone. Imagine…A possible use is to give compatible coding agents inspectable package procedures and validation checks before a research task. The paper…The authors report that skills lowered the PaperBench scores on two tasks. They suggest that retrieved material may sometimes distract the…
Test whether AI agents finish professional computer workThe study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain…The authors present ALE as a way to evaluate agents on software-based professional work. A team considering an agent could use that…The authors identify overlap with training data or task-specific optimization as a threat to a public benchmark. They release only 150 of…

ARCHIVE

Previous issues

The last two issues. Every earlier edition is in the archive.

01
Model Memory, Scene Prediction, Scientific Software and Agent Workflows01-10-2026
↗
30
Chinese Answers, Agent Memory, Engineering Tests, Open Models and Simulators30-09-2026
↗