Choosing which agent failures to revisit can improve testing results
The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order…
ISSUE 06/2026 · 02-10-2026
The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain experts. In each task, an agent receives instructions and works in an environment with the relevant…

6 paper · Original sources linked in every story
5 MINUTES ON PtoP
An AI assistant can be useful without being ready to finish a job alone. Anthropic says more than 16,000 Barclays colleagues have adopted a Claude-powered assistant that helps staff find information for customer questions. That is a picture of AI woven into someone’s workday. The authors of Agents’ Last Exam ask a different question: can an agent leave behind a completed file, report or design? In their computer-work benchmark, full completion varied sharply by task tier. Finding an answer and finishing a deliverable are worth treating as different tests. See: Barclays scales Claude to upgrade operations and improve… · Test whether AI agents finish professional computer work
Several items today suggest that what surrounds an agent matters as much as the request it receives. The ActiveSaddler authors tested a way to choose practice tasks as an agent’s weaknesses changed; on two benchmarks, they report higher held-out scores than when the task order was fixed in advance. ServiceNow says its AutoSynthData system similarly creates and checks practice tasks based on a workplace helper’s mistakes, with improved performance in its own controlled tests. Neither result tells us how much improvement a particular workplace would see. See: Choosing which agent failures to revisit can improve… · AutoSynthData: Generating Training Data for Enterprise…
Practice is only one part of that surrounding structure. The DisCo authors gave a research agent reusable instructions drawn from software repositories and papers: how to set things up, check results and recover from mistakes. With task-oriented guidance, they report higher aggregate scores on research benchmarks than for the same agent without it, while noting that guidance lowered scores on two PaperBench tasks. Instructions could help a team avoid repeatedly explaining a workflow, but choosing the wrong ones might get in the way. See: Reusable skills may help research agents follow software…
Then comes the question of who decides what counts as done. In the development project described by the Atria Dawn Preview authors, agents often proposed methods and revisions, while people usually made the final choices. The Agents’ Last Exam authors, meanwhile, check finished work against reference outputs or task-specific rules rather than accepting a plausible account of the work. Together, these papers offer a useful distinction: an agent might carry out steps, but people still need to choose the aim and decide what evidence would make the result trustworthy. See: AI agents can propose research methods while people choose… · Test whether AI agents finish professional computer work
If that distinction holds in your own work, the first step may be smaller than choosing a new tool. This week, could you pick one task you might give an AI assistant and write down the finished thing you would expect to see—and the check you would use before relying on it? See: Test whether AI agents finish professional computer work
Just here for the stories? They are below, by topic.
PtoP · NEWSLETTER
Six papers and the news, explained in plain English, each morning after the editor approves the issue. Free, and you can unsubscribe at any time.
News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.
THIS ISSUE AT A GLANCE
↗02 / NEW / EDITOR PICKSHumanoid robots can reuse learned movements through revised task instructionsThe study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to…
↗03 / NEW / EDITOR PICKSOverlapping slices could help preserve holes in generated 3D objectsThe authors ask whether a model can generate detailed 3D shapes without losing how their parts connect…
↗04 / TOP VOTED · 6 MONTHSAI agents can propose research methods while people choose directionsAn AI agent can carry out extended research tasks without taking over the decisions that set their…
↗05 / TOP VOTED · 6 MONTHSReusable skills may help research agents follow software instructionsThe study asks whether a research agent does better when it can consult practical instructions for a…
↗06 / TOP VOTED · 6 MONTHSTest whether AI agents finish professional computer workThe study asks whether AI agents can complete extended professional tasks on a computer. The authors…
↗PODCAST / ENGLISH EDITION
The front-page episode covers the papers and news; each paper also has its own short conversation.
Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.
The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order…
The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than…
The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a…
An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the…
The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its…
The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE…
FROM THE NEWS DESK
Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.
Tools and building Hugging Face Blog · 02-10-2026 · company announcement
ServiceNow says it built a way to create practice tasks for workplace AI helpers based on what they get wrong. Imagine a helper learning to update a support ticket while following company rules, then practicing that skill on a different ticket. The system, AutoSynthData, makes and checks new tasks before using them to train the helper.
As an illustration, if a helper wrongly closes a ticket before an approval arrives, a practice task could ask it to handle a different ticket that also needs approval. A check would reject a result that skips that step.
A company could use this approach to train an AI helper on difficult tasks in its own support workflow.
These are results from ServiceNow's own tests in EnterpriseOps Gym, a controlled workplace-task test environment. They do not establish that the approach works equally well in a company's live systems.
Practice tasks may be more useful when they target a helper's current mistakes and are checked for valid solutions. ServiceNow reports improved test performance after training on such tasks.
Was this explanation easy to understand?
Business and strategy Anthropic · 01-10-2026 · company announcement
Anthropic says Barclays is expanding its use of Claude to help staff answer customer questions, handle emails and develop software. For example, a bank worker can use an existing assistant to find information for a customer. Anthropic says more than 16,000 Barclays colleagues have adopted that assistant.
If a customer asks about a banking service, an employee can use the Colleague Knowledge Assistant to find relevant information before replying. This illustrates the use Anthropic describes; it does not show how much time any particular customer saves.
Anthropic reports that Barclays uses Claude to help route incoming emails in its Global Markets business, so staff can identify requests that need action.
These are claims in an Anthropic announcement, not independently measured results presented here. The stated goal for Claude Code to reach 50% of Barclays developers by the end of 2026 is a forecast, not an achieved outcome.
Barclays already uses a Claude-powered knowledge assistant and is extending the technology to other work, but the promised gains and future reach remain uncertain.
Was this explanation easy to understand?
AI agents Google · 30-09-2026 · company announcement
Google announced Gemini 4 Argon, an AI model it says can carry out long coding and security tasks with less step-by-step help. Google is initially sharing it with selected cyber defenders rather than releasing it broadly.
For example, a hospital software team could ask an AI system to look for a vulnerability, suggest a fix, and let people review that fix before using it. This illustrates the kind of task Google describes, not a verified use of Argon.
Selected security teams could use Argon to help find and repair weaknesses in software.
The performance claims come from Google’s announcement. Google says it is still testing guardrails and gathering feedback before wider access.
Argon is a limited rollout of an AI model designed to take on longer, more independent work; its broader usefulness and safety remain to be tested.
Was this explanation easy to understand?
Tools and building xAI · 18-09-2026 · company announcement
xAI says it has released Grok Voice Transcribe 2.0, a tool that turns speech into written words. The company says it handles noisy calls and multiple languages better than its earlier version, at the same price.
Imagine a customer-support call with a poor connection. The tool could write down what was said and label which person spoke.
A support team could use the resulting transcription to review calls without listening to each recording.
The broad accuracy claims rely on xAI's own evaluations, though it also cites a public ranking. Performance on a particular call may differ.
This is a company-reported improvement aimed at making everyday audio easier to turn into usable text.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS
Why it matters to youIf you help improve an AI assistant at work, you may need to decide which failed practice tasks deserve another attempt. This paper suggests that changing that choice as the assistant changes can improve results in benchmark tests.
The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order beforehand. The authors built ActiveSaddler to make that choice during offline optimization, before any deployment. A harness is the set of prompts, available tools, and operating rules around an AI model. ActiveSaddler does not replace the method that repairs this harness; it chooses which training tasks provide evidence for the next repair. It groups related failures, estimates which known weakness may be worth another attempt, and sometimes tries an unseen task to look for a new one. The authors tested this approach with the AutoSaddler harness optimizer on GAIA2 and Terminal-Bench 2.0.
The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order beforehand. The authors built ActiveSaddler to make that choice during offline optimization, before any deployment. A harness is the set of prompts, available tools, and operating rules around an AI model. ActiveSaddler does not replace the method that repairs this harness; it chooses which training tasks provide evidence for the next repair. It groups related failures, estimates which known weakness may be worth another attempt, and sometimes tries an unseen task to look for a new one. The authors tested this approach with the AutoSaddler harness optimizer on GAIA2 and Terminal-Bench 2.0.
Hypothetical illustration, not a reported result: An assistant repeatedly omits attachments from different calendar tasks. Instead of treating each failed task as unrelated, a curriculum could group them as one possible weakness, revisit it after a proposed repair, and then try an unseen calendar task to check for other failures.
The authors report mean test Pass@1 over three test-time executions; this score is the share of held-out tasks passed on a single attempt. With ActiveSaddler added to AutoSaddler, they report 59.8% on the GAIA2 test split and 80.0% on the Terminal-Bench 2.0 test split. These are, respectively, 4.4 and 7.5 percentage points above the same optimizer using a randomly shuffled scenario order fixed before optimization. In separate component-removal tests, the authors report lower test scores when failure-pattern grouping, arm prioritization, or adaptive exploration is replaced. Those tests concern the named variants, not a different deployed system.
A fixed practice schedule can revisit a weakness that has already been addressed or move past one that remains. The authors' results suggest that the choice of training tasks matters alongside the method used to change the harness.
A possible use is planning which practice tasks should guide improvements to an agent's prompts, tools, or operating rules. The paper tests that idea in benchmark environments, not in a workplace deployment.
These are controlled benchmark results, not a production evaluation. The authors used public and synthetic task data, not real user data, and say the experiments do not establish safety, security, privacy, or governance readiness. GAIA2 training, development, and test tasks came from disjoint simulated environments; Terminal-Bench 2.0 used a uniform random partition because it lacks a natural grouping axis. The reported test scores average three test-time executions. ActiveSaddler also adds optimizer-side computation. The paper says a project website and code will be available, but the supplied text does not verify an available installation. General PtoP note: this arXiv version is a preprint; the supplied text does not establish peer review.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation is unverified, and the paper's linked project website and code are described as forthcoming. Prerequisites: a few fictional practice tasks and a way to record outcomes. No account, payment, model access, or real execution traces are needed for this exercise.
An n8n workflow would not reproduce this study: the paper does not specify a verified n8n integration, and its method depends on harness updates, execution records, and repeated evaluation.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS
Why it matters to youIf you plan tasks for a warehouse or research robot, this study suggests a possible way to try new box-handling tasks without retraining its movement controller. The authors tested the approach mainly in simulation, not as a warehouse deployment.
The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than retraining how it moves. The authors’ answer is to keep its movement controller fixed and revise a reward program: a set of goals for successive parts of a task. Imagine moving a box onto a table. One part can favor lifting and carrying it; a later part can favor letting go. A completion condition tells the system when to switch parts. InterEvolve has a language-model agent revise those parts after simulated attempts, while a numerical search adjusts settings such as how strongly each goal counts. A separate, fixed verifier checks whether the attempt met the task’s requirements. Successful programs can be kept in a skill library for later tasks.
The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than retraining how it moves. The authors’ answer is to keep its movement controller fixed and revise a reward program: a set of goals for successive parts of a task. Imagine moving a box onto a table. One part can favor lifting and carrying it; a later part can favor letting go. A completion condition tells the system when to switch parts. InterEvolve has a language-model agent revise those parts after simulated attempts, while a numerical search adjusts settings such as how strongly each goal counts. A separate, fixed verifier checks whether the attempt met the task’s requirements. Successful programs can be kept in a skill library for later tasks.
Hypothetical illustration, not a reported trial: for a request to place a parcel on a shelf, one stage could favor getting hold of it and moving it upward. A second could favor releasing it once it reaches the shelf. A verifier would separately check whether the parcel stayed there and the robot let go.
On the authors’ eight simulated box-task families, InterEvolve’s evolved programs achieved 86.5% success, compared with 34.6% for the agent’s initial program after numerical tuning. Success means every task criterion and physical constraint held in the same attempt. The authors also show a physical Unitree G1 executing simulation-evolved kicking and pushing tasks using onboard sensing and an off-board workstation; they do not give a physical-task success rate in the supplied text.
The authors separate task strategy from movement training. In their experiments, changing the goals and the order in which they apply let the same controller use movements that a fixed program did not reliably bring out.
A possible use is preparing new task instructions for a robot that already has relevant movements, then checking candidate instructions in simulation before physical execution. The paper does not establish a general-purpose deployment process.
The authors state that search cannot bring out behavior the controller never learned and is limited by its stored examples of body-and-object states and by available measurements. Simulation and language-model use take GPU time and add latency, ruling out real-time replanning during physical execution. The main tracking test uses held-out clips from objects represented in training; the additional box types tested for transfer were also seen during controller pretraining, though not during program search. The physical demonstration covers two tasks, with no reported success rate. In that setup, prompts prepared from a simulated run are replayed on the robot while its fixed controller responds to sensed state; the paper does not report physical test-time program evolution. PtoP note: this arXiv version is a preprint, not evidence of peer review.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation is unverified. Prerequisites: paper or pencil and a familiar multi-step object-moving task. No software access or cost is required for this exercise; reproducing the study would require resources not established by verified installation instructions here.
An n8n workflow is not appropriate here: the paper does not describe a verified automation interface for its robot, simulator or program search.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS
Why it matters to youIf you make 3D assets from reference images, this study offers a way to think about a familiar problem: a shape can look close to the image but lose a hole or a thin support. The authors test a method designed to keep those structures intact.
The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a shape through overlapping cross-sections viewed from three directions, rather than many tiny 3D cells. Imagine slicing a bicycle wheel from front to back: successive slices reveal where the open center and spokes begin and end. SILSA compresses such slices into tokens, small packets of shape information. A reconstruction model learns to turn those packets back into a 3D surface. A separate image-guided generator learns to produce the packets from a single picture. The authors also train the reconstruction model to notice connected parts and holes, while a shared 3D workspace helps slices from different directions describe the same object.
The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a shape through overlapping cross-sections viewed from three directions, rather than many tiny 3D cells. Imagine slicing a bicycle wheel from front to back: successive slices reveal where the open center and spokes begin and end. SILSA compresses such slices into tokens, small packets of shape information. A reconstruction model learns to turn those packets back into a 3D surface. A separate image-guided generator learns to produce the packets from a single picture. The authors also train the reconstruction model to notice connected parts and holes, while a shared 3D workspace helps slices from different directions describe the same object.
Hypothetical illustration, not a reported test: A designer supplies one image of a bicycle wheel. A useful 3D result would keep its center open and its thin spokes connected to the rim, rather than filling the opening or breaking a spoke.
The authors trained SILSA on Trellis-500K and evaluated using 200 randomly sampled Toys4K assets and 50 in-the-wild images, which they say did not overlap with training. In their image-to-3D comparison, SILSA had leading scores on the reported FD, PSNR, coverage and MMD measures; it tied the best KD and LPIPS scores. These measures assess aspects of generated-shape quality, image fidelity or how well outputs represent the reference set. Its PSNR, an image-fidelity score where higher is better, was 32.74 versus 30.12 for SparseFlex. SparseFlex had a slightly higher CLIP score, the paper’s input-image alignment measure. Separately, in the SliceVAE reconstruction comparison, the authors report a lower Betti error—a measure of mismatched connected parts and holes—than the reconstruction baselines. SILSA used 384 fixed slice tokens. In the reported efficiency comparison, its training memory and end-to-end inference time were lower than Dora’s under the stated test settings.
The distinction between a close-looking surface and a correctly connected object matters when a missing opening or broken support changes the shape itself. The authors report both structural measures and generation-cost measures, so readers can see what happened in their experiments rather than infer structure from appearance alone.
Possible uses, not tested deployments, include making draft 3D assets from images for creative or instructional work. The paper reports generation and reconstruction experiments, not an end-to-end workplace workflow.
The authors say SILSA depends on the diversity and scale of its training data: shapes with unfamiliar structures may be reconstructed less faithfully. Ambiguous or low-information input images may also reduce generation quality. Their open-surface garment evaluation notes that topology measures are near zero for methods that recover the rough surface, limiting what that measure distinguishes there. The authors also identify risks of unauthorized replicas and displacement of manual modeling work. General PtoP note: this arXiv source is a preprint; the reported benchmarks are not a tested deployment.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and model access are unverified. Prerequisites: paper and pencil, plus an everyday object with an opening, such as a mug. No software cost is needed for this exercise; access or cost for running SILSA is not established here.
An n8n workflow is not appropriate here: the supplied paper does not verify a runnable SILSA interface for a proposed integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS
Why it matters to youIf you lead research or engineering work, this study may help you think about which parts of a project to give an AI agent—and which decisions to keep with people. It describes one model-development project, not a tested way to run every team.
An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the authors describe in their development of Atria Dawn Preview, a model designed to use tools across research and engineering work. They trained it on tasks in working environments, keeping records of its steps and checking outcomes against evidence such as tests or file state. They then compared its performance on set tasks and examined task records and agent logs from the people developing it. In those records, agents often proposed methods and made revisions; people usually made the final choices.
An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the authors describe in their development of Atria Dawn Preview, a model designed to use tools across research and engineering work. They trained it on tasks in working environments, keeping records of its steps and checking outcomes against evidence such as tests or file state. They then compared its performance on set tasks and examined task records and agent logs from the people developing it. In those records, agents often proposed methods and made revisions; people usually made the final choices.
Hypothetical example, not a paper result: an engineering lead asks an agent to investigate a failing data-processing job. The agent checks logs, suggests a change and runs a test. The lead decides whether that test answers the right question and whether the change should be kept.
The authors report that Atria Dawn Preview had the highest reported score on five of 16 benchmarks; rankings use the available entries, and some comparison results were unreported. In their project review, participants rated 151 of 455 completed AI-assisted tasks with usable responses as infeasible without AI under the same scope and resources. Among participants in execution roles, humans made the final choice in 85.5% of recorded method-or-parameter decisions. That figure concerns who selected an option, including options proposed by AI; it is not a measure of task success.
The authors’ distinction is between doing more work and deciding what work is worth doing. Their project records show substantial agent involvement alongside continued human selection and intervention. They say completing research tasks does not, by itself, show that an AI system can repeatedly discover worthwhile improvements to future models.
As a possible use, a research team could ask an agent to run bounded experiments and bring the results to a human decision-maker. The paper does not establish a general deployment procedure or show that an agent can choose valuable research directions on its own.
This is an arXiv preprint, not a claim of peer review. The collaboration findings come from the authors’ own development project, and the infeasibility finding is based on participants’ ratings of their tasks. The paper calls its tool-use cases selected demonstrations rather than estimates of average success. Evaluation coverage also varies: SkillsBench uses a self-contained subset that excludes multimodal tasks, and one optimization case had no complete formal mean at the end of its published trace. The authors say the weather interface does not report forecast accuracy and the displayed computer-aided designs are not mechanical or manufacturing validation.
FROM PAPER TO PRACTICE
The official repository links to Atria Dawn Preview’s model page and hosted access, but its README gives no complete local installation commands. The following is an untested, text-only exploration; pricing and the computing needed for local use are not specified there.
An n8n automation is not appropriate for this first exercise: the point is to inspect the agent’s evidence and make a deliberate human choice, rather than automate that choice.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS
Why it matters to youIf you use a coding agent for machine-learning experiments, this study offers a way to give it relevant setup steps and checks before it starts. The authors tested that approach on research benchmarks, not in an everyday workplace deployment.
The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its general knowledge alone. Imagine asking an agent to use an unfamiliar software package: it needs to know not just what the package does, but how to set it up, check its output and recover from mistakes. The authors call this practical know-how operational knowledge. Their DisCo system turns material from repositories and papers into reusable instruction bundles called skills. An entry file says when a skill applies; supporting files can hold detail and executable helpers. A router helps the agent open only the relevant bundle. The authors compared a Codex research agent with and without these skills while keeping its GPT-5.5 model, research harness—the software coordinating its work—and downstream running budget fixed.
The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its general knowledge alone. Imagine asking an agent to use an unfamiliar software package: it needs to know not just what the package does, but how to set it up, check its output and recover from mistakes. The authors call this practical know-how operational knowledge. Their DisCo system turns material from repositories and papers into reusable instruction bundles called skills. An entry file says when a skill applies; supporting files can hold detail and executable helpers. A router helps the agent open only the relevant bundle. The authors compared a Codex research agent with and without these skills while keeping its GPT-5.5 model, research harness—the software coordinating its work—and downstream running budget fixed.
Hypothetical illustration, not a study result: a researcher asks an agent to compare two unfamiliar model-serving packages. A relevant skill might tell it to use the same workload for both, retain the commands it ran and check the measurements before reporting them.
The authors report that, on the full MLE-bench suite of 75 machine-learning competitions, GPT-5.5 Codex with task-oriented AREX-Skill guidance received an Any-Medal score of 72.89%, versus 31.11% without skills. That score reports how often the agent attained any medal. On PaperBench's 20 paper-reproduction tasks, they report an average replication score of 39.59% with the selected paper-derived skills, versus 29.45% without them; that score is assigned by the benchmark's reproduction grader. The paper also reports higher aggregate scores with skills on FrontierCS and PassNet. Those evaluations used their own task-oriented graphs, not simply the public repository collection.
The authors' comparison separates access to distilled instructions from changes to the agent's model or research harness. It therefore speaks to what practical guidance may add under their benchmark settings, while the separate effort of creating that guidance still matters.
A possible use is to give compatible coding agents inspectable package procedures and validation checks before a research task. The paper measures benchmark work; it does not establish that this approach succeeds in a particular workplace.
The authors report that skills lowered the PaperBench scores on two tasks. They suggest that retrieved material may sometimes distract the agent from a better task-specific approach, but present this as a possible explanation, not an established cause. Skill construction happened before benchmark execution and used a separate budget; matched downstream running budgets do not include that construction cost. The public repository snapshot and the skill collections used in the evaluations are distinct. For PaperBench, the authors excluded each target paper and its released artifacts as skill sources, while selecting related prior work. The repository collection is described as curated rather than exhaustive. General PtoP note: this is an arXiv preprint, not a peer-reviewed finding, and benchmark scores are not deployment results.
FROM PAPER TO PRACTICE
Advanced coding exercise using the official repository's README; these steps are not a tested procedure here. No-execution alternative included.
n8n is not necessary here: the paper studies coding-agent research procedures, not a tested n8n integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS
Why it matters to youIf you choose AI tools for work that ends in a file, report or design, this study offers a way to ask whether an agent finishes the job, not just whether it answers questions. Its results may help you frame tests for your own work; they are not evidence of a deployment in your workplace.
The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain experts. In each task, an agent receives instructions and works in an environment with the relevant software and input files. The authors then check what it produced against a reference output or a task-specific set of rules. The test focuses on finished work rather than a plausible account of how to do it.
The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain experts. In each task, an agent receives instructions and works in an environment with the relevant software and input files. The authors then check what it produced against a reference output or a task-specific set of rules. The test focuses on finished work rather than a plausible account of how to do it.
Hypothetical illustration, not a reported test result: an agent receives source figures and instructions for a financial workbook. It opens the relevant software, fills in the workbook and saves the required file. A check of the saved cells against a reference would tell you more about completion than the agent’s description of the steps it took.
In the authors’ main ALE evaluation, Codex paired with GPT-5.5 and given desktop tools fully passed 38.1% of Near-Term tasks, the tier containing 67 task instances. The same configuration fully passed 0.0% of Last-Exam tasks, the tier containing 38 task instances. These percentages count tasks receiving full credit, not the amount of partial work completed. The authors report separate scores for other agent-and-model pairings; these findings should not be read as scores for GPT-5.5 on its own.
The authors argue that success on short tests does not establish whether an agent can carry a longer job through to a usable result. ALE makes the final deliverable central to the test. For a working reader, the distinction is between seeing an agent perform promising steps and confirming that it produced the requested work.
The authors present ALE as a way to evaluate agents on software-based professional work. A team considering an agent could use that approach as a model for specifying a deliverable and checking it, but the paper does not show that an ALE score predicts success in that team’s own setting.
The authors identify overlap with training data or task-specific optimization as a threat to a public benchmark. They release only 150 of the 1,490 task instances publicly and hold others privately. Some checks of visual or written outputs still use an AI model as a judge rather than a fixed code-based check. The paper says repeated runs were available for only a subset of configurations because of compute limits. Each evaluation run also had a five-hour limit. General PtoP note: this source is an arXiv preprint, not a claim of peer review, and a benchmark result is not a tested workplace deployment.
FROM PAPER TO PRACTICE
The official repository links to a setup guide, but its supplied README does not give the commands needed to run the test here; installation and execution are therefore unverified in this guide.
n8n is not a useful substitute for this test: ALE evaluates an agent working inside a software-equipped computer environment, not just a sequence of connected automation steps.
Was this explanation easy to understand?
COMPARISON
Compare topics, possible uses and limits. Each result remains attributed to its own paper.
| Paper | Central idea | Possible use | Limit |
|---|---|---|---|
| Choosing which agent failures to revisit can improve testing results | The study asks whether an AI agent improves more when its practice tasks are chosen as its weaknesses change, rather than put in order beforehand. The authors built… | A possible use is planning which practice tasks should guide improvements to an agent's prompts, tools, or operating rules. The paper tests… | These are controlled benchmark results, not a production evaluation. The authors used public and synthetic task data, not real user data… |
| Humanoid robots can reuse learned movements through revised task instructions | The study asks whether a humanoid can tackle a new object-handling task by changing what it is asked to achieve, rather than retraining how it moves. The authors’ answer… | A possible use is preparing new task instructions for a robot that already has relevant movements, then checking candidate instructions in… | The authors state that search cannot bring out behavior the controller never learned and is limited by its stored examples of… |
| Overlapping slices could help preserve holes in generated 3D objects | The authors ask whether a model can generate detailed 3D shapes without losing how their parts connect. Their answer is to describe a shape through overlapping… | Possible uses, not tested deployments, include making draft 3D assets from images for creative or instructional work. The paper reports… | The authors say SILSA depends on the diversity and scale of its training data: shapes with unfamiliar structures may be reconstructed less… |
| AI agents can propose research methods while people choose directions | An AI agent can carry out extended research tasks without taking over the decisions that set their direction. That is the pattern the authors describe in their… | As a possible use, a research team could ask an agent to run bounded experiments and bring the results to a human decision-maker. The paper… | This is an arXiv preprint, not a claim of peer review. The collaboration findings come from the authors’ own development project, and the… |
| Reusable skills may help research agents follow software instructions | The study asks whether a research agent does better when it can consult practical instructions for a task, rather than work from its general knowledge alone. Imagine… | A possible use is to give compatible coding agents inspectable package procedures and validation checks before a research task. The paper… | The authors report that skills lowered the PaperBench scores on two tasks. They suggest that retrieved material may sometimes distract the… |
| Test whether AI agents finish professional computer work | The study asks whether AI agents can complete extended professional tasks on a computer. The authors built Agents’ Last Exam, or ALE, around work contributed by domain… | The authors present ALE as a way to evaluate agents on software-based professional work. A team considering an agent could use that… | The authors identify overlap with training data or task-specific optimization as a threat to a public benchmark. They release only 150 of… |
ARCHIVE
The last two issues. Every earlier edition is in the archive.
01