Chinese-Jev sorts Chinese text into preset answers
The study asks whether a compact model can make fixed-answer decisions in Chinese instead of writing open-ended replies. The authors…
ISSUE 04/2026 · 30-09-2026
The study asks whether AI agents can complete engineering work across professional software, not just operate the software. The authors built EngiWorld, a benchmark—a set of tasks and checks—with 1,301 tasks across six engineering fields and 26 software platforms or…

6 paper · Original sources linked in every story
News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.
THIS ISSUE AT A GLANCE
↗02 / NEW / EDITOR PICKSShared Memory Could Help AI Assistants Reuse Visual LessonsThe study asks whether AI assistants can reuse lessons from earlier questions, even when a different…
↗03 / NEW / EDITOR PICKSCheck engineering files before trusting an AI agent’s completed workThe study asks whether AI agents can complete engineering work across professional software, not just…
↗04 / TOP VOTED · 6 MONTHSKimi K3 offers another model to consider for complex coding workThe study asks how far an open-weight model can go on long, complicated tasks. The authors built Kimi K3…
↗05 / TOP VOTED · 6 MONTHSPredicting experiment outcomes may cut research-agent training timeThe authors report that an agent can be trained with fewer real program runs by predicting most outcomes…
↗06 / TOP VOTED · 6 MONTHSStudent Simulators Could Help Tutors Test Different GuidanceThe authors ask whether a simulated student can do two things at once: make the kinds of mistakes a…
↗PODCAST / ENGLISH EDITION
The front-page episode covers the papers and news; each paper also has its own short conversation.
Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.
The study asks whether a compact model can make fixed-answer decisions in Chinese instead of writing open-ended replies. The authors…
The study asks whether AI assistants can reuse lessons from earlier questions, even when a different system learned them. The authors’…
The study asks whether AI agents can complete engineering work across professional software, not just operate the software. The…
The study asks how far an open-weight model can go on long, complicated tasks. The authors built Kimi K3 to work across extended text…
The authors report that an agent can be trained with fewer real program runs by predicting most outcomes and checking a small share…
The authors ask whether a simulated student can do two things at once: make the kinds of mistakes a particular student makes, and…
FROM THE NEWS DESK
Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.
Anthropic · 29-09-2026
Anthropic reports that GLM-5.3 built working cyberattacks in controlled tests and argues that its safeguards are easy to bypass. In a separate test using harmful requests, simple techniques got the model to respond 64% to 100% of the time.
Imagine visiting a booby-trapped web page that reads a file from your computer. Anthropic says a researcher used GLM-5.3 to build such an exploit against a browser running in an isolated test environment, not against public users.
Security teams could use capable models to find vulnerabilities in software and help fix them before attackers do.
The harmful-request test was a simulation: no code generated by the model was executed. Its response rates do not show how often real attacks would succeed. The browser exploit was demonstrated on a Linux test system; its effect on other systems remains uncertain.
Anthropic’s measured exploit results and separate safeguard tests support its warning about misuse, but the article’s predictions of real-world harm are an assessment, not a measured outcome.
Was this explanation easy to understand?
Hugging Face Blog · 29-09-2026
NVIDIA says it has released a tool that predicts missing answers in tables by learning from examples already in the table. Called Kumo Tabular, it is pretrained on artificial tables, so it does not need to be trained from scratch for each new question. NVIDIA reports that it ranked first on several benchmark comparisons; these are company-reported tests, not independent findings.
Imagine a shop’s order table showing which past orders received refunds. The tool could estimate whether a new order is likely to receive one. That estimate would not, by itself, justify issuing a refund without review.
A business could use it to flag orders for staff to examine.
NVIDIA warns that accuracy may fall when new rows differ from the examples or a table is far outside the sizes used in training. Users should test it on their own data before relying on it.
NVIDIA’s tests suggest faster, accurate predictions for some table-based tasks, but results for a particular business still need checking.
Was this explanation easy to understand?
xAI · 22-09-2026
SpaceXAI says Grok Bot helped its customer support team handle more requests without hiring more people. It checks each support ticket—a request for help—and can answer customers or flag problems for the team.
If someone asks for a refund, Grok Bot can follow the company’s instructions. SpaceXAI says 99% of refund requests are resolved without human intervention, meaning without a person stepping in—not that every request is.
A support team could use a similar tool to spot repeated customer problems and bring them to engineers sooner.
This is a company account, not an independent study. It does not say how often the bot makes mistakes or whether another team would get the same results. People still set guardrails and handle cases that need judgment.
The idea is to let the bot handle routine requests and spot patterns, while people check its work and take on harder cases.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS
Why it matters to youIf you sort Chinese customer messages or route requests at work, this research offers a possible way to make a preset choice without asking a model to write a reply. The paper tests that approach on research datasets, not in a customer-service deployment.
The study asks whether a compact model can make fixed-answer decisions in Chinese instead of writing open-ended replies. The authors built Chinese-Jev to read a question and its possible answers together, then assign each answer a probability. They trained a general version on Chinese-language decisions and separately adapted copies for medicine, law, and finance.
The study asks whether a compact model can make fixed-answer decisions in Chinese instead of writing open-ended replies. The authors built Chinese-Jev to read a question and its possible answers together, then assign each answer a probability. They trained a general version on Chinese-language decisions and separately adapted copies for medicine, law, and finance.
Hypothetical illustration, not a paper result: A service desk receives a Chinese message asking to move tomorrow’s appointment. It could offer “reschedule,” “cancel,” and “other” as candidates and use a decision model to score them. A person or separate system would still decide what to do next.
On the 100,000-decision General part of Chinese-Jev Bench, the authors report 69.20% accuracy for Chinese-Jev General, compared with 68.35% for the hosted Jev model. Accuracy here is the share of eligible single-label decisions where the highest-probability candidate agrees with the dataset label; examples with soft targets are excluded. After separate domain training, the authors report that the medical specialist exceeds Jev’s medical accuracy, while the legal and financial specialists remain below it. They measured 14 milliseconds per General decision locally on one graphics processor, versus 284 milliseconds for a request to Jev’s hosted service. The latter includes network travel, so these figures do not directly compare model computation speed. They also report roughly one second per decision for a compressed Chinese-Jev model in a browser prototype on an iPhone 15 Pro.
Many work tasks need a label rather than a written answer. The authors’ results show how they tested that distinction for Chinese-language decisions, while also showing that general training did not transfer evenly to specialist tasks.
A possible use is sorting requests where the allowed outcomes are known in advance. The paper measures answer selection on its benchmark; it does not establish how Chinese-Jev would perform in a live workplace workflow.
The authors report remaining accuracy gaps against Jev in law and finance. Confidence calibration—how closely stated probabilities track observed correctness—did not improve consistently after specialist training; it worsened on their legal benchmark. Accuracy and calibration calculations exclude soft-target examples. The authors grouped decisions from shared source material into the same data partition and screened against available held-out text, but say paraphrases across the full training corpus were not exhaustively merged. The hosted-service timing includes network travel, unlike the local timing. Some source datasets retain noncommercial terms. The authors say they will release models, data, the benchmark, and construction code; the supplied text does not establish that these are downloadable now. General PtoP note: this source is an arXiv preprint, not an indication of peer review, and benchmark results are not a tested workplace deployment.
FROM PAPER TO PRACTICE
Safe paper-only exercise. Prerequisites: the supplied paper text, a few Chinese-language example messages you are permitted to use, and predefined labels. No model installation or paid access is needed for this exercise; downloadable model access, installation steps, and any service cost are unverified.
An n8n automation example would be premature: the supplied paper does not verify an accessible Chinese-Jev model or callable service to connect to a workflow.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS
Why it matters to youIf you use AI assistants to work through documents, charts or diagrams, this research may help you understand how one assistant’s lessons could guide another. The authors tested shared memory on set questions, not in an everyday workplace.
The study asks whether AI assistants can reuse lessons from earlier questions, even when a different system learned them. The authors’ answer is that, in their tests, a stored bank of guidance and image details helped some other systems solve new questions without changing their answering models. Imagine an assistant learning that a transit-map question requires checking a particular junction: it could save both that advice and the relevant part of the map. EpiCon has one small model that revises such notes and image regions across attempts at a question, and another that groups lessons for later retrieval. Surrounding software checks the proposed changes and manages storage. In the main tests, the stored bank was frozen, and each question received just one solving attempt.
The study asks whether AI assistants can reuse lessons from earlier questions, even when a different system learned them. The authors’ answer is that, in their tests, a stored bank of guidance and image details helped some other systems solve new questions without changing their answering models. Imagine an assistant learning that a transit-map question requires checking a particular junction: it could save both that advice and the relevant part of the map. EpiCon has one small model that revises such notes and image regions across attempts at a question, and another that groups lessons for later retrieval. Surrounding software checks the proposed changes and manages storage. In the main tests, the stored bank was frozen, and each question received just one solving attempt.
Hypothetical illustration, not a measured result: an assistant answering a new map question retrieves a note to check whether a route crosses a junction, along with a saved image region showing an earlier junction. It uses the note to decide what to inspect in the new map; the earlier image does not replace the new map.
The authors evaluated the trained two-model EpiCon variant on 11 benchmarks with four pairings: Codex or DeepSeek-Harness as the harness, each with Qwen3.8-27B or Gemma4-31B as the answering backbone. With one attempt per question, its macro-average scores improved by 1.7 to 4.9 points over No Memory across those pairings. A macro-average gives each benchmark equal weight; a point is a difference on that average, not one additional correct answer. In a separate frozen-bank transfer test, a bank built by Codex with Qwen3.8-27B raised the 11-benchmark macro-average by 3.2 points when reused by Codex with Gemma4-31B, and by 4.2 points when reused by DeepSeek-Harness with Qwen3.8-27B. The authors also report that the trained small memory models reduced memory-operation time by 67% to 74% relative to EpiCon configurations using the answering backbone for memory operations; their average scores were lower by 0.9 to 3.6 points. These are comparisons between memory-model variants, not claims that the small answering systems outscored larger agents.
The distinction is between changing an answering model and giving it a record it can consult. The authors tested whether that record could remain useful when the answering model or its surrounding software changed.
A possible use is letting teams of AI assistants retain lessons about recurring visual tasks without retraining the answering model. The paper measures question-solving on benchmarks; it does not demonstrate a workplace integration.
The authors’ construction questions came from the same benchmarks as the evaluation questions, though the project-defined question splits were disjoint; ReasonMap was also separated by source map. The bank-evolution comparison added a second set of construction questions, so it measures evolution and expansion together. Improvements varied by task, and some individual benchmark scores fell. The separate question-level memory tests allowed up to five attempts for memory-enabled settings versus one for No Memory, so they measure refinement including extra attempts. Training checks established structural validity for tree-operation examples, but individual operations were not verified through later assistant runs. General PtoP note: this arXiv source is a preprint, not evidence of peer review or a deployed service.
FROM PAPER TO PRACTICE
Safe conceptual exercise; the paper gives a project page, but installation, model access and executable steps are unverified here. Prerequisites: two sample image-based questions and somewhere to write notes; no paid service is needed for this paper exercise.
An n8n workflow is not appropriate to specify from this source because it does not document a verified n8n integration or installation procedure.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS
Why it matters to youIf you use design, simulation, or manufacturing software, this study could help you decide what to check when an AI agent says a job is finished. Its tests show why an apparently completed workflow still needs the delivered files checked.
The study asks whether AI agents can complete engineering work across professional software, not just operate the software. The authors built EngiWorld, a benchmark—a set of tasks and checks—with 1,301 tasks across six engineering fields and 26 software platforms or workbenches. An agent might receive a drawing, choose or operate tools, make a design, and submit files. Instead of accepting its claim that it is done, the authors’ checking programs reopen the files and test requirements such as dimensions, physical outputs, connections, and whether details survive a move between applications.
The study asks whether AI agents can complete engineering work across professional software, not just operate the software. The authors built EngiWorld, a benchmark—a set of tasks and checks—with 1,301 tasks across six engineering fields and 26 software platforms or workbenches. An agent might receive a drawing, choose or operate tools, make a design, and submit files. Instead of accepting its claim that it is done, the authors’ checking programs reopen the files and test requirements such as dimensions, physical outputs, connections, and whether details survive a move between applications.
Hypothetical illustration, not a tested paper result: an agent draws a mounting bracket in one application and opens it in another for a strength calculation. A picture of the bracket might look right, but a file check could reveal a missing bolt hole or an incorrect load setting.
The authors tested seven models on the same 300-task subset of EngiWorld, not on all 1,301 tasks. Claude Opus 5 (Max) had the highest overall EngiScore, 44.3 out of 100. This score gives equal weight to each tested task: most receive credit only if every required check passes, while feasible quantitative designs can receive partial credit for quality. Across all seven models, 6 of 168 attempts at the subset’s multi-software tasks succeeded—about 3.6%. The authors also report that all seven models scored zero on the same 128 tasks in the subset.
The authors’ failure analysis separates two problems a working team would want to notice: an agent may declare completion without files that pass the checks, or it may run out of permitted decisions before finishing. A visible result or a completion message is therefore not the same measure as a checked engineering deliverable.
As a possible use, a team considering AI assistance for engineering work could use the paper’s distinction between operating tools and checking deliverables to plan its own review. The paper does not demonstrate a workplace deployment.
The main comparison uses a 300-task subset selected to control evaluation cost, so it is not a result for every EngiWorld task. The authors’ verifier audit covers a completed group of reference files and selected alternative and defective submissions; tolerance-boundary tests are outside those completed audit populations. Their paired comparison of prescribed multi-software and open-ended work covers two shared design cases, which have different delivery requirements. General PtoP note: the supplied source is an arXiv preprint, not evidence of peer review, and benchmark results are not a demonstrated workplace deployment.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and runnable benchmark access are unverified from the supplied text. Prerequisites: an engineering requirement and a sample file you are permitted to inspect. Access to professional software may require a licence or payment; no benchmark setup or cost is verified here.
An n8n workflow—a proposed way to connect automated steps—is not appropriate as a paper example here: the study tests agents inside native engineering software and checks their files, but does not describe or test an n8n integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS
Why it matters to youIf you build software, this paper may help you decide whether Kimi K3 belongs on your list of coding assistants to investigate. It reports benchmark tests and development examples, not a guarantee about work in your own codebase.
The study asks how far an open-weight model can go on long, complicated tasks. The authors built Kimi K3 to work across extended text, code and visual input, then trained it on tasks that involve acting, checking outcomes and revising. Picture a coding assistant reading a project, changing a file, running a test and using the failure message to try again. That is an illustration of the kind of work the training targets; the paper does not show that every such task will succeed.
The study asks how far an open-weight model can go on long, complicated tasks. The authors built Kimi K3 to work across extended text, code and visual input, then trained it on tasks that involve acting, checking outcomes and revising. Picture a coding assistant reading a project, changing a file, running a test and using the failure message to try again. That is an illustration of the kind of work the training targets; the paper does not show that every such task will succeed.
Hypothetical example, not a reported result: a developer asks for a small change to a website form. An assistant edits the validation rule, runs the existing tests, reads a failing test and revises the edit. A person then checks the behavior and the final change.
In the authors’ in-house Kimi Webdev Bench, blind expert judges compared Kimi K3 with Claude Opus 4.8. Both ran at maximum reasoning effort with the Claude Code harness. Judges preferred Kimi K3’s output on 58.6% of prompts and Claude Opus 4.8’s on 27.6%; these percentages count judged prompts, not successful workplace deployments. The authors also report coding benchmark and case-study results, but those use their named tasks and setups rather than the web-development comparison’s judging method.
The paper connects a released set of model weights to tests of extended coding and other multi-step work. For a working reader, the distinction is between a model that scores well in the authors’ reported settings and one that has been shown to fit a specific team’s tools, costs and review process.
A possible use is to investigate Kimi K3 for coding tasks that require several rounds of tool use and revision. The reported scores and examples are a starting point for that investigation, not evidence that a particular workplace setup has been tested.
The authors say Kimi K3 still trails Claude Fable 5 and GPT-5.6 Sol overall in their evaluated suite. Some comparisons use different coding harnesses, selected public task subsets, internally maintained tests or third-party scores; the SWE-Marathon tasks used an H20-calibrated branch before its final release. The authors note fallback behavior for Claude Fable 5 and potential cyberguards for GPT-5.6 Sol. Their hardware examples also have specific scope: one kernel-optimization case covers an NVIDIA Hopper GPU and an alternative-vendor GPU, while the reported MiniTriton comparison uses an NVIDIA L20. Those findings should not be read as measurements on every GPU. General PtoP note: this arXiv technical report is a preprint, not evidence of peer review, and a benchmark is not a workplace trial.
FROM PAPER TO PRACTICE
No-install conceptual exercise; it does not run Kimi K3. Prerequisites: a hypothetical coding task, a description of the expected behavior and a test you would want to pass. No model access or paid service is needed for this exercise; access and running costs for the released weights are not established here.
An n8n workflow is not needed for the no-install edit–test–revision exercise; wiring a coding agent into an automation tool would introduce an integration the paper does not test.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS
Why it matters to youIf you build tools that help data scientists test models, this paper describes a possible way to spend less computing time training those tools. The evidence comes from benchmark tasks, not a tested workplace deployment.
The authors report that an agent can be trained with fewer real program runs by predicting most outcomes and checking a small share for real. Imagine an agent writing several programs for a data task. Normally, each program must run in its own isolated workspace before the agent receives a score. The authors instead use a language model to predict most outcomes; they call this predictor a world model. Real runs remain in the training loop to help correct its mistakes. The authors call the full method World Model Reinforcement Learning: reinforcement learning means improving an agent using scores from its attempts.
The authors report that an agent can be trained with fewer real program runs by predicting most outcomes and checking a small share for real. Imagine an agent writing several programs for a data task. Normally, each program must run in its own isolated workspace before the agent receives a score. The authors instead use a language model to predict most outcomes; they call this predictor a world model. Real runs remain in the training loop to help correct its mistakes. The authors call the full method World Model Reinforcement Learning: reinforcement learning means improving an agent using scores from its attempts.
Hypothetical illustration, not a paper result: An agent proposes two ways to classify customer comments. A predictor estimates which program would score better. The team runs a small selection of programs for real, compares the scores, and adjusts how it uses future predictions.
On held-out MLE-Dojo test tasks, the authors report that Qwen3.5-4B-Ours used 286 GPU-hours of training, compared with 883 GPU-hours for the same-scale Qwen3.5-4B-GRPO trained with real execution. A GPU-hour is one hour of use of one graphics processor. The authors report higher average scores for their corrected method than for the same-scale real-execution method on both MLE-Dojo (test) and DSBench, at both the 4B and 9B scales. These scores are leaderboard percentiles: they describe where a submission ranks among entries in its competition, with a higher rank scoring better. Specifically, the authors report that Qwen3.5-4B-Ours outscored Kimi-48B-A3B on both benchmark averages, while Qwen3.5-9B-Ours outscored Nemotron-120B-A12B on both averages. In the separate LIBERO-Long simulated robot-arm test, they report a higher overall task-success rate for MiniVLA-1B-Ours than for MiniVLA-1B-GRPO.
The authors identify a practical mismatch: many proposed solutions can be generated together, but each real program run needs its own workspace and machine time. Their approach changes where most training scores come from while retaining some real runs as a check.
A possible use is training agents for data-science competitions or similar modeling tasks where running every proposed solution takes time and computing resources. The paper tests competition-based benchmarks; it does not establish performance in an organization’s own workflow.
The authors note that older public competitions and winning solutions may have appeared in models’ earlier training material. Their MLE-Dojo split has no identical task on both sides, but one held-out task shares a competition family with a training task; the authors say its data and scoring measure differ. DSBench is described as disjoint from their training set, and the authors evaluated the tasks that ran end to end in their workspace rather than the entire available pool. In the robot-arm test, training used only the first 16 official starting states per task, while evaluation used all 50. The paper’s mathematical guarantee depends on stated assumptions about the scores and prediction errors. PtoP note: this arXiv source is a preprint, not evidence of peer review; benchmark results are not a deployment test.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and runnable code are unverified from the supplied source. Prerequisites: paper and pencil, with no model access, hardware, or cost required. Reproducing the experiments would require access to models, task data, and substantial computing resources; availability and cost are not established here.
An n8n workflow is not appropriate here: the paper studies a specialized agent-training loop with isolated program runs, not an automation workflow or a tested n8n integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS
Why it matters to youIf you build learning tools or plan lessons, you may want to know which guidance might help a particular learner. This paper explores whether a computer-made stand-in for that learner could help test guidance before it is tried with real students.
The authors ask whether a simulated student can do two things at once: make the kinds of mistakes a particular student makes, and change its answer when given guidance. Think of a chess learner who usually chooses one move, then sees a coach’s hint. A useful stand-in would first resemble that learner’s recorded choices and then respond to the hint. The authors trained StudentSim in two stages: first on records pooled from many learners in a subject, then on records from one learner. They tested separate simulators for chess, second-language English writing and mathematics.
The authors ask whether a simulated student can do two things at once: make the kinds of mistakes a particular student makes, and change its answer when given guidance. Think of a chess learner who usually chooses one move, then sees a coach’s hint. A useful stand-in would first resemble that learner’s recorded choices and then respond to the hint. The authors trained StudentSim in two stages: first on records pooled from many learners in a subject, then on records from one learner. They tested separate simulators for chess, second-language English writing and mathematics.
Hypothetical illustration, not a study result: A chess coach has records showing that one player often overlooks a threat to a piece. A simulator might be asked to choose that player’s move on a new board, then choose again after a hint about the threat. The first answer illustrates behavioral fidelity; the second illustrates guidance responsiveness.
On the authors’ held-out chess test, the per-player StudentSim simulators matched their players’ recorded moves with an average score of 0.5150, versus 0.2316 for GPT-5.4 prompted to role-play the players. In everyday terms, about half of recorded moves matched, versus about a quarter for GPT-5.4; these are averages across players, not the score of every player. The authors also report higher guidance-responsiveness scores for StudentSim than for GPT-5.4 in chess, English writing and mathematics. In a separate chess tutor experiment, expert raters scored the tutor trained with a pooled, trained StudentSim reward above both a no-reinforcement-learning tutor and one trained with a GPT-5.4 simulator reward on the study’s three rating measures. That tutor experiment used the pooled simulator, not the individual per-player simulators.
Feedback from real learners is slow and costly to collect, according to the authors. Their approach offers a way to study both a learner-like starting answer and a response to guidance, rather than testing only one of those properties.
A possible research use is to compare candidate tutor messages using simulated responses before seeking feedback from learners. The paper also tests a narrower use: training a chess tutor with feedback from a trained simulator. It does not report a live tutoring deployment.
The guidance-responsiveness test checks whether a simulator reaches a predetermined correction; it does not establish that the real student would make that revision. Chess and mathematics tutor messages were generated under templates, while the English-writing corrections came from teacher annotations. In the chess comparison, StudentSim’s input included a Maia2-derived list of likely moves that was omitted from the GPT-5.4 prompt. The paper says the chess learners used for per-student training were nested within the pooled-training player group, although their evaluation records were held out. The tutor experiment evaluated chess guidance through expert ratings, not learning outcomes from students using the tutor over time; the authors do not run that tutor experiment for English writing or mathematics. This source is an arXiv preprint. As a general PtoP note, a benchmark result is not evidence of deployment performance.
FROM PAPER TO PRACTICE
The official repository provides installation instructions, but these steps have not been tested here. Prerequisites: a local copy of the repository, Python with pip, and the ability to install its chess dependencies. Training requires suitable computing resources; steps that call a language model require Azure OpenAI credentials and may incur costs.
An n8n message-and-automation workflow is not appropriate for reproducing this paper’s model training and held-out evaluation.
Was this explanation easy to understand?
COMPARISON
Compare topics, possible uses and limits. Each result remains attributed to its own paper.
| Paper | Central idea | Possible use | Limit |
|---|---|---|---|
| Chinese-Jev sorts Chinese text into preset answers | The study asks whether a compact model can make fixed-answer decisions in Chinese instead of writing open-ended replies. The authors built Chinese-Jev to read a question… | A possible use is sorting requests where the allowed outcomes are known in advance. The paper measures answer selection on its benchmark… | The authors report remaining accuracy gaps against Jev in law and finance. Confidence calibration—how closely stated probabilities track… |
| Shared Memory Could Help AI Assistants Reuse Visual Lessons | The study asks whether AI assistants can reuse lessons from earlier questions, even when a different system learned them. The authors’ answer is that, in their tests, a… | A possible use is letting teams of AI assistants retain lessons about recurring visual tasks without retraining the answering model. The… | The authors’ construction questions came from the same benchmarks as the evaluation questions, though the project-defined question splits… |
| Check engineering files before trusting an AI agent’s completed work | The study asks whether AI agents can complete engineering work across professional software, not just operate the software. The authors built EngiWorld, a benchmark—a… | As a possible use, a team considering AI assistance for engineering work could use the paper’s distinction between operating tools and… | The main comparison uses a 300-task subset selected to control evaluation cost, so it is not a result for every EngiWorld task. The… |
| Kimi K3 offers another model to consider for complex coding work | The study asks how far an open-weight model can go on long, complicated tasks. The authors built Kimi K3 to work across extended text, code and visual input, then… | A possible use is to investigate Kimi K3 for coding tasks that require several rounds of tool use and revision. The reported scores and… | The authors say Kimi K3 still trails Claude Fable 5 and GPT-5.6 Sol overall in their evaluated suite. Some comparisons use different coding… |
| Predicting experiment outcomes may cut research-agent training time | The authors report that an agent can be trained with fewer real program runs by predicting most outcomes and checking a small share for real. Imagine an agent writing… | A possible use is training agents for data-science competitions or similar modeling tasks where running every proposed solution takes time… | The authors note that older public competitions and winning solutions may have appeared in models’ earlier training material. Their… |
| Student Simulators Could Help Tutors Test Different Guidance | The authors ask whether a simulated student can do two things at once: make the kinds of mistakes a particular student makes, and change its answer when given guidance… | A possible research use is to compare candidate tutor messages using simulated responses before seeking feedback from learners. The paper… | The guidance-responsiveness test checks whether a simulator reaches a predetermined correction; it does not establish that the real student… |
ARCHIVE
Find earlier editions of the newspaper.
30