ISSUE 03/2026 · 29-09-2026

A smaller search model for page images, trained without reading training pages

Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s question-reading part while keeping its page index—the stored representations used for search. They…

Editorial illustration of research being examined at a desk.
AI-generated editorial illustration: a visual metaphor, not a figure from the paper.
Read the lead story ↘

6 paper · Original sources linked in every story

From the news desk

News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.

  1. NVIDIA Technical Blog · 28-09-2026NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring ↗Read the explainer ↓
  2. Hugging Face Blog · 28-09-2026Holo4: powering generalist computer-use agents ↗Read the explainer ↓
  3. xAI · 28-09-2026Team Bots: AI coworkers that learn from your team ↗Read the explainer ↓
  4. Google · 24-09-2026Introducing Gemini 3.8 Live with Live Avatar ↗Read the explainer ↓

PODCAST / ENGLISH EDITION

The full issue, then each paper

The front-page episode covers the papers and news; each paper also has its own short conversation.

Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.

EP 01ARXIV:2609.35505

Teaching a smaller AI model with answers it has already written

Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training?…

EP 03ARXIV:2609.34899

A smaller search model for page images, trained without reading training pages

Can a small model search an existing collection of page images without running a large model for every question? The authors test a…

EP 04ARXIV:2609.17488

Can one model learn to predict and fill gaps in tables?

The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry…

EP 05ARXIV:2608.09888

Can a model learn a visual rule from examples without writing out its reasoning?

The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its…

EP 06ARXIV:2609.11638

Vidu S2 makes a case for video that responds while it runs

Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present…

FROM THE NEWS DESK

The news, explained

Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.

NVIDIA Technical Blog · 28-09-2026

NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring

In brief

NVIDIA announced a design for watching and limiting computer helpers that act on their own. The central idea is to keep some safety checks beyond a helper’s reach, including on separate hardware.

An example

Imagine a travel-booking helper that may read your calendar but not your bank files. A separate guard checks its requests and blocks a request for a bank file. This is an illustration, not a reported test.

Application

A company could use NVIDIA’s proposed setup to give an agent limited access to work tools. NVIDIA says its OpenShell runtime puts the agent in a sandbox and enforces a policy; an optional, out-of-band hardware layer can provide another check.

The limitation

This is a company announcement, not measured proof that the design prevents harmful actions. The article does not establish how well it works in everyday use.

Takeaway

NVIDIA argues that agents need independent limits and monitoring, rather than being trusted to police themselves.

Original source ↗

Was this explanation easy to understand?

Hugging Face Blog · 28-09-2026

Holo4: powering generalist computer-use agents

In brief

Hcompany says it has released Holo4, an AI model that can work across computer programs by clicking, typing, writing code, or using software connections. Its central idea is to use one model for tasks that cross different programs and ways of controlling them.

An example

Imagine copying figures from a website into an office record. The model might use the website’s on-screen controls, then send the figures to the record system through a direct software connection. This is an illustration, not a task the post says it tested.

Application

A business could try Holo4 for work that spans an older program with only a GUI and a newer service with an API.

The limitation

This is a company announcement, not independent proof of reliability. Hcompany reports a 61.7% score for Holo4 27B on a desktop benchmark, below the 81.8% it cites for Opus 5.5. Test setups can differ, and benchmark results do not guarantee success at work.

Takeaway

Holo4 is meant to handle several kinds of computer work with one model, but its performance on everyday business tasks remains uncertain.

Original source ↗

Was this explanation easy to understand?

xAI · 28-09-2026

Team Bots: AI coworkers that learn from your team

In brief

xAI says it has launched shared AI helpers that teams can give relevant information and use together. It says Team Bots are available in public beta on Teams and Enterprise plans.

An example

Imagine a sales team giving one helper its account notes. It could prepare a morning briefing for the team. This is an illustration, not a reported result.

Application

One possible use is checking marketing drafts against a team’s writing guidelines. xAI describes this as a use for its Bot.

The limitation

This is a company announcement, not an independent test. The article does not establish how reliably a Bot remembers information or gets answers right.

Takeaway

The central idea is a shared helper that can work from a team’s information and retain what it learns. Its real-world reliability remains uncertain.

Original source ↗

Was this explanation easy to understand?

Google · 24-09-2026

Introducing Gemini 3.8 Live with Live Avatar

In brief

Google announced a version of its conversational assistant that can appear as a moving, speaking character during a live exchange. Google says Gemini 3.8 Live with Live Avatar responds to what it sees and hears, pairing speech with video. This is a company announcement, not an independent test of how well it works.

An example

Imagine, hypothetically, a hotel guest speaking to an on-screen character about check-in. The character could respond aloud while the guest watches it speak.

Application

A business could use the feature for customer service. Google says it is available in Gemini Enterprise.

The limitation

Creating a custom avatar currently requires enterprise allowlisting. The article does not provide independent evidence of performance in everyday use.

Takeaway

Google is offering businesses a more visually present way to use its conversational assistant, but its claims still need real-world scrutiny.

Original source ↗

Was this explanation easy to understand?

01NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS

Teaching a smaller AI model with answers it has already written

Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training? The authors study this question in mathematical reasoning. Their method, Least Square Policy Distillation (LSPD), has the smaller student write answers and asks a larger teacher how likely it would have been to choose the words in them. Training brings the student's choices closer to the teacher's while encouraging a range of possible answers. A version called LSPD-RB also keeps earlier student answers for reuse.

Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, and Weitong Zhang · An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning · arXiv:2609.35505Read the paper ↗

In one sentence

Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training? The authors study this question in mathematical reasoning. Their method, Least Square Policy Distillation (LSPD), has the smaller student write answers and asks a larger teacher how likely it would have been to choose the words in them. Training brings the student's choices closer to the teacher's while encouraging a range of possible answers. A version called LSPD-RB also keeps earlier student answers for reuse.

Key concepts

  • Distillation means training a student model with guidance from a teacher model. Here, the student generates the answers that the teacher helps score.
  • A token is a small piece of text. At each token in a student answer, LSPD compares how likely the student and teacher were to choose it. It reduces the gap, but softens the penalty for unusually large gaps.
  • Entropy measures how spread out a model's possible next-token choices are. LSPD's training objective encourages that spread rather than concentrating every choice on one option.
  • A rollout batch is a newly collected set of student answers; an optimization update is one adjustment to the student using answers already collected. LSPD-RB stores past answers in a replay buffer so later updates can draw from them.

A concrete example

Hypothetical illustration, not a paper result: A student writes a solution to a math problem. For a word in one reasoning step, the teacher assigns a different likelihood than the student did. Training adjusts the student using that difference. When the next batch arrives, the replay-buffer version can also revisit this older solution.

What the researchers measured

In the authors' main table, a Qwen3-1.7B-Base student trained with LSPD-RB using a Qwen3-4B non-thinking teacher scores 36.60% Avg@16 on AMC23 after 10 training steps. The on-policy distillation baseline—which learns from newly collected student answers—scores 34.64% for that same model pair and benchmark; the authors say methods other than LSPD-RB required at least 40 steps to converge stably. Avg@16 averages correctness across 16 generated answers per problem, then across problems. These are headline score comparisons, not counts of model updates. Separately, in the authors' Figure 2 replay-curve experiment with that model pair on AMC23, AIME24 and AIME25, each collection step gathers answers to 64 prompts, with four answers per prompt. For LSPD-RB in that experiment, the authors perform 256 optimization updates after adding each new batch to the replay buffer. They report that its curves reach saturated performance within approximately 10 collection steps. The 256-update detail describes that replay-curve experiment; it should not be inferred from the main table's 36.60% entry. Across the authors' six math benchmarks and three teacher–student pairs, LSPD's Avg@16 exceeds the strongest baseline by 0.91 percentage points on average.

Why it matters

The authors distinguish two costs that can be easy to confuse: collecting new student answers and repeatedly updating a model with answers already collected. Their experiments examine whether reuse reduces the need for new collection. They also test whether a trained student can find a correct answer when allowed several attempts.

Where it might help

Possible use, not a demonstrated deployment: adapting a smaller reasoning model when collecting fresh student answers is costly. The paper tests training methods, not a finished user-facing service.

Impact across sectors

  • Possible education use, not a proven deployment: a developer might investigate whether the method helps train a model for math-practice tools.
  • Possible software-development use, not a proven deployment: a team might investigate this training approach for a smaller code-assistance model. The paper includes a limited code benchmark, not a deployed assistant.

Where the evidence stops

The authors train on a mathematical-reasoning dataset. Their additional tests outside math cover four named benchmarks, rather than every possible task or deployment setting. Their mathematical regret guarantee—a bound on accumulated distance from an ideal policy—applies to an idealized optimistic algorithm, not directly to the practical training recipe; it assumes bounded reward differences, adequate model coverage and a particular unbiased teacher-feedback model. The supplied text does not report an analysis of overlap between training and evaluation questions. General PtoP note: this is an arXiv preprint; the supplied text does not establish peer review.

FROM PAPER TO PRACTICE

How to try it

The official repository offers a command-only preview; this is not a training or score reproduction, and the procedure has not been tested here. Prerequisites: obtain https://github.com/UNCSciML/LSPD and use a shell with Bash and Python 3.12. The preview needs no model download or graphics processor. Full training has a substantial hardware barrier: the README specifies four NVIDIA RTX PRO 6000 graphics processors.

  1. Obtain the repository through its web page and open a shell in its root directory.
  2. Run `bash train_lspd.sh --dry-run`.
  3. Run `bash train_lspd_rb.sh --dry-run`.
  4. Compare the printed launch commands for the ordinary and replay-buffer methods. Expect commands, not trained models or benchmark scores. Installation and actual training are unverified here.

n8n example

An n8n workflow is not appropriate here: the paper studies local model training and benchmark evaluation, not an automation workflow.

Was this explanation easy to understand?

02NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS

Can a model tell what changed when it learned?

Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers adjusted during training—can reveal what it learned or how its behavior changed. They trained an Imprint Reader to describe changes made to Qwen3-14B. They then tested whether signals from that Reader could guide changes to the original model. This is an arXiv preprint, not a peer-reviewed publication.

Guanxu Chen, Qihao Lin, Jing Shao · Imprint Reader: From Weight-Update Readout to Behavioral Intervention · arXiv:2609.35261Read the paper ↗

In one sentence

Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers adjusted during training—can reveal what it learned or how its behavior changed. They trained an Imprint Reader to describe changes made to Qwen3-14B. They then tested whether signals from that Reader could guide changes to the original model. This is an arXiv preprint, not a peer-reviewed publication.

Key concepts

  • A weight update is the difference between a model’s weights before and after a small training task. The authors made separate updates for individual facts or response tendencies.
  • The Imprint Reader is a trained copy of the model. The authors temporarily added each frozen update to it, asked what had changed, and removed the update before training the Reader further.
  • An anchor-free meta-query asks about the change without naming the fact or behavior. Training also included no-change and random-change cases, for which the Reader was taught not to make a specific claim.
  • MetaEdit uses gradients of the Reader’s likelihood for a written target self-report. A gradient is a signal showing which weight changes would affect how likely the Reader is to produce those words. The authors used these signals to select rows of weights to remove or adjust in the original Qwen3-14B.

A concrete example

Hypothetical illustration, not a paper result: a temporary copy learns a new fact about a made-up planet. The Reader sees only the resulting weight update and a question such as “What changed?” It must describe the fact without seeing the training questions. For a behavior change, MetaEdit instead starts with a written target such as “I check my intermediate steps” and uses the Reader’s gradients to choose weights to adjust; it does not need the Reader to generate a successful description first.

What the researchers measured

On 100 held-out knowledge updates and 100 held-out behavior updates, the joint Imprint Reader reached judge-based Pass@100 of 2% and 16%, respectively. Pass@100 counts the share of updates for which at least one of 100 generated descriptions was judged to state the complete target. In a separate safety test on harmful prompts, MetaEdit (Reader) applied to the original Qwen3-14B raised responses classified as refusals from 57.9% to 64.1% when its safety-maintenance target selected 0.5% of eligible weight rows for removal. That classification used fixed refusal-expression patterns. In separate tool-use tests, the signed MetaEdit edit raised Qwen3-14B’s BFCL Overall score from 41.69% to 44.60%; BFCL Overall is the benchmark’s combined score across tool-use cases. The authors also report more backtracking and sub-goal expressions in mathematical reasoning traces. Those are separate experiments, not additional Reader description successes.

Why it matters

The authors connect two questions: can a model describe a change recorded in its weights, and can the signal used to assess a desired description help guide an edit? Their experiments suggest the signal can be useful even when the Reader rarely gives a complete description on its own.

Where it might help

The authors study controlled readout of individual changes and interventions guided by written behavior descriptions. These are research settings, not demonstrated deployments or autonomous self-improvement.

Impact across sectors

  • Possible safety research use, not a proven deployment: study whether targeted weight changes alter measured refusal behavior.
  • Possible education use, not a proven deployment: examine whether an assistant’s mathematical explanations show more checking and intermediate goals.
  • Possible software-tool use, not a proven deployment: explore adjustments to how an assistant chooses and calls supplied tools.

Where the evidence stops

The authors say complete, freely generated descriptions remain infrequent and can miss details. They tested controlled updates tied to one fact or behavior, not recovery of the training examples behind an update. Knowledge items were extracted from existing benchmark materials; the paper does not establish how this approach would fare on arbitrary real-world model updates. The safety figures count pattern-classified refusals on harmful prompts, not overall safety. Benign-prompt garbling measurements are missing for the new signed MetaEdit conditions. The reported task scores are benchmark results, not evidence of deployment. As a general PtoP note, an arXiv preprint should not be taken to imply peer review.

FROM PAPER TO PRACTICE

How to try it

Conceptual exercise only: the supplied text names code and a model, but provides no verifiable link or README commands here. Installation, access conditions, and any download cost are unverified; reproducing the reported Reader training used eight graphics processors.

  1. Prerequisite: use only this article and paper for a reading exercise; no software or account is needed.
  2. Write a made-up fact and a separate made-up response tendency.
  3. For each, imagine a model trained on examples and note that its weight update—not those examples—is what the Reader would receive.
  4. Write a neutral question that gives away neither target, then ask what evidence a complete answer would need.
  5. Compare your standard with Pass@100: observe that one passing answer among 100 tries for an update counts as a pass, even if most answers fail.

n8n example

An n8n workflow is not appropriate here: the reported method depends on access to model weights and their gradients, not a demonstrated workflow integration.

Was this explanation easy to understand?

03NEW / EDITOR PICKS

ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS

A smaller search model for page images, trained without reading training pages

Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s question-reading part while keeping its page index—the stored representations used for search. They train the replacement using only the large model’s representations of questions, not its representations of training pages. The source is an arXiv preprint; its findings should not be presented as peer-reviewed results.

Zhuchenyang Liu, Ziyi Wang, Yao Zhang, Yu Xiao · ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport · arXiv:2609.34899Read the paper ↗

In one sentence

Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s question-reading part while keeping its page index—the stored representations used for search. They train the replacement using only the large model’s representations of questions, not its representations of training pages. The source is an arXiv preprint; its findings should not be presented as peer-reviewed results.

Key concepts

  • A visual document retriever compares a written question directly with page images, without first turning the pages into text.
  • The large teacher represents a question as several token embeddings: numerical descriptions of pieces of text. Its existing page index holds comparable descriptions of page-image pieces.
  • MaxSim is the matching rule: each question piece finds its best-matching page piece, and those matches contribute to the page’s score.
  • The student can split a question into different pieces from the teacher. The authors’ optimal transport method makes soft matches between the two sets of pieces and learns how much weight to give each student piece.
  • The authors prove that a low alignment cost limits the difference between student and teacher page scores. This is a sufficient condition for similar scores, not a promise that every search ranking will agree.

A concrete example

Hypothetical illustration, not a study result: Someone asks, “Which chart shows revenue growth?” One model may treat “revenue growth” as one piece; another may split it. During training, the small model learns to place weight on pieces that stand in for the large model’s pieces. During search, its question pieces are compared with the already indexed page images.

What the researchers measured

The authors report that, on the ViDoRe v3 benchmark, the 149-million-parameter student trained from ColQwen3.5-4.5B scored 55.1 NDCG@5 points against that teacher’s unchanged index; its teacher scored 58.7. NDCG@5 rewards a search system for putting relevant pages near the top of its first five results. For that same student and teacher, the authors measured 87 versus 2,290 milliseconds to encode one question on a single central-processor thread. This timing excludes scoring pages. Across five separately trained teacher–student pairs, the authors report that the students retained about 95% of their own teachers’ benchmark scores. In the controlled objective comparison on two teachers, they describe their page-free method as on par with the strongest page-dependent comparison; they report less cached teacher data read during training.

Why it matters

The authors focus on a repeated cost: a page can be indexed once, but every new question must be encoded. Their approach moves that question-time work to a smaller model while leaving the teacher’s page index in place. Training also avoids reading cached representations of training pages, unlike the score-distillation comparison.

Where it might help

Possible application, not a demonstrated deployment: An organization that has already indexed page images with a compatible large teacher could use its matching small student to encode new search questions. The paper does not show that this removes the cost of storing or searching the page index.

Impact across sectors

  • Possibility, not a proven deployment: A finance-document search tool could use this approach to find relevant page images in an existing index.
  • Possibility, not a proven deployment: A human-resources archive could try it for questions about scanned policy pages.

Where the evidence stops

The authors used one student encoder family and one fixed-seed run per configuration, with no variance reported. They say small differences between page-free methods on the first benchmark cannot be resolved from a single run. Training-objective comparisons cover only ColQwen3.5-4.5B and Tomoro-ColQwen3-8B, not the other three teachers. Training uses questions drawn from question–page pairs, even though this method does not read the paired pages; training on unpaired question logs was not tested. The translated question variants reuse pages from the base training pairs. Only the question encoder is made smaller: the teacher’s page index and its storage remain. The authors say their score bound is loose and sufficient rather than necessary. They also do not report results by language, despite translated training questions and a French evaluation subset.

FROM PAPER TO PRACTICE

How to try it

A local exploration based on the paper’s official repository README, not a procedure verified here. Prerequisites: Python, sentence-transformers version 6.0 or later for the multi-vector example, access to the published model download, and matching teacher page embeddings if you want to score pages. Download or compute costs are not specified; obtaining a teacher page index may be a substantial barrier. Installation of sentence-transformers is not verified by the supplied instructions.

  1. Read the ColNanoVDR guide linked from https://github.com/Ryenhails/NanoVDR and identify its matching teacher; a different teacher’s page index will not do.
  2. Before loading a model, note that `trust_remote_code=True` permits code supplied by the model publisher to run locally. If you are unsure about granting that permission, stop before loading.
  3. If you accept that permission and have the prerequisite library, follow the README’s ColNanoVDR example: load `nanovdr/ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML` with `MultiVectorEncoder(..., trust_remote_code=True)`, then use its shown `encode_query` call on the example revenue-growth question. Observe that it produces a question representation; this alone is not a page-search result.
  4. Only if you already have `page_embeddings` from the matching teacher, use the README’s `query.similarity(q, page_embeddings)` example to obtain page scores. Without those embeddings, stop at step 3; the README snippet does not create them.

n8n example

An n8n workflow is not appropriate to present as a paper feature: the paper and supplied README do not provide a ready-to-connect n8n search service.

Was this explanation easy to understand?

04TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS

Can one model learn to predict and fill gaps in tables?

The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry, without being retrained for each new table. Their LimiX-2 model practices on computer-generated tables. During practice, it predicts both a chosen answer column and hidden entries elsewhere in the table. The authors then test its predictions on public collections of real-world tables and examine whether its internal patterns help identify relationships between columns.

LimiX Team · LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence · arXiv:2609.17488 · 807 HF votes at selectionRead the paper ↗

In one sentence

The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry, without being retrained for each new table. Their LimiX-2 model practices on computer-generated tables. During practice, it predicts both a chosen answer column and hidden entries elsewhere in the table. The authors then test its predictions on public collections of real-world tables and examine whether its internal patterns help identify relationships between columns.

Key concepts

  • A table contains rows of cases and columns of facts. LimiX-2 keeps a separate internal representation for each cell, rather than reducing a whole row to one representation.
  • In-context learning means using examples supplied with a new table without changing the model’s trained settings. LimiX-2 uses rows with known answers as context for rows whose answers it must predict.
  • Masked modeling means hiding entries during training and asking the model to infer them. The authors hide individual cells, columns or blocks of cells, so training is not limited to predicting one answer column.
  • The authors generate training tables from simulated systems in which variables depend on one another. They use these synthetic data to vary relationships, values and missing-entry patterns.
  • A causal skeleton is a map of which variables are directly connected, without arrows showing the direction of influence. The authors test whether LimiX-2’s feature attention—internal scores for how it uses columns—can help recover such a map.

A concrete example

Hypothetical illustration, not a study result: imagine a table of homes with size, age and sale price. Given some rows with prices, the model could predict a price for a new row. If that row’s age were hidden, the training approach also asks it to infer the missing age from the available entries and example rows.

What the researchers measured

On the full 51-dataset TabArena prediction benchmark, the authors report an Elo rating of 1935 for LimiX-2 in its default configuration, versus 1818 for the runner-up, TabFM+. Elo is a rating fitted from models’ pairwise results across datasets; it is not the percentage of answers a model got right. The authors also report the highest overall Elo for LimiX-2 among the compared methods on TALENT and BCCO. In a separate test on six named causal-discovery datasets, they report that skeletons derived from LimiX-2’s feature attention ranked first by their edge-recovery score on all six.

Why it matters

Many table-prediction methods require separate training for each dataset. The authors study whether practice across varied generated tables can instead support several kinds of inference with one model.

Where it might help

Possible uses include predicting a category, predicting a numerical value and suggesting values for missing table entries. These are possible applications, not deployments established by the paper.

Impact across sectors

  • Possibility, not a proven deployment: a healthcare analyst could explore predictions from a table of patient measurements, subject to appropriate data safeguards.
  • Possibility, not a proven deployment: a finance team could explore predictions from a table of account records, subject to its normal checks.
  • Possibility, not a proven deployment: a scientific team could explore missing measurements in a research table.

Where the evidence stops

The authors excluded 12 TALENT datasets with more than 10 answer classes, so that evaluation covers the remaining 288. LimiX-2 was pretrained exclusively on synthetic tables. The causal test recovers connections without their directions and covers six datasets; some comparison methods timed out or did not apply to certain datasets. The authors say their projected results for a larger, two-billion-parameter model are forecasts, not measurements, and that model comparisons cannot isolate architecture from differences in training data and compute. General PtoP note: this source is an arXiv preprint, not evidence of peer review or deployment performance.

FROM PAPER TO PRACTICE

How to try it

The official repository’s README describes a local inference route; the steps below are not reported here as tested. Prerequisites are Python 3.12 or newer, the README’s software dependencies, a compatible computing setup, a LimiX-2 checkpoint and a table in its required layout. Check the linked non-commercial license before use; obtaining compute or access to a suitable machine may involve a cost.

  1. Read the README and license at https://github.com/limix-ldm/LimiX, and locate the LimiX-2 checkpoint linked there. Its installation example names a different clone address, so check the repository address before installing.
  2. Prepare a classification dataset in the README’s layout: a dataset subdirectory containing matching train and test CSV files. Each file needs a header, with the answer column last.
  3. From a local copy of the verified repository, follow the README’s installation instruction `python -m pip install -e .`. Installation and hardware compatibility are unverified here.
  4. Run the README’s `LimiX-infer --help` to inspect the available flags. If your dataset, checkpoint and computing setup are ready, adapt only the paths in its documented classification example: `LimiX-infer --task_type Classification --data_dir /path/to/class202512_527 --model_path /path/to/LimiX-2.ckpt --gpuid 0 --save_name limix_cls`.
  5. Observe whether the command produces a saved classification result for your dataset. This is a local trial, not a reproduction of the paper’s benchmark scores.

n8n example

An n8n workflow is not useful to specify here: the paper does not describe or test an n8n integration, and its reported work is model evaluation rather than workflow automation.

Was this explanation easy to understand?

05TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS

Can a model learn a visual rule from examples without writing out its reasoning?

The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its intermediate steps. Imagine being shown several small colored grids in which a shape moves to a marked square. You then have to move a new shape in the same way. The authors’ system, BDH-CQ, stores information from the examples in a memory that changes as it sees them. It then repeatedly works on an internal representation of the new grid before producing an answer. Its trained parameters do not change while it solves that task.

Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, Richard Zhong · BDH-CQ: In-Context Learning with Recurrent Latent Reasoning · arXiv:2608.09888 · 792 HF votes at selectionRead the paper ↗

In one sentence

The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its intermediate steps. Imagine being shown several small colored grids in which a shape moves to a marked square. You then have to move a new shape in the same way. The authors’ system, BDH-CQ, stores information from the examples in a memory that changes as it sees them. It then repeatedly works on an internal representation of the new grid before producing an answer. Its trained parameters do not change while it solves that task.

Key concepts

  • Learning from examples at use time: The example grids tell BDH-CQ what rule to apply to the next grid, rather than naming the rule in words.
  • Recurrent memory: As BDH-CQ reads each example, it updates a persistent internal state. That state carries information from earlier examples to the new question.
  • Latent reasoning: The system revises an internal representation of its answer instead of printing a sequence of intermediate reasoning steps.
  • Exact whole-task scoring: On the visual puzzle tests, an output grid must match the target exactly. For tasks with several test grids, the whole task counts as solved only when all its test grids are correct.

A concrete example

Hypothetical illustration, not a reported test result: Draw two examples in which a red square is copied to every gray marker. Then draw a new grid with different marker positions. A solver would need to infer the copying rule from the examples and apply it to the new positions.

What the researchers measured

On the 400-task public ARC-AGI-1 evaluation set, the authors report that the 150-million-parameter BDH-CQ system solved 29.5% of tasks when either of up to two ranked answers could count. They calculate a cost of $0.00070 per task from measured hardware time and an assumed rental rate. The authors describe this score-and-cost point as beyond the previously reported benchmark frontier. In separate controlled grid tests, they report that results depended on the operation: some rules transferred across the tested range, while ordering longer sequences and applying some combinations of operations were harder. A separate comparison of STANDARD and MIN effort settings reports costs of $0.00265246 and $0.00088399 per task, respectively; the authors say the score difference in that comparison is statistically unresolved. Those effort-comparison costs come from a separate setting from the headline cost figure.

Why it matters

The authors study two parts of solving a new problem together: learning its rule from examples and spending more computation on the answer. Their colored-grid tests also let them check whether a rule works across several new inputs, not just one.

Where it might help

Possible applications, not demonstrated deployments, include tools that learn a narrowly defined visual transformation from examples. The paper tests colored-grid puzzles; it does not report use in a workplace system.

Impact across sectors

  • Possibility for education: Example-based grid puzzles could help illustrate how a rule learned from a few cases may fail on a harder case; the paper does not test a teaching tool.
  • Possibility for software tools: A future visual-editing assistant might take before-and-after examples as instructions; the paper does not demonstrate such an assistant.
  • Possibility for industrial inspection: Example-led image transformations might be worth studying for inspection workflows; the paper does not evaluate real inspection images.

Where the evidence stops

The authors say the exact memory updates, implementation details and full training recipe remain proprietary. Their training mix includes ConceptARC data, so the ConceptARC test is not a fresh test of previously unexposed material: changing task identifiers and batch composition does not rule out exposure through training or model selection. In the authors’ generated-task analyses, task construction can affect apparent difficulty; they also found at least one task with a target output that contradicted its examples, with unknown prevalence. Some controlled comparisons use only one puzzle family, so the authors do not claim those differences hold across visual operations. The system returns answer grids, not its internal reasoning, so an incorrect grid cannot reveal exactly where its reasoning went wrong. General PtoP context: this source is an arXiv preprint, not a claim of peer review; a puzzle benchmark is not a deployment test.

FROM PAPER TO PRACTICE

How to try it

Safe conceptual exercise; installation and access to BDH-CQ are unverified because the supplied paper gives no verified code or model installation link. Prerequisite: paper and colored pens, or any grid editor. The paper gives no access price for trying the model; this exercise needs no model access.

  1. Draw two small input grids with a shape and a gray marker in different places.
  2. Draw an output for each that copies the shape to the marker.
  3. Draw a third input with the marker somewhere new, but leave its output blank.
  4. Ask someone to infer the rule and fill in the blank grid. Observe whether the same rule gives the exact expected output; this illustrates the task format, not BDH-CQ’s measured performance.

n8n example

A proposed n8n integration is not appropriate to specify here: the paper provides no verified public model interface for an automation workflow.

Was this explanation easy to understand?

06TOP VOTED · 6 MONTHS

ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS

Vidu S2 makes a case for video that responds while it runs

Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present two parts of Vidu S2. Vidu S2-Avatar generates a character that can respond to new instructions and images during a stream. Vidu S2-Editing changes an incoming video stream, such as its clothing or background. The authors also explore turning their output into spatial video: separate views for the left and right eyes that create a sense of depth. This is an arXiv preprint, not a claim of peer review.

Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu · Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation · arXiv:2609.11638 · 701 HF votes at selectionRead the paper ↗

In one sentence

Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present two parts of Vidu S2. Vidu S2-Avatar generates a character that can respond to new instructions and images during a stream. Vidu S2-Editing changes an incoming video stream, such as its clothing or background. The authors also explore turning their output into spatial video: separate views for the left and right eyes that create a sense of depth. This is an arXiv preprint, not a claim of peer review.

Key concepts

  • Streaming generation: Instead of finishing a whole video before showing it, Vidu S2-Avatar makes successive short sections. That lets a later instruction affect what comes next.
  • Changing references: A reference image shows an appearance to use. For Vidu S2-Avatar, the authors describe adding one during a stream—for example, to show an object the character should pick up. Their vision-and-language agent turns instructions and images into prompts and checks generated frames to guide later prompts.
  • Frame-aligned attention: If someone raises an arm in an incoming video, Vidu S2-Editing pairs each source frame with the edited frame for that same moment. This is the authors’ method for keeping the action and timing while changing the requested appearance.
  • Self-Replay Forcing: The models practise on sections they generated themselves, add noise to those sections, and replay them during training so errors in earlier sections can inform learning in later ones. It is a training method for streams that build on their own past output.

A concrete example

Hypothetical illustration, not a reported test: A person starts a character stream with an image, then shows a picture of a blue mug and asks the character to pick it up. Separately, someone could send a camera stream to Vidu S2-Editing with an instruction to change the background. The first task generates a character’s next actions; the second edits frames arriving from an existing video.

What the researchers measured

The authors report that Vidu S2-Avatar generates 720p video at 25–42 frames per second: the range counts newly generated video images each second, not playback speed. On Sparkle-Bench, a public video-editing evaluation, they report an Overall judged-editing-quality score of 3.74 for Vidu S2-Editing, the highest in that table. This is a benchmark score, not a speed measurement. The authors also report leading available results for Vidu S2-Avatar on the reported StreamAV-Bench subset and describe human-preference comparisons on internal tests. Those findings belong to the named model and test settings; the Avatar generation-speed figure is not an Editing speed figure.

Why it matters

Editorial context: A video that responds as it is made calls for a different experience from ordering a clip and waiting for it to finish. The paper addresses both making new frames and changing incoming ones, while treating depth-view video as an exploration rather than a settled deployment.

Where it might help

Possible applications, not proven deployments, include responsive digital characters, live visual effects on incoming video, and depth-view video experiences. The paper describes a playable online demo but does not establish that these uses are deployed in any particular sector.

Impact across sectors

  • Possible entertainment use: an interactive character could respond to a viewer’s new instruction during a stream.
  • Possible video-production use: a live incoming feed could be given a new visual style or background.
  • Possible immersive-media use: separate left- and right-eye streams could be shown in a headset; the authors say resolution and delay remain practical challenges.

Where the evidence stops

The authors identify several bounds on these findings. Their StreamAV-Bench table reports the available Gemini-MLLM subset, with some comparison measurements unavailable; it is not a complete set of results for every system. Commercial-system comparisons use internal benchmarks and paired human judgments. The editing training videos come from videos filtered through the Avatar data pipeline, although the four editing-task training subsets are mutually disjoint; the paper does not state whether training data overlap with evaluation sets. Spatial-video results shown are representative examples, not a quantified headset evaluation. The authors say practical headset use needs higher resolution and lower end-to-end delay, especially when a camera view must stay in step with head movement. General PtoP note: a benchmark result does not by itself establish performance in a deployed service.

FROM PAPER TO PRACTICE

How to try it

The paper links an online demo, not verified installation instructions or code. This is a suggested observation exercise, not a tested procedure. Prerequisites: a browser and, if you choose to supply one, an image you have permission to use. Current access, account requirements, and cost are not stated in the paper.

  1. Visit https://vidu.com/vidu-stream and check whether the paper’s demo is available.
  2. If it is, look for the option to start a character stream with an initial image and a text or spoken instruction, as described in the paper. Do not upload a private image without permission.
  3. If the demo permits it, give a new instruction or reference image while the stream runs.
  4. Observe whether later generated frames respond to that change, and note any visible delay or inconsistency. Do not treat a single demonstration as a benchmark result. Local installation is unverified.

n8n example

An n8n workflow is not appropriate here: the paper provides an online demo but does not document an automation interface or a tested n8n integration.

Was this explanation easy to understand?

COMPARISON

What these studies show together

Compare topics, possible uses and limits. Each result remains attributed to its own paper.

PaperCentral ideaPossible useLimit
Teaching a smaller AI model with answers it has already writtenCan a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training? The authors study this question in…Possible use, not a demonstrated deployment: adapting a smaller reasoning model when collecting fresh student answers is costly. The paper…The authors train on a mathematical-reasoning dataset. Their additional tests outside math cover four named benchmarks, rather than every…
Can a model tell what changed when it learned?Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers adjusted during training—can reveal what…The authors study controlled readout of individual changes and interventions guided by written behavior descriptions. These are research…The authors say complete, freely generated descriptions remain infrequent and can miss details. They tested controlled updates tied to one…
A smaller search model for page images, trained without reading training pagesCan a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s…Possible application, not a demonstrated deployment: An organization that has already indexed page images with a compatible large teacher…The authors used one student encoder family and one fixed-seed run per configuration, with no variance reported. They say small differences…
Can one model learn to predict and fill gaps in tables?The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry, without being retrained for each new…Possible uses include predicting a category, predicting a numerical value and suggesting values for missing table entries. These are…The authors excluded 12 TALENT datasets with more than 10 answer classes, so that evaluation covers the remaining 288. LimiX-2 was…
Can a model learn a visual rule from examples without writing out its reasoning?The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its intermediate steps. Imagine being shown…Possible applications, not demonstrated deployments, include tools that learn a narrowly defined visual transformation from examples. The…The authors say the exact memory updates, implementation details and full training recipe remain proprietary. Their training mix includes…
Vidu S2 makes a case for video that responds while it runsCan computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present two parts of Vidu S2. Vidu…Possible applications, not proven deployments, include responsive digital characters, live visual effects on incoming video, and depth-view…The authors identify several bounds on these findings. Their StreamAV-Bench table reports the available Gemini-MLLM subset, with some…

ARCHIVE

Issue archive

Find earlier editions of the newspaper.

29
AI studies test reasoning, weight readouts, page search, tables and live video29-09-2026
↗
27
Two Texts, a Tutor and Agent Judgment27-09-2026
↗
26
Models and the World26-09-2026
↗
Search every paper by keyword or tag ↗