Teaching a smaller AI model with answers it has already written
Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training?…
ISSUE 03/2026 · 29-09-2026
Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s question-reading part while keeping its page index—the stored representations used for search. They…

6 paper · Original sources linked in every story
News and analysis from the sources, separate from the papers. Headlines and dates link to the originals.
THIS ISSUE AT A GLANCE
↗02 / NEW / EDITOR PICKSCan a model tell what changed when it learned?Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language…
↗03 / NEW / EDITOR PICKSA smaller search model for page images, trained without reading training pagesCan a small model search an existing collection of page images without running a large model for every…
↗04 / TOP VOTED · 6 MONTHSCan one model learn to predict and fill gaps in tables?The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an…
↗05 / TOP VOTED · 6 MONTHSCan a model learn a visual rule from examples without writing out its reasoning?The authors ask whether a model can pick up a new rule from a few examples and work out an answer…
↗06 / TOP VOTED · 6 MONTHSVidu S2 makes a case for video that responds while it runsCan computer-made video respond to a person while it is playing, rather than making them wait for a…
↗PODCAST / ENGLISH EDITION
The front-page episode covers the papers and news; each paper also has its own short conversation.
Both voices are AI-generated with OpenAI. This is a scripted editorial conversation, not a recording of the authors.
Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training?…
Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers…
Can a small model search an existing collection of page images without running a large model for every question? The authors test a…
The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry…
The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its…
Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present…
FROM THE NEWS DESK
Editorial explanations of the sources: examples and possible uses are distinguished from what the source establishes.
NVIDIA Technical Blog · 28-09-2026
NVIDIA announced a design for watching and limiting computer helpers that act on their own. The central idea is to keep some safety checks beyond a helper’s reach, including on separate hardware.
Imagine a travel-booking helper that may read your calendar but not your bank files. A separate guard checks its requests and blocks a request for a bank file. This is an illustration, not a reported test.
A company could use NVIDIA’s proposed setup to give an agent limited access to work tools. NVIDIA says its OpenShell runtime puts the agent in a sandbox and enforces a policy; an optional, out-of-band hardware layer can provide another check.
This is a company announcement, not measured proof that the design prevents harmful actions. The article does not establish how well it works in everyday use.
NVIDIA argues that agents need independent limits and monitoring, rather than being trusted to police themselves.
Was this explanation easy to understand?
Hugging Face Blog · 28-09-2026
Hcompany says it has released Holo4, an AI model that can work across computer programs by clicking, typing, writing code, or using software connections. Its central idea is to use one model for tasks that cross different programs and ways of controlling them.
Imagine copying figures from a website into an office record. The model might use the website’s on-screen controls, then send the figures to the record system through a direct software connection. This is an illustration, not a task the post says it tested.
A business could try Holo4 for work that spans an older program with only a GUI and a newer service with an API.
This is a company announcement, not independent proof of reliability. Hcompany reports a 61.7% score for Holo4 27B on a desktop benchmark, below the 81.8% it cites for Opus 5.5. Test setups can differ, and benchmark results do not guarantee success at work.
Holo4 is meant to handle several kinds of computer work with one model, but its performance on everyday business tasks remains uncertain.
Was this explanation easy to understand?
xAI · 28-09-2026
xAI says it has launched shared AI helpers that teams can give relevant information and use together. It says Team Bots are available in public beta on Teams and Enterprise plans.
Imagine a sales team giving one helper its account notes. It could prepare a morning briefing for the team. This is an illustration, not a reported result.
One possible use is checking marketing drafts against a team’s writing guidelines. xAI describes this as a use for its Bot.
This is a company announcement, not an independent test. The article does not establish how reliably a Bot remembers information or gets answers right.
The central idea is a shared helper that can work from a team’s information and retain what it learns. Its real-world reliability remains uncertain.
Was this explanation easy to understand?
Google · 24-09-2026
Google announced a version of its conversational assistant that can appear as a moving, speaking character during a live exchange. Google says Gemini 3.8 Live with Live Avatar responds to what it sees and hears, pairing speech with video. This is a company announcement, not an independent test of how well it works.
Imagine, hypothetically, a hotel guest speaking to an on-screen character about check-in. The character could respond aloud while the guest watches it speak.
A business could use the feature for customer service. Google says it is available in Gemini Enterprise.
Creating a custom avatar currently requires enterprise allowlisting. The article does not provide independent evidence of performance in everyday use.
Google is offering businesses a more visually present way to use its conversational assistant, but its claims still need real-world scrutiny.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 01 · NEW / EDITOR PICKS
Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training? The authors study this question in mathematical reasoning. Their method, Least Square Policy Distillation (LSPD), has the smaller student write answers and asks a larger teacher how likely it would have been to choose the words in them. Training brings the student's choices closer to the teacher's while encouraging a range of possible answers. A version called LSPD-RB also keeps earlier student answers for reuse.
Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training? The authors study this question in mathematical reasoning. Their method, Least Square Policy Distillation (LSPD), has the smaller student write answers and asks a larger teacher how likely it would have been to choose the words in them. Training brings the student's choices closer to the teacher's while encouraging a range of possible answers. A version called LSPD-RB also keeps earlier student answers for reuse.
Hypothetical illustration, not a paper result: A student writes a solution to a math problem. For a word in one reasoning step, the teacher assigns a different likelihood than the student did. Training adjusts the student using that difference. When the next batch arrives, the replay-buffer version can also revisit this older solution.
In the authors' main table, a Qwen3-1.7B-Base student trained with LSPD-RB using a Qwen3-4B non-thinking teacher scores 36.60% Avg@16 on AMC23 after 10 training steps. The on-policy distillation baseline—which learns from newly collected student answers—scores 34.64% for that same model pair and benchmark; the authors say methods other than LSPD-RB required at least 40 steps to converge stably. Avg@16 averages correctness across 16 generated answers per problem, then across problems. These are headline score comparisons, not counts of model updates. Separately, in the authors' Figure 2 replay-curve experiment with that model pair on AMC23, AIME24 and AIME25, each collection step gathers answers to 64 prompts, with four answers per prompt. For LSPD-RB in that experiment, the authors perform 256 optimization updates after adding each new batch to the replay buffer. They report that its curves reach saturated performance within approximately 10 collection steps. The 256-update detail describes that replay-curve experiment; it should not be inferred from the main table's 36.60% entry. Across the authors' six math benchmarks and three teacher–student pairs, LSPD's Avg@16 exceeds the strongest baseline by 0.91 percentage points on average.
The authors distinguish two costs that can be easy to confuse: collecting new student answers and repeatedly updating a model with answers already collected. Their experiments examine whether reuse reduces the need for new collection. They also test whether a trained student can find a correct answer when allowed several attempts.
Possible use, not a demonstrated deployment: adapting a smaller reasoning model when collecting fresh student answers is costly. The paper tests training methods, not a finished user-facing service.
The authors train on a mathematical-reasoning dataset. Their additional tests outside math cover four named benchmarks, rather than every possible task or deployment setting. Their mathematical regret guarantee—a bound on accumulated distance from an ideal policy—applies to an idealized optimistic algorithm, not directly to the practical training recipe; it assumes bounded reward differences, adequate model coverage and a particular unbiased teacher-feedback model. The supplied text does not report an analysis of overlap between training and evaluation questions. General PtoP note: this is an arXiv preprint; the supplied text does not establish peer review.
FROM PAPER TO PRACTICE
The official repository offers a command-only preview; this is not a training or score reproduction, and the procedure has not been tested here. Prerequisites: obtain https://github.com/UNCSciML/LSPD and use a shell with Bash and Python 3.12. The preview needs no model download or graphics processor. Full training has a substantial hardware barrier: the README specifies four NVIDIA RTX PRO 6000 graphics processors.
An n8n workflow is not appropriate here: the paper studies local model training and benchmark evaluation, not an automation workflow.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 02 · NEW / EDITOR PICKS
Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers adjusted during training—can reveal what it learned or how its behavior changed. They trained an Imprint Reader to describe changes made to Qwen3-14B. They then tested whether signals from that Reader could guide changes to the original model. This is an arXiv preprint, not a peer-reviewed publication.
Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers adjusted during training—can reveal what it learned or how its behavior changed. They trained an Imprint Reader to describe changes made to Qwen3-14B. They then tested whether signals from that Reader could guide changes to the original model. This is an arXiv preprint, not a peer-reviewed publication.
Hypothetical illustration, not a paper result: a temporary copy learns a new fact about a made-up planet. The Reader sees only the resulting weight update and a question such as “What changed?” It must describe the fact without seeing the training questions. For a behavior change, MetaEdit instead starts with a written target such as “I check my intermediate steps” and uses the Reader’s gradients to choose weights to adjust; it does not need the Reader to generate a successful description first.
On 100 held-out knowledge updates and 100 held-out behavior updates, the joint Imprint Reader reached judge-based Pass@100 of 2% and 16%, respectively. Pass@100 counts the share of updates for which at least one of 100 generated descriptions was judged to state the complete target. In a separate safety test on harmful prompts, MetaEdit (Reader) applied to the original Qwen3-14B raised responses classified as refusals from 57.9% to 64.1% when its safety-maintenance target selected 0.5% of eligible weight rows for removal. That classification used fixed refusal-expression patterns. In separate tool-use tests, the signed MetaEdit edit raised Qwen3-14B’s BFCL Overall score from 41.69% to 44.60%; BFCL Overall is the benchmark’s combined score across tool-use cases. The authors also report more backtracking and sub-goal expressions in mathematical reasoning traces. Those are separate experiments, not additional Reader description successes.
The authors connect two questions: can a model describe a change recorded in its weights, and can the signal used to assess a desired description help guide an edit? Their experiments suggest the signal can be useful even when the Reader rarely gives a complete description on its own.
The authors study controlled readout of individual changes and interventions guided by written behavior descriptions. These are research settings, not demonstrated deployments or autonomous self-improvement.
The authors say complete, freely generated descriptions remain infrequent and can miss details. They tested controlled updates tied to one fact or behavior, not recovery of the training examples behind an update. Knowledge items were extracted from existing benchmark materials; the paper does not establish how this approach would fare on arbitrary real-world model updates. The safety figures count pattern-classified refusals on harmful prompts, not overall safety. Benign-prompt garbling measurements are missing for the new signed MetaEdit conditions. The reported task scores are benchmark results, not evidence of deployment. As a general PtoP note, an arXiv preprint should not be taken to imply peer review.
FROM PAPER TO PRACTICE
Conceptual exercise only: the supplied text names code and a model, but provides no verifiable link or README commands here. Installation, access conditions, and any download cost are unverified; reproducing the reported Reader training used eight graphics processors.
An n8n workflow is not appropriate here: the reported method depends on access to model weights and their gradients, not a demonstrated workflow integration.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 03 · NEW / EDITOR PICKS
Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s question-reading part while keeping its page index—the stored representations used for search. They train the replacement using only the large model’s representations of questions, not its representations of training pages. The source is an arXiv preprint; its findings should not be presented as peer-reviewed results.
Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s question-reading part while keeping its page index—the stored representations used for search. They train the replacement using only the large model’s representations of questions, not its representations of training pages. The source is an arXiv preprint; its findings should not be presented as peer-reviewed results.
Hypothetical illustration, not a study result: Someone asks, “Which chart shows revenue growth?” One model may treat “revenue growth” as one piece; another may split it. During training, the small model learns to place weight on pieces that stand in for the large model’s pieces. During search, its question pieces are compared with the already indexed page images.
The authors report that, on the ViDoRe v3 benchmark, the 149-million-parameter student trained from ColQwen3.5-4.5B scored 55.1 NDCG@5 points against that teacher’s unchanged index; its teacher scored 58.7. NDCG@5 rewards a search system for putting relevant pages near the top of its first five results. For that same student and teacher, the authors measured 87 versus 2,290 milliseconds to encode one question on a single central-processor thread. This timing excludes scoring pages. Across five separately trained teacher–student pairs, the authors report that the students retained about 95% of their own teachers’ benchmark scores. In the controlled objective comparison on two teachers, they describe their page-free method as on par with the strongest page-dependent comparison; they report less cached teacher data read during training.
The authors focus on a repeated cost: a page can be indexed once, but every new question must be encoded. Their approach moves that question-time work to a smaller model while leaving the teacher’s page index in place. Training also avoids reading cached representations of training pages, unlike the score-distillation comparison.
Possible application, not a demonstrated deployment: An organization that has already indexed page images with a compatible large teacher could use its matching small student to encode new search questions. The paper does not show that this removes the cost of storing or searching the page index.
The authors used one student encoder family and one fixed-seed run per configuration, with no variance reported. They say small differences between page-free methods on the first benchmark cannot be resolved from a single run. Training-objective comparisons cover only ColQwen3.5-4.5B and Tomoro-ColQwen3-8B, not the other three teachers. Training uses questions drawn from question–page pairs, even though this method does not read the paired pages; training on unpaired question logs was not tested. The translated question variants reuse pages from the base training pairs. Only the question encoder is made smaller: the teacher’s page index and its storage remain. The authors say their score bound is loose and sufficient rather than necessary. They also do not report results by language, despite translated training questions and a French evaluation subset.
FROM PAPER TO PRACTICE
A local exploration based on the paper’s official repository README, not a procedure verified here. Prerequisites: Python, sentence-transformers version 6.0 or later for the multi-vector example, access to the published model download, and matching teacher page embeddings if you want to score pages. Download or compute costs are not specified; obtaining a teacher page index may be a substantial barrier. Installation of sentence-transformers is not verified by the supplied instructions.
An n8n workflow is not appropriate to present as a paper feature: the paper and supplied README do not provide a ready-to-connect n8n search service.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 04 · TOP VOTED · 6 MONTHS
The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry, without being retrained for each new table. Their LimiX-2 model practices on computer-generated tables. During practice, it predicts both a chosen answer column and hidden entries elsewhere in the table. The authors then test its predictions on public collections of real-world tables and examine whether its internal patterns help identify relationships between columns.
The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry, without being retrained for each new table. Their LimiX-2 model practices on computer-generated tables. During practice, it predicts both a chosen answer column and hidden entries elsewhere in the table. The authors then test its predictions on public collections of real-world tables and examine whether its internal patterns help identify relationships between columns.
Hypothetical illustration, not a study result: imagine a table of homes with size, age and sale price. Given some rows with prices, the model could predict a price for a new row. If that row’s age were hidden, the training approach also asks it to infer the missing age from the available entries and example rows.
On the full 51-dataset TabArena prediction benchmark, the authors report an Elo rating of 1935 for LimiX-2 in its default configuration, versus 1818 for the runner-up, TabFM+. Elo is a rating fitted from models’ pairwise results across datasets; it is not the percentage of answers a model got right. The authors also report the highest overall Elo for LimiX-2 among the compared methods on TALENT and BCCO. In a separate test on six named causal-discovery datasets, they report that skeletons derived from LimiX-2’s feature attention ranked first by their edge-recovery score on all six.
Many table-prediction methods require separate training for each dataset. The authors study whether practice across varied generated tables can instead support several kinds of inference with one model.
Possible uses include predicting a category, predicting a numerical value and suggesting values for missing table entries. These are possible applications, not deployments established by the paper.
The authors excluded 12 TALENT datasets with more than 10 answer classes, so that evaluation covers the remaining 288. LimiX-2 was pretrained exclusively on synthetic tables. The causal test recovers connections without their directions and covers six datasets; some comparison methods timed out or did not apply to certain datasets. The authors say their projected results for a larger, two-billion-parameter model are forecasts, not measurements, and that model comparisons cannot isolate architecture from differences in training data and compute. General PtoP note: this source is an arXiv preprint, not evidence of peer review or deployment performance.
FROM PAPER TO PRACTICE
The official repository’s README describes a local inference route; the steps below are not reported here as tested. Prerequisites are Python 3.12 or newer, the README’s software dependencies, a compatible computing setup, a LimiX-2 checkpoint and a table in its required layout. Check the linked non-commercial license before use; obtaining compute or access to a suitable machine may involve a cost.
An n8n workflow is not useful to specify here: the paper does not describe or test an n8n integration, and its reported work is model evaluation rather than workflow automation.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 05 · TOP VOTED · 6 MONTHS
The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its intermediate steps. Imagine being shown several small colored grids in which a shape moves to a marked square. You then have to move a new shape in the same way. The authors’ system, BDH-CQ, stores information from the examples in a memory that changes as it sees them. It then repeatedly works on an internal representation of the new grid before producing an answer. Its trained parameters do not change while it solves that task.
The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its intermediate steps. Imagine being shown several small colored grids in which a shape moves to a marked square. You then have to move a new shape in the same way. The authors’ system, BDH-CQ, stores information from the examples in a memory that changes as it sees them. It then repeatedly works on an internal representation of the new grid before producing an answer. Its trained parameters do not change while it solves that task.
Hypothetical illustration, not a reported test result: Draw two examples in which a red square is copied to every gray marker. Then draw a new grid with different marker positions. A solver would need to infer the copying rule from the examples and apply it to the new positions.
On the 400-task public ARC-AGI-1 evaluation set, the authors report that the 150-million-parameter BDH-CQ system solved 29.5% of tasks when either of up to two ranked answers could count. They calculate a cost of $0.00070 per task from measured hardware time and an assumed rental rate. The authors describe this score-and-cost point as beyond the previously reported benchmark frontier. In separate controlled grid tests, they report that results depended on the operation: some rules transferred across the tested range, while ordering longer sequences and applying some combinations of operations were harder. A separate comparison of STANDARD and MIN effort settings reports costs of $0.00265246 and $0.00088399 per task, respectively; the authors say the score difference in that comparison is statistically unresolved. Those effort-comparison costs come from a separate setting from the headline cost figure.
The authors study two parts of solving a new problem together: learning its rule from examples and spending more computation on the answer. Their colored-grid tests also let them check whether a rule works across several new inputs, not just one.
Possible applications, not demonstrated deployments, include tools that learn a narrowly defined visual transformation from examples. The paper tests colored-grid puzzles; it does not report use in a workplace system.
The authors say the exact memory updates, implementation details and full training recipe remain proprietary. Their training mix includes ConceptARC data, so the ConceptARC test is not a fresh test of previously unexposed material: changing task identifiers and batch composition does not rule out exposure through training or model selection. In the authors’ generated-task analyses, task construction can affect apparent difficulty; they also found at least one task with a target output that contradicted its examples, with unknown prevalence. Some controlled comparisons use only one puzzle family, so the authors do not claim those differences hold across visual operations. The system returns answer grids, not its internal reasoning, so an incorrect grid cannot reveal exactly where its reasoning went wrong. General PtoP context: this source is an arXiv preprint, not a claim of peer review; a puzzle benchmark is not a deployment test.
FROM PAPER TO PRACTICE
Safe conceptual exercise; installation and access to BDH-CQ are unverified because the supplied paper gives no verified code or model installation link. Prerequisite: paper and colored pens, or any grid editor. The paper gives no access price for trying the model; this exercise needs no model access.
A proposed n8n integration is not appropriate to specify here: the paper provides no verified public model interface for an automation workflow.
Was this explanation easy to understand?
ENGLISH EDITION / PAPER 06 · TOP VOTED · 6 MONTHS
Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present two parts of Vidu S2. Vidu S2-Avatar generates a character that can respond to new instructions and images during a stream. Vidu S2-Editing changes an incoming video stream, such as its clothing or background. The authors also explore turning their output into spatial video: separate views for the left and right eyes that create a sense of depth. This is an arXiv preprint, not a claim of peer review.
Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present two parts of Vidu S2. Vidu S2-Avatar generates a character that can respond to new instructions and images during a stream. Vidu S2-Editing changes an incoming video stream, such as its clothing or background. The authors also explore turning their output into spatial video: separate views for the left and right eyes that create a sense of depth. This is an arXiv preprint, not a claim of peer review.
Hypothetical illustration, not a reported test: A person starts a character stream with an image, then shows a picture of a blue mug and asks the character to pick it up. Separately, someone could send a camera stream to Vidu S2-Editing with an instruction to change the background. The first task generates a character’s next actions; the second edits frames arriving from an existing video.
The authors report that Vidu S2-Avatar generates 720p video at 25–42 frames per second: the range counts newly generated video images each second, not playback speed. On Sparkle-Bench, a public video-editing evaluation, they report an Overall judged-editing-quality score of 3.74 for Vidu S2-Editing, the highest in that table. This is a benchmark score, not a speed measurement. The authors also report leading available results for Vidu S2-Avatar on the reported StreamAV-Bench subset and describe human-preference comparisons on internal tests. Those findings belong to the named model and test settings; the Avatar generation-speed figure is not an Editing speed figure.
Editorial context: A video that responds as it is made calls for a different experience from ordering a clip and waiting for it to finish. The paper addresses both making new frames and changing incoming ones, while treating depth-view video as an exploration rather than a settled deployment.
Possible applications, not proven deployments, include responsive digital characters, live visual effects on incoming video, and depth-view video experiences. The paper describes a playable online demo but does not establish that these uses are deployed in any particular sector.
The authors identify several bounds on these findings. Their StreamAV-Bench table reports the available Gemini-MLLM subset, with some comparison measurements unavailable; it is not a complete set of results for every system. Commercial-system comparisons use internal benchmarks and paired human judgments. The editing training videos come from videos filtered through the Avatar data pipeline, although the four editing-task training subsets are mutually disjoint; the paper does not state whether training data overlap with evaluation sets. Spatial-video results shown are representative examples, not a quantified headset evaluation. The authors say practical headset use needs higher resolution and lower end-to-end delay, especially when a camera view must stay in step with head movement. General PtoP note: a benchmark result does not by itself establish performance in a deployed service.
FROM PAPER TO PRACTICE
The paper links an online demo, not verified installation instructions or code. This is a suggested observation exercise, not a tested procedure. Prerequisites: a browser and, if you choose to supply one, an image you have permission to use. Current access, account requirements, and cost are not stated in the paper.
An n8n workflow is not appropriate here: the paper provides an online demo but does not document an automation interface or a tested n8n integration.
Was this explanation easy to understand?
COMPARISON
Compare topics, possible uses and limits. Each result remains attributed to its own paper.
| Paper | Central idea | Possible use | Limit |
|---|---|---|---|
| Teaching a smaller AI model with answers it has already written | Can a smaller model learn from a larger one while reusing its earlier answers, rather than continually writing new ones for training? The authors study this question in… | Possible use, not a demonstrated deployment: adapting a smaller reasoning model when collecting fresh student answers is costly. The paper… | The authors train on a mathematical-reasoning dataset. Their additional tests outside math cover four named benchmarks, rather than every… |
| Can a model tell what changed when it learned? | Sometimes, but not reliably yet, according to the authors. They ask whether a change to a language model’s weights—the numbers adjusted during training—can reveal what… | The authors study controlled readout of individual changes and interventions guided by written behavior descriptions. These are research… | The authors say complete, freely generated descriptions remain infrequent and can miss details. They tested controlled updates tied to one… |
| A smaller search model for page images, trained without reading training pages | Can a small model search an existing collection of page images without running a large model for every question? The authors test a way to replace the large model’s… | Possible application, not a demonstrated deployment: An organization that has already indexed page images with a compatible large teacher… | The authors used one student encoder family and one fixed-seed run per configuration, with no variance reported. They say small differences… |
| Can one model learn to predict and fill gaps in tables? | The LimiX Team asks whether one pretrained model can use the rows it sees in a table to predict an answer or infer a missing entry, without being retrained for each new… | Possible uses include predicting a category, predicting a numerical value and suggesting values for missing table entries. These are… | The authors excluded 12 TALENT datasets with more than 10 answer classes, so that evaluation covers the remaining 288. LimiX-2 was… |
| Can a model learn a visual rule from examples without writing out its reasoning? | The authors ask whether a model can pick up a new rule from a few examples and work out an answer without spelling out its intermediate steps. Imagine being shown… | Possible applications, not demonstrated deployments, include tools that learn a narrowly defined visual transformation from examples. The… | The authors say the exact memory updates, implementation details and full training recipe remain proprietary. Their training mix includes… |
| Vidu S2 makes a case for video that responds while it runs | Can computer-made video respond to a person while it is playing, rather than making them wait for a finished clip? The authors present two parts of Vidu S2. Vidu… | Possible applications, not proven deployments, include responsive digital characters, live visual effects on incoming video, and depth-view… | The authors identify several bounds on these findings. Their StreamAV-Bench table reports the available Gemini-MLLM subset, with some… |
ARCHIVE
Find earlier editions of the newspaper.
29