HOST: So what changes if video is made while I'm watching it? EXPERT: So you can actually give a new instruction before the stream ends. The authors built V2S to Avatar to generate a character's next sections in response. HOST: So what might that look like? EXPERT: As a hypothetical example, you show it a picture of a mug and ask the character to pick it up. The paper describes new reference images arriving during a stream. HOST: And the editing part, is it making that character too? EXPERT: No, Video S2 editing changes an incoming video. It pairs each arriving frame with the edited frame for the same moment, and it tries to keep the original action and timing. HOST: Do the results tell us how fast both parts generate video? EXPERT: So the paper reports 25 to 42 new frames each second at 720p for Avatar's real-time generation speed. That's not an editing speed or just a playback rate. On the SparkleBench editing test, editing gets an overall judged quality score of 3.74. HOST: So what should I keep in mind about those tests? EXPERT: An Avatar benchmark table covers an available subset, and commercial comparisons use internal tests. The authors also say headset video still needs better resolution and lower delay. The practical takeaway is to keep the two models' results separate and treat the spatial video work as an exploration.