Abstract
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Community
When a camera turns away and comes back, most video world models forget what was there. LOCI keeps both kinds of memory: half of the blocks retain a KV cache of past observations, the other half add a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, and its readout steers the queries of the cache-backed blocks.
- On the MIND memory benchmark and held-out Unreal Engine trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax baseline.
- With full history it lowers peak memory by ~30% at equal length; with a bounded bank it streams long videos at constant memory.
- We release LOCI-revisit-data: 24 h of game-engine video with exact camera poses, depth and revisit pairs.
🔗 Project page: https://xiaji2021.github.io/LOCI/
📦 Dataset: https://huggingface.co/datasets/sum0214/LOCI-revisit-data
💻 Code: https://github.com/xiaji2021/LOCI
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation (2026)
- Honeycomb: Constant-Size Scene Memory Representation for Video World Models (2026)
- WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory (2026)
- ReWorld: An Interactive World Model with Long-Horizon Memory (2026)
- Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation (2026)
- Alaya-EVOKE: From Linear-Scaling Supervision to Endless World (2026)
- Addressable Memory for Video World Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.40222 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper