Never Look Back: Understanding Persistencein 3D Object Memory from Egocentric Videos

Shravan S Chaudhari1·William Paul2·Suchi Saria1·Rama Chellappa1*·Homanga Bharadhwaj1*
what you saw
hover anything
0:00
drag to look aroundhover a dotpause, then hover the video▮ an object moved
Overview

Overview Video Walkthrough

LEDGER (Long-horizon Egocentric Descriptions, Geometry, and Event Records) builds a memory of objects from an egocentric video. It finds the objects in sampled frames, places each one in 3D using the camera pose, and links repeated sightings of the same object over time. Objects stay in memory after they leave view, including ones the wearer never touches. Each object's history is split into rest segments: a move is recorded only when several observations agree, so small localization errors are not counted as moves. Each segment also gets a short description, for example what a container held. Questions are answered later from these records as text, without the video.

Stitched streams

Multiple scenes.
Single memory.

Here, recordings from three different kitchens are joined into one 21-minute stream, and LEDGER builds one memory for the whole stream. The kitchens are not placed in a shared 3D frame, so each appears as its own island. Objects from earlier kitchens stay in memory as the wearer moves on. The questions are the ones from Figure 3 of the paper, answered from this memory without any frames. In our study of 100 such streams, returning to the same scene cost the memory little, but moving to a new scene lowered its accuracy; building one memory per scene recovered part of this loss.

drag the timelinedrag to orbit◆ a question, answered from the memoryislands are not to scale with each other

Tracking across a cut.

Point trackers follow pixels through a continuous video. At a cut, they keep an object only if the new view happens to look like the old one. In this example, AllTracker keeps the projector, shelf and flowers on day 6 but loses the microwave. After a change of place, it lost all 9 objects we tested, so nothing from before the cut can be asked about later. LEDGER keeps the objects in memory across the cut, and here finds all four again on day 6.

another stitched stream: the same shared house on two different days (EgoLife, in UCS-Bench) · faces blurred
from the paper's study of 100 stitched streams (3,244 questions)

See every stage.

Qwen3.5-9B lists the objects in each sampled frame. YOLO-World finds them in the frame (SAM 3 on Ego4D, where localization accuracy is measured). WildDet3D turns each detection into a 3D box, and the camera pose, given by the dataset or estimated with FastVGGT, places it in the scene. Move the lens over the frame to see each stage.

move the lensspace lens on/off←→ frames“Ledger”: hover an object
this frame

Re-identification
through clustering and tracking.

The same object is seen many times, sometimes under different names. LEDGER matches each new sighting to an existing object with the same label, close to where that object was last seen. A second pass uses the SAM 3 video tracker to merge objects that were wrongly split into several tracks. Objects stay in memory after they leave view, though a large move can still create a new entry.

Ask LEDGER anything

Watch LEDGER look things up.

Here, ReMEmbR's question-answering agent (with GPT-5.4) reads LEDGER's memory. It searches the memory by meaning, time or place, reads what comes back, and answers from that text only. LEDGER's own reader, used for the main results, answers in a single step instead. The questions shown are ones the agent answered correctly in at least two of three runs; below each answer are the results of other methods on the same question.

Results

On HD-EPIC, LEDGER's memory helps the answering model more than any of the memories we rebuilt (AMEGO, DirectMe, OSNOM and ReMEmbR), most of all for objects the wearer never touched. On UCS-Bench, reading the memory is better than reading DirectMe's scene graph; sampled frames still do better on their own, but adding the memory to them helps further. On Ego4D VQ3D, it places named objects closer to their true positions than the other memories. The benefit holds as recordings get longer, while changes of scene remain a weakness.

HD-EPIC · accuracy
Ego4D VQ3D · median 3D error (m) · lower is better
UCS-Bench · accuracy vs video length

Citation

@misc{chaudhari2026lookbackunderstandingpersistence,
      title={Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos}, 
      author={Shravan Chaudhari and William Paul and Suchi Saria and Rama Chellappa and Homanga Bharadhwaj},
      year={2026},
      eprint={2610.10538},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.10538}, 
}