
Embodied AI systems learn from continuous, first-person visual streams — long-form video that provides naturalistic spatiotemporal grounding, multimodality, and object motion, but also constant distribution shift. Iconic single-subject image datasets don't prepare models for this regime. We develop methods for embodied representation learning, world-model prediction and planning, unsupervised continual video learning from egocentric streams, and incremental recognition in open-world environments. Our vision is efficient, adaptable, on-device visual learning that grounds perception in real-world action — with applications in embodied planning, robotics, and visual assistance.

CoRR · 2026-07-02
AdaJEPA adapts a latent world model inside closed-loop MPC, using each observed transition as a self-supervised signal before the next replan.

CoRR · 2026-06-03
BabyCL is a continual multimodal learning framework that processes a child's SAYCam egocentric stream in a single chronological pass, jointly learning visual representations and word semantics.

ICML 2026 · 2026-03-12
Inspired by the perceptual straightening hypothesis in human vision, we introduce temporal straightening to improve representation learning for latent planning.

ICLR 2026 · 2025-10-07
Midway Network is a new self-supervised learning architecture that learns strong visual representations for both object recognition and motion understanding solely from natural videos by modeling latent dynamics.

CoRR · 2025-08-21
StreamMem is a query-agnostic KV cache memory mechanism for streaming video understanding.

CoLLAs 2025 · 2025-06-04
We provide a theoretical analysis of sample replay in over-parameterized continual linear regression, and we show that replay can provably increase forgetting in the worst case even though the network has the capacity to memorize all tasks.

CoLLAs 2025 · 2025-01-21
Memory Storyboard groups recent past frames into temporal segments and provides effective summarization of the past visual streams for memory replay.

ICLR 2025 · 2024-08-20
We propose PooDLe, a self-supervised learning method that combines an invariance-based objective on pooled representations with a dense SSL objective that enforces equivariance to optical flow warping.

CoLLAs 2024 · 2024-04-29
We formulate Osiris, a unifying framework for unsupervised continual learning (UCL), which disentangles learning objectives that encompass stability, plasticity, and cross-task consolidation.

CogSci 2024 · 2024-02-01
We train self-supervised video models on longitudinal, egocentric headcam recordings collected from a child over a two year period in their early development.

CoRR · 2023-12-07
LifelongMemory is a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval.