app
Learning from Embodied Visual Experience

Embodied AI systems learn from continuous, first-person visual streams — long-form video that provides naturalistic spatiotemporal grounding, multimodality, and object motion, but also constant distribution shift. Iconic single-subject image datasets don't prepare models for this regime. We develop methods for embodied representation learning, world-model prediction and planning, unsupervised continual video learning from egocentric streams, and incremental recognition in open-world environments. Our vision is efficient, adaptable, on-device visual learning that grounds perception in real-world action — with applications in embodied planning, robotics, and visual assistance.

Research Works in the Area

design

AdaJEPA: An Adaptive Latent World Model

CoRR · 2026-07-02

AdaJEPA adapts a latent world model inside closed-loop MPC, using each observed transition as a self-supervised signal before the next replan.

Learn more
design

Continual Visual and Verbal Learning Through a Child's Egocentric Input

CoRR · 2026-06-03

BabyCL is a continual multimodal learning framework that processes a child's SAYCam egocentric stream in a single chronological pass, jointly learning visual representations and word semantics.

Learn more
design

Temporal Straightening for Latent Planning

ICML 2026 · 2026-03-12

Inspired by the perceptual straightening hypothesis in human vision, we introduce temporal straightening to improve representation learning for latent planning.

Learn more
design

Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics

ICLR 2026 · 2025-10-07

Midway Network is a new self-supervised learning architecture that learns strong visual representations for both object recognition and motion understanding solely from natural videos by modeling latent dynamics.

Learn more
design

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

CoRR · 2025-08-21

StreamMem is a query-agnostic KV cache memory mechanism for streaming video understanding.

Learn more
design

Replay Can Provably Increase Forgetting

CoLLAs 2025 · 2025-06-04

We provide a theoretical analysis of sample replay in over-parameterized continual linear regression, and we show that replay can provably increase forgetting in the worst case even though the network has the capacity to memorize all tasks.

Learn more
design

Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos

CoLLAs 2025 · 2025-01-21

Memory Storyboard groups recent past frames into temporal segments and provides effective summarization of the past visual streams for memory replay.

Learn more
design

PooDLe: Pooled and Dense Self-Supervised Learning from Naturalistic Videos

ICLR 2025 · 2024-08-20

We propose PooDLe, a self-supervised learning method that combines an invariance-based objective on pooled representations with a dense SSL objective that enforces equivariance to optical flow warping.

Learn more
design

Integrating Present and Past in Unsupervised Continual Learning

CoLLAs 2024 · 2024-04-29

We formulate Osiris, a unifying framework for unsupervised continual learning (UCL), which disentangles learning objectives that encompass stability, plasticity, and cross-task consolidation.

Learn more
design

Self-Supervised Learning of Video Representations from a Child's Perspective

CogSci 2024 · 2024-02-01

We train self-supervised video models on longitudinal, egocentric headcam recordings collected from a child over a two year period in their early development.

Learn more
design

LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos

CoRR · 2023-12-07

LifelongMemory is a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval.

Learn more