Wenzhao Zheng

Wenzhao Zheng

Postdoctoral Fellow, UC Berkeley · BAIR

World Models and Agents for
Reliable Physical Intelligence

LLMs provide an important foundation for general intelligence, but realizing Physical Superintelligence (PSI) in the real world requires a complete physical-world harness system around them. Building on the cognitive, reasoning, and decision-making capabilities of LLMs, we connect spatial intelligence, world modeling, and embodied intelligence into a system of continuous perception, prediction, action, and feedback, and study efficiency and safety issues throughout the system.

Physical Superintelligence (PSI) system concept: shared spatial intelligence on the left connects bidirectionally to three levels on the right: an LLM/VLM cognitive core, a world model, and embodied intelligence. Reciprocal links connect the levels, including a direct cognition-to-action path. AI safety and reliability surround the system, which interacts with the physical environment through observation, action, and feedback.
Shared foundations Large Models AI Safety

I'm actively looking for interns and collaborators who are passionate in this direction. Feel free to drop me an email!

This page only shows work where I am a first/co-first author, Project Leader, or corresponding author.

Spatial Intelligence

Turn the real world into digital twins for PSI's general-purpose understanding and interactive training

Spatial intelligence provides the scene understanding and simulation data that Physical Superintelligence (PSI) needs. We recover geometry, semantics, and motion from images and video, and study efficient, complete representations of the 3D world. Our past work spans semantic occupancy prediction, feed-forward Gaussian prediction, and streaming reconstruction. We now focus on:

  • Continuous scene geometry and surface reconstruction
  • Structured motion and 4D scene understanding
  • Generative geometry foundation models for simulation
StreamVGGT scene reconstructions and comparisons of inference time and memory
StreamVGGT: streaming scene reconstructionCached memory integrates new observations into the scene. Reconstruction results and efficiency comparisons are shown here; the project provides an interactive demo.
Reconstruction results · July 2025StreamVGGT · Interactive demo
Related work & resources

Research timeline

Choosing a scene representation

Our view is not a simple progression in which one representation replaces another. Depth, occupancy, points, Gaussians, and implicit fields offer different trade-offs. We want to retain the geometry, motion, and uncertainty that an agent needs to predict and act.

  1. Explored in our workRepresent occupied space
  2. Explored in our workMaintain geometry and motion
  3. Open research goalBuild generative geometry for simulation

Open question How can a representation capture continuous surfaces and structured motion while remaining efficient enough for simulation and online interaction?

  1. 2023

    3D occupancy representations

    TPVFormer maps surround-camera images to a tri-perspective representation and semantic occupancy.
    3D semantic occupancy

    Make scene geometry and semantics explicit, while reducing supervision costs.

  2. 2024

    Gaussian scene representations

    S3Gaussian combines spatial-temporal features with Gaussians to reconstruct streets, render novel views, and separate static and dynamic content.
    Self-supervised dynamic street reconstruction

    Represent scenes with sparse Gaussians and extend them across time.

    More work · 1
  3. 2025

    Streaming scene reconstruction

    StreamVGGT incrementally reconstructs an indoor scene using streaming images and cached memory.
    Online reconstruction

    Integrate new observations without reconstructing the scene from scratch.

    More work · 1
  4. 2026

    Complete scene understanding

    SM4RT reconstruction of a moving swing with colored motion trajectories.
    Structured 4D motion

    Connect continuous surfaces and structured motion in a more complete scene representation.

    More coming soon

    Generative geometry foundation models for simulation.

World Modeling

Predict action consequences and future evolution

We study how to represent and generate a changing world. OccWorld models 3D scene dynamics, SVG explores semantic representations for generation, and Astra predicts visual sequences under different actions. We now focus on:

  • Unified multimodal world models for generation and understanding
  • From explorable to strongly interactive world models
  • Build Harnesses for Agentic World Models
Astra: generated left-turn sequenceA prerecorded, action-conditioned video generation from Astra. The right-turn example explores a different trajectory from the same scene.
10-second clip · Generated videoAstra · Full demos
Related work & resources

Research timeline

What makes a prediction useful?

Our work brings together explicit 3D dynamics, persistent latent states, semantic generative representations, and action-conditioned video. Camera-explorable scenes are not the same as worlds that respond to embodied actions. We aim to connect generation and understanding, then build harnesses that let agents test predictions through interaction.

  1. Explored in our workGenerate scene dynamics
  2. Explored in our workExplore scenes and condition futures on actions
  3. Open research goalBuild embodied world models

Open question What must a model represent to predict contact, manipulation, and persistent changes caused by an embodied agent?

  1. 2024

    3D world models

    OccSora encodes 4D occupancy and generates scenes for different trajectory conditions.
    Trajectory-conditioned 4D generation

    Predict scene evolution through explicit 3D states and persistent latent dynamics.

  2. 2025

    Interactive world models

    MoRe4D / MoGe4D generates 4D trajectories and multi-view video from a single image.
    MoRe4D / MoGe4D generates 4D trajectories and multi-view video from a single image.

    Connect generative representations, explorable 4D scenes, and action-conditioned futures.

  3. 2026

    Embodied world models

    Move from exploring a generated scene to changing it through embodied interaction.

    More coming soon

    Model how an embodied agent's actions change the world, and use feedback to guide the next interaction.

Embodied Intelligence

Act on the world and learn from feedback

Our past work in autonomous driving spans generative end-to-end systems, world-action models, multimodal closed-loop driving foundation models, and a vision-geometry-action paradigm. We are now turning to more general-purpose robots, with two questions:

  • How can humanoid robots acquire general whole-body intelligence?
  • How can touch enable more dexterous grasping?
DVGT-2: reconstruction and trajectory planningDVGT-2 jointly reconstructs dense geometry and predicts trajectories from streaming visual observations.
10-second excerpt · Planning visualizationDVGT-2 · Full demo
Related work & resources

Research timeline

Planning and feedback

Our driving work connects generation, spatial understanding, language, and action. These results motivate a broader research direction in humanoid robots, but driving performance alone does not establish general whole-body control or tactile dexterity.

  1. Explored in our workGenerate driving plans end to end
  2. Explored in our workConnect world models to actions
  3. Open research goalDevelop whole-body and tactile intelligence

Open question How can a humanoid coordinate locomotion and dexterous manipulation, using touch and visual feedback to adapt its actions?

  1. 2024

    End-to-end autonomous driving

    GenAD contrasts a sequential perception-prediction-planning pipeline with joint future generation.
    Generative driving paradigm

    Connect perception, future generation, and planning within one driving model.

  2. 2025

    World-action models

    Vega produces different driving trajectories and future images for different instructions in the same scene.
    Instruction-conditioned plans

    Ground driving decisions in predicted futures, language, and visual geometry.

  3. 2026

    General whole-body intelligence

    Extend our focus from driving to humanoid locomotion and dexterous manipulation.

    More coming soon

    General-purpose humanoid locomotion and whole-body coordination, together with tactile dexterous manipulation.

Large Models and AI Safety

Large models and AI safety are shared foundations for the Physical Superintelligence (PSI) system. Our research began with similarity metric learning and visual representations, and has gradually expanded to large multimodal models, spatial agents, and AI-generated content verification.

Large Models: the Cognitive Brain of PSI Agents

We aim to make large models the cognitive core of PSI agents, connecting perception, language, reasoning, and decision-making. We currently focus on two foundational capabilities: efficient multimodal inference and spatial agents in 3D environments.

  • SparseVLMReduce visual tokens while retaining evidence relevant to the question.
  • Proxy3DConnect compact 3D representations to a vision-language model's spatial reasoning.

More coming soon

We are exploring multi-agent evolution and desktop agents.

AI Safety: the Safety Net for PSI Agents

We study safeguards for reliable, safe, and human-friendly PSI agents. Our current entry point is generated-content verification: detecting synthetic images and identifying forgery cues in generated videos, with evidence people can inspect.

  • UniGenDetJointly learn image generation and generated-image detection, with interpretable feedback.
  • SkyraLocate video artifacts in space and time, and explain what went wrong.

More coming soon

We aim to extend safety from digital AI to physical AI, and from passive protection to proactive prevention.

Research timeline

  1. 2017–2022

    Similarity metric learning

    IDML compares images using semantic differences and uncertainty.
    IDML compares images using semantic differences and uncertainty.

    Learn representations that capture structure, similarity, and uncertainty.

    • HDMLPreprint CVPR 2019

      Learn representations with informative examples

      Earlier work on metric learning generated label-preserving examples with adaptive difficulty, exploring how training samples shape the geometry of a learned representation.

      HDML manipulates sample difficulty in embedding space and generates corresponding features.
      Earlier representation learning
    • SDMLPublication ECCV 2020

      Learn similarity between spatial layouts

      SDML embeds images and room layouts in a shared space whose distances reflect structural differences, then decodes the scene layout. The date shown is the publisher's online publication date; no earlier preprint is verified.

      SDML learns structural similarity between images and room layouts.
      SDML learns structural similarity between images and room layouts.
    • IDMLPreprint TPAMI 2024

      Make similarity uncertainty-aware

      IDML combines semantic and uncertainty embeddings for image comparison. The date and citation count refer to the original 2022 preprint; the extended study appeared in TPAMI 2024.

      IDML compares images using semantic differences and uncertainty.
      IDML compares images using semantic differences and uncertainty.
  2. 2023

    Visual foundation models

    SpatialFormer exchanges information between image tokens and explicit spatial tokens.
    SpatialFormer exchanges information between image tokens and explicit spatial tokens.

    Study transferable visual representations, supervision, and spatial structure.

    • OPERAPreprint ICCV 2023

      Combine complementary visual supervision

      OPERA brings supervised and self-supervised learning into a hierarchical representation framework, exploring how different supervision signals can support transferable visual features.

      OPERA combines different supervision signals through hierarchical representations.
      OPERA combines different supervision signals through hierarchical representations.
    • TL-AlignPreprint ICCV 2023

      Align visual tokens and training labels

      TL-Align accounts for token mixing in vision transformers when assigning labels to augmented images, improving the alignment between visual evidence and training supervision.

      TL-Align aligns mixed-image labels with the visual tokens used by a transformer.
      TL-Align aligns mixed-image labels with the visual tokens used by a transformer.
    • SpatialFormerPublication ECCV 2024

      Give visual backbones explicit spatial structure

      SpatialFormer couples spatial tokens with image tokens to learn transferable scene representations. The date shown is the publisher's online publication date; no earlier preprint is verified.

      SpatialFormer exchanges information between image tokens and explicit spatial tokens.
      SpatialFormer exchanges information between image tokens and explicit spatial tokens.
      4
  3. 2024

    Efficient large models

    SparseVLM retains different visual tokens according to the question being asked.
    Efficient multimodal inference

    Preserve task-relevant visual evidence while reducing inference cost.

  4. 2025

    AI-generated content verification

    Skyra compares video artifact explanations and grounds its findings with bounding boxes and timestamps.
    Video artifact localization and explanation

    Detect generated content and inspect its spatial and temporal inconsistencies.

    • SkyraPreprint CVPR 2026

      Make generated-video failures inspectable

      Detected human-perceivable artifacts in generated videos, grounding explanations in spatial regions and timestamps. This provides evidence for content verification, not a guarantee of physical-agent safety.

      Skyra compares video artifact explanations and grounds its findings with bounding boxes and timestamps.
      Video artifact localization and explanation
    • UniGenDetPreprint CVPR 2026

      Learn generation and detection together

      Unified image generation and generated-image detection, using interpretable detector feedback to improve generation while strengthening detection. The focus is generated-content analysis, not general agent alignment.

      UniGenDet jointly trains an image generator and a detector with real-or-fake judgments and explanatory feedback.
      Joint generation and detection
    • GenWorldPreprint arXiv 2025

      Examine physical consistency in generated video

      GenWorld studies whether generated videos are distinguishable from real-world observations. Its benchmark and detector examine cross-view consistency, linking generated-content verification to physical realism.

      GenWorld evaluates real and generated videos through spatial consistency.
      GenWorld evaluates real and generated videos through spatial consistency.
  5. 2026

    Spatial agents

    Proxy3D connects semantic and geometry encoders, a 3D proxy generator, and a language model for spatial questions and grounding.
    3D proxies for vision-language reasoning

    Ground large-model reasoning in spatial evidence.

    • Proxy3DPreprint CVPR 2026

      Ground multimodal reasoning in compact 3D structure

      Combined semantic clustering and geometric features into compact 3D proxies, then aligned them with a vision-language model for spatial question answering and grounding. This develops a spatial reasoning ingredient for agents.

      Proxy3D connects semantic and geometry encoders, a 3D proxy generator, and a language model for spatial questions and grounding.
      3D proxies for vision-language reasoning

    More coming soon

    Multi-agent evolution and desktop agents.

Looking ahead

Spatial intelligence as a foundation

  • Physical laws have a simpler, more fundamental expression in the 3D world we inhabit (Allegory of the Cave). Directly mapping high-dimensional 2D images to a low-dimensional action space is inefficient scaling, prone to learning correlations that fail to generalize and lead to model hallucinations.
  • Realizing true PSI requires not just scaling up passively collected data, but building large-scale, diverse digital twin worlds for closed-loop training, ultimately combined with LLMs to enable recursive self-improvement of PSI.

Integrating world models and LLMs

  • Language is a high-level representation of the world developed over human history, providing a strong foundation for the autoregressive objective of LLMs. Yet language alone cannot fully capture the dynamics of physical interactions. LLMs and world models need to be integrated.
  • World representations that support both generation and understanding could unify language, vision, and action at the levels of representation and architecture. We aim for understanding, prediction, and decision-making to develop together within a single model, ultimately leading to a fully unified model.

The rise of humanoid robots

  • Humanoid robots are the most likely path to PSI. Human activity provides abundant videos of human actions. The key is to transfer this experience to robot locomotion and manipulation, then learn body control through the robots' own feedback.
  • Much of the built environment is designed around the human body. We see humanoid robots as a promising way to perform general-purpose tasks in existing human spaces without extensive redesign, while collaborating with specialized robots of other forms.

About me

I am a postdoctoral fellow in EECS at UC Berkeley, affiliated with BAIR, working with Kurt Keutzer. I received my Ph.D. in Automation from Tsinghua University, advised by Jie Zhou and Jiwen Lu, and my B.S. in Physics from Tsinghua.

Contact and collaboration

I welcome conversations with students and collaborators interested in world models, embodied intelligence, spatial intelligence, and large models. For in-person or remote research internships at BAIR or Tsinghua University, please get in touch.

wenzhao.zheng@outlook.com

Research figure