University of Pennsylvania · Fall 2026

World Models

This course introduces world models—learned representations and predictors of environment dynamics for perception, planning, and decision making. We study how they support control and reasoning across reinforcement learning, video and 3D, multimodal agents, and robotics.

Instructor
Jiatao Gu ↗
Time
Tuesdays & Thursdays · 12:00–1:29 PM
Room
AGH 105A & 105B ↗
Office hours
Thursdays · 4:30–5:30 PM · AGH 423
Section
001 · CRN 90075

Temporary schedule update

Travel-related class changes

Updated

Tue · Sep 15Held virtually via Zoom

Thu · Sep 17Class tentatively canceled

Tue · Sep 22Tentative: virtual via Zoom · link on Canvas

What is next

Course staff

Meet the instructional team. Instructor office hours are Thursdays, 4:30–5:30 PM in AGH 423.

Tentative Schedule

23 lectures, one final project presentation, and one final exam, plus four confirmed no-class dates and one tentative cancellation, following the Penn academic calendar. Assignment and project deadlines are collected under Coursework. Slides and readings appear here as the semester goes on.

Representation, prediction, and interaction are recurring themes, not sequential phases. Later application units—including 3D and robotics—bring these questions back together.

Download course calendar (.ics) ↓
  • Latent World ModelsL3–L10
  • Generative & Video ModelsL11–L14
  • Spatial & Physical ModelsL15–L17
  • Robotics & AgentsL18–L22
  1. 01

    World Models: An Overview

    Observation, state, transition, memory, action, prediction, simulation, planning, and reasoning.

    Slides (PDF)
    Lecture 1
  2. 02

    World Models: History, Foundations, and Probabilistic Formulation

    Trajectory distributions, latent state, partial observability, and one-step versus rollout objectives.

    Lecture 2
  3. 03

    Environments, Simulators, and Rollouts

    Environment interfaces, known transition rules, reset and step, observations, actions, feedback, and trajectory collection.

    Lecture 3
  4. 04

    State-Space Models

    Transition and emission models, filtering, smoothing, observability, and belief-state inference.

    Lecture 4
  5. 05

    Self-supervised Representation Learning I

    What representations preserve, how probes test them, and how reconstruction, masked prediction, autoregressive prediction, and contrastive objectives shape features.

    Lecture 5
  6. 06

    Self-supervised Representation Learning II

    Joint-embedding prediction, target encoders, masked latent prediction, and action-conditioned extensions.

    Lecture 6
  7. 07

    Latent-Variable and Adversarial Models

    Held virtually via Zoom. VAEs, the ELBO, posterior collapse, GANs, and explicit versus implicit generation.

    Lecture 7Zoom
  8. Tentative — No class

    Class is tentatively canceled due to urgent instructor travel. Confirmation will be posted on Canvas.

    Tentative
  9. 08

    Latent World Models

    Tentative format change: virtual via Zoom; the link will be posted on Canvas. Autoregressive prediction, history and action conditioning, training versus rollout, and the World Models (2018) system.

    Lecture 8ZoomTentative
  10. 09

    Latent Dynamics and Planning

    Sequential VAEs, recurrent state-space models, prior and posterior inference, and latent-space planning with PlaNet, MPC, and CEM.

    Lecture 9
  11. 10

    Policy and Value Learning with World Models

    Policy and value learning in imagination with Dreamer; task-oriented model learning and short-horizon planning with TD-MPC; model bias and closed-loop evaluation.

    Lecture 10
  12. No class — Fall Term Break

    University break.

  13. No class — COLM 2026

    Instructor conference travel.

  14. No class — COLM 2026

    Instructor conference travel.

  15. 11

    Diffusion and Flow Matching

    Denoising, scores, velocity objectives, probability-flow views, and training–sampling tradeoffs.

    Lecture 11
  16. 12

    Video World Models I: Generation and Prediction

    Pixel, token, and latent video representations; spatiotemporal generative architectures, conditional prediction, and frame, chunk, and block-level rollout.

    Lecture 12
  17. 13

    Video World Models II: Interaction and Long-Horizon Rollouts

    Action conditioning, persistent memory, long-context consistency, closed-loop drift, and interactive inference efficiency.

    Lecture 13
  18. 14

    Normalizing and Autoregressive Flows

    Change of variables, invertible transformations, triangular Jacobians, TARFlow, and STARFlow.

    Lecture 14
  19. 15

    Spatial World Models I: Geometry and 3D Representations

    Coordinate frames, depth, point clouds, occupancy, radiance fields, and Gaussian representations.

    Lecture 15
  20. 16

    Spatial World Models II: 4D Dynamics and Interaction

    Scene flow, tracking, dynamic occupancy, contact, action conditioning, and spatial consistency.

    Lecture 16
  21. 17

    Neural Physics and Learned Physical Dynamics

    Learned simulators for particles, meshes, fluids, and deformable objects; graph networks, neural operators, physical priors, and long-horizon stability.

    Lecture 17
  22. 18

    World Models for Robot Learning I: Simulation, Control, and Sim-to-Real

    Classical and learned dynamics for model-based control, system identification, domain randomization, and sim-to-real transfer.

    Lecture 18
  23. 19

    World Models for Robot Learning II: Vision-Language-Action and World-Action Models

    Vision-language-action policies, latent actions, video pretraining, and world-action models that jointly predict futures and robot actions.

    Lecture 19
  24. 20

    LLMs as World Models: Simulation, State Tracking, and Grounding

    Language and multimodal context as observation, belief state, action, feedback, and memory.

    Lecture 20
  25. 21

    Reasoning Models: Deliberation, Recurrence, and Test-Time Computation

    Sequential deliberation, branching search, verification, recurrent depth, and adaptive computation.

    Lecture 21
  26. 22

    World Models for Digital Agents

    World prediction, reasoning, tools, and feedback in games, GUIs, software, and multi-agent systems.

    Lecture 22
  27. 23

    Evaluating World Models: Utility, Robustness, and Open Problems

    Penn follows a Thursday schedule on Tuesday. We close with utility, controllability, calibration, OOD behavior, intervention, drift, latency, and failure.

    Lecture 23
  28. No class — Thanksgiving Break

    University break.

  29. 24

    Final Project Presentations

    Student project presentations, discussion, and course synthesis.

  30. 25

    Final Exam

    In-class final exam covering the core concepts developed throughout the course.

Lecture slides

Lecture 1World Models: An Overview

Resources

A growing collection of useful blogs, tutorials, seminars, and systems. New links will be added throughout the semester.

Updated throughout Fall 2026

Coursework

All assignment and project dates are collected here, separate from the lecture schedule. Specifications and grading weights are posted before each release.

01

Assignment 1

Out
Sep 10
Due
Sep 24
02

Assignment 2

Out
Sep 30
Due
Oct 19
03

Assignment 3

Out
Oct 21
Due
Nov 18

Final project

Explore a world-modeling method in one application domain, making the modeled state, transition, action interface, and evaluation criteria explicit.

Physical systemsVideo & spatial worldsRobot learningLanguage & digital agents
  1. Proposal

    Question, domain, method, and evaluation plan.

  2. Checkpoint

    Working system, early evidence, and risks.

  3. Final project presentations

    Results, failure analysis, and discussion.

  4. Report + code

    Final submission; exact format to be announced.

Logistics

Everything marked “to be announced” is confirmed and posted here before classes begin.

Office hours & course contact

Instructor office hours: Thursdays, 4:30–5:30 PM, AGH 423. Course questions and temporary meeting links are posted on Canvas.

Grading, late work & AI use

To be announced

Academic integrity

Penn’s Code of Academic Integrity applies to all work in this course.

Read the code ↗
Is this a reinforcement learning course?

No. We teach the RL and control machinery needed to understand how learned world models support decisions, but the course also covers representation learning, generative modeling, video and 3D, robotics, language, and digital agents.

Is the schedule final?

Class dates and major milestones are set. Topic order, readings, and guest speakers may still change, and per-session materials are linked from the schedule as they are ready.

What will the final project look like?

You explore a world-modeling method in one application domain. The rubric, team policy, and submission format are posted before proposals are due.