Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton. ACM 1991; Key: dyna architecture; ExpEnv: None
A curated list of awesome model based RL resources (continually updated)
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Richard S. Sutton. ACM 1991; Key: dyna architecture; ExpEnv: None
Marc Peter Deisenroth, Carl Edward Rasmussen. ICML 2011; Key: probabilistic dynamics model; ExpEnv: cart-pole system, robotic unicycle
Sergey Levine, Vladlen Koltun. ICML 2014; Key: guided policy search; ExpEnv: mujoco
Nicolas Heess, Greg Wayne, David Silver, Timothy Lillicrap, Yuval Tassa, Tom Erez. NIPS 2015; Key: backpropagation through paths, gradient on real trajectory; ExpEnv: mujoco
Junhyuk Oh, Satinder Singh, Honglak Lee. NIPS 2017; Key: value-prediction model; ExpEnv: collect domain, atari
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, Honglak Lee. NIPS 2018; Key: ensemble model and Qnet, value expansion; ExpEnv: mujoco, roboschool
David Ha, Jürgen Schmidhuber. NIPS 2018; Key: vae(representation), rnn(predictive model); ExpEnv: car racing, vizdoom
Kurtland Chua, Roberto Calandra, Rowan McAllister, Sergey Levine. NIPS 2018; Key: probabilistic ensembles with trajectory sampling; ExpEnv: cartpole, mujoco
Michael Janner, Justin Fu, Marvin Zhang, Sergey Levine. NeurIPS 2019; Key: ensemble model, sac, k-branched rollout; ExpEnv: mujoco
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, Tengyu Ma. ICLR 2019; Key: Discrepancy Bounds Design, ME-TRPO with multi-step, Entropy regularization; ExpEnv: mujoco
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, Pieter Abbeel. ICLR 2018; Key: ensemble model, TRPO; ExpEnv: mujoco
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, Mohammad Norouzi. ICLR 2019; Key: DreamerV1, latent space imagination; ExpEnv: deepmind control suite, atari, deepmind lab
Tingwu Wang, Jimmy Ba. ICLR 2020; Key: model-based policy planning in action space and parameter space; ExpEnv: mujoco
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, David Silver. Nature 2020; Key: MCTS, value equivalence; ExpEnv: chess, shogi, go, atari
Shenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik, Seonghyeon Ye, Sihyun Yu, Wei-Cheng Tseng, Yuzhu Dong, Kaichun Mo, Chen-Hsuan Lin, Jiannan Xiang, Yuqi Xie, Ruijie Zheng, Dantong Niu, Pooya Jannaty, Jinwei Gu, Jun Zhang, Jitendra Malik, Pieter Abbeel, Ming-Yu Liu, Yuke Zhu, Joel Jang, Jim…
Jiankai Zuo, Yang Zhang, Yu Zhang, Jiarui Liang, Yaying Zhang. ICML 2026; Key: continuous-time dynamics, latent dynamics, irregular events, neural ODE; ExpEnv: continuous-time event prediction
Yaxuan Li, Junjie Wen, Zhongyi Zhou, Yefei Chen, Chaomin Shen, Yaxin Peng, Yichen Zhu. ICML 2026; Key: world model, policy evaluation, discrete diffusion, robotics; ExpEnv: LIBERO, RoboTwin
Chaokang Jiang, Desen Zhou, Jiuming Liu, Li Sun. ICML 2026; Key: world model, vector graphics, diffusion flow, streaming; ExpEnv: Waymo
Yuchen Wang, Jiangtao Kong, Sizhe Wei, Xiaochang Li, Haohong Lin, Hongjue Zhao, Tianyi Zhou, Lu Gan, Huajie Shao. ICML 2026; Key: robot world model, trajectory model, cross-embodiment, robotics; ExpEnv: diverse robotic systems
Chengrui Li, Yunmiao Wang, Yule Wang, Weihan Li, Dieter Jaeger, Anqi Wu. ICML 2026; Key: latent dynamics, RNN, neural data, low-rank; ExpEnv: neural data
Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Marc Bellemare, Alessandro Lazaric, Ahmed Touati. ICML 2026; Key: world model, jumpy prediction, compositional planning, MBRL; ExpEnv: compositional planning benchmarks
Wei-Di Chang, Mikael Henaff, Brandon Amos, Gregory Dudek, Scott Fujimoto. ICML 2026; Key: model-based RL, search, planning, empirical study; ExpEnv: MBRL benchmarks
Jiayu Chen, Le Xu, Aravind Venugopal, Jeff Schneider. ICML 2026; Key: offline MBRL, world model adaptation, policy-driven, robustness; ExpEnv: D4RL, MuJoCo
Tianwei Ni, Esther Derman, Vineet Jain, Vincent Taboga, Siamak Ravanbakhsh, Pierre-Luc Bacon. ICML 2026; Key: long-horizon RL, model-based, offline RL, conservatism-free; ExpEnv: D4RL
Guojian Zhan, Likun Wang, Feihong Zhang, Yang Guan, Shengbo Li. ICML 2026; Key: model-based RL, policy improvement, harmonized objective; ExpEnv: DMControl
Jiafei Lyu, Zichuan Lin, Scott Fujimoto, Kai Yang, Yangkun Chen, Saiyong Yang, Zongqing Lu, Deheng Ye. ICML 2026; Key: model-based representations, debiasing, sample efficiency, continuous control; ExpEnv: continuous control
Jonathan Spieler, Sven Behnke. ICML 2026; Key: latent imagination, gradient-based MPC, world model; ExpEnv: continuous control
Xingyu Jiang, Yuheng Pan, Mukang You, Xiuhui Zhang, Ning Gao, Guanwei Yan, Hao Li, Yue Deng. ICML 2026; Key: world model, latent-space value alignment, MBRL; ExpEnv: Atari, DMControl
Muxi Tao, Jiangtao Wen, Yuxing Han. ICML 2026; Key: experience replay, model-based RL, prioritization; ExpEnv: MuJoCo
Hojun Chung, Junseo Lee, Songhwai Oh. ICML 2026; Key: universal horizon models, offline RL, model-based; ExpEnv: D4RL
Yongchao Huang. ICML 2026; Key: JEPA, variational, probabilistic world model; ExpEnv: representation learning benchmarks
Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun, Randall Balestriero. ICML 2026; Key: JEPA, causal, latent interventions, world model; ExpEnv: object-centric video
Samo Hromadka, Kai Biegun, Lior Fox, James Heald, Maneesh Sahani. ICML 2026; Key: latent dynamics, maximum likelihood, reconstruction-free
Yaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan, Wendelin Boehmer. ICML 2026; Key: tree search, sequential Monte Carlo, MCTS alternative, planning; ExpEnv: planning benchmarks
Michael Psenka, Michael Rabbat, Aditi Krishnapriyan, Yann LeCun, Amir Bar. ICML 2026; Key: parallel planning, stochastic gradient, world model; ExpEnv: world-model planning benchmarks
Zhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng, Haonan Wang, Zhichao Wu, Haoxin Lin, Pierre-Luc Bacon, Yang Yu. ICML 2026; Key: world model, VLA, model-based RL; ExpEnv: VLA benchmarks
Yanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen, Percy Liang, Chelsea Finn. ICML 2026; Key: VLA, world model, iterative co-improvement; ExpEnv: VLA robotics
Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Jimmy Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, Franziska Meier. ICML 2026; Key: world model, cross-embodiment, foundation model, latent actions; ExpEnv: cross-embodiment robotics
Quentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas, Yann LeCun, Michael Rabbat. ICML 2026; Key: latent actions, world model, in-the-wild video; ExpEnv: real-world video
Zhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai, Yilun Du. ICML 2026; Key: 4D latent world model, robot planning, structured representations; ExpEnv: robot planning
Ruiqi Wu, Xuanhua He, Meng Cheng, Tianyu Yang, Yong Zhang, Zhuoliang Kang, Xunliang Cai, Xiaoming Wei, Chunle Guo, Chongyi Li, Ming-Ming Cheng. ICML 2026; Key: long-horizon, interactive world model, hierarchical memory; ExpEnv: interactive world modeling
Jianjie Fang, Yingshan Lei, Qin Wan, Ziyou Wang, Yuchao Huang, Yongyan Xu, Baining Zhao, Weichen Zhang, Chen Gao, Xinlei Chen, Yong Li. ICML 2026; Key: world model benchmark, interactive, unified action generation; ExpEnv: interactive world modeling
Marco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric, Andrea Tirinzoni. ICLR 2026; Key: zero-shot RL, unsupervised RL, self-predictive representations, JEPA, world modeling; ExpEnv: zero-shot RL benchmarks
Emre Adabag, Marcus Greiff, John Subosits, Thomas Jonathan Lew. ICLR 2026; Key: differentiable optimization, model predictive control, GPU acceleration, robotics; ExpEnv: RL benchmark control tasks
Tal Daniel, Carl Qi, Dan Haramati, Amir Zadeh, Chuan Li, Aviv Tamar, Deepak Pathak, David Held. ICLR 2026; Key: world model, object-centric, latent dynamics, self-supervised, video prediction; ExpEnv: real-world multi-object video, goal-conditioned imitation
Jiahan Zhang, et al. ICLR 2026; Key: world models, embodied AI, closed-loop evaluation, generative WM benchmark; ExpEnv: closed-loop embodied agent benchmarks
Naoki Morihira, Amal Nahar, Kartik Bharadwaj, Yasuhiro Kato, Akinobu Hayashi, Tatsuya Harada. ICLR 2026; Key: Dreamer, MBRL, world model, decoder-free, representation learning; ExpEnv: DeepMind Control Suite, Meta-World
Nicklas Hansen, Hao Su, Xiaolong Wang. ICLR 2026; Key: world models, multi-task RL, continuous control, language-conditioned; ExpEnv: 200-task multi-domain continuous control
Boxuan Zhang, Weipu Zhang, Zhaohan Feng, Wei Xiao, Jian Sun, Jie Chen, Gang Wang. ICLR 2026; Key: multi-task RL, world model, mixture-of-experts, latent dynamics, transformer; ExpEnv: Atari 100k, multi-task RL
Mehran Aghabozorgi, Alireza Moazeni, Yanshu Zhang, Ke Li. ICLR 2026; Key: model-based RL, world model, uncertainty quantification, IMLE; ExpEnv: DeepMind Control, MyoSuite
Weipu Zhang, Adam Jelley, Trevor McInroe, Amos Storkey, Gang Wang. ICLR 2026; Key: model-based RL, object-centric, world model, segmentation; ExpEnv: Atari, Hollow Knight
Lior Cohen, Ofir Nabati, Kaixin Wang, Navdeep Kumar, Shie Mannor. ICLR 2026; Key: world models, diffusion, model-based RL, on-policy rollout; ExpEnv: Atari 100K, Craftium
Junha Chun, Youngjoon Jeong, Taesup Kim. ICLR 2026; Key: world model, planning, MPC, vision transformer, latent rollout; ExpEnv: visual robotics planning tasks
Yihong Guo, Yu Yang, Pan Xu, Anqi Liu. ICLR 2026; Key: model-based RL, off-dynamics RL, domain adaptation, offline RL; ExpEnv: off-dynamics offline RL benchmarks
Kwanyoung Park, Seohong Park, Youngwoon Lee, Sergey Levine. ICLR 2026; Key: offline RL, world models, model-based RL, action chunking, long-horizon; ExpEnv: long-horizon offline RL
Wolfgang Lehrach, et al. ICLR 2026; Key: code world model, LLM-generated WM, information-set MCTS, planning, AlphaZero-family; ExpEnv: general-game-playing board/card games
Yuan Pu, Yazhe Niu, Jia Tang, Junyu Xiong, Shuai Hu, Hongsheng Li. ICLR 2026; Key: multi-task RL, UniZero, MCTS, latent-space planning, world model, MoE; ExpEnv: multi-task UniZero benchmarks
Jiayu Chen, Le Xu, Wentse Chen, Jeff Schneider. ICLR 2026; Key: offline MBRL, Bayes-adaptive MDP, MCTS, model uncertainty; ExpEnv: offline RL benchmarks
Yun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang, Ti-Rong Wu. ICLR 2026; Key: AlphaZero, MCTS, regret prioritization, search control; ExpEnv: board games
Fangqi Zhu, Zhengyang Yan, Zicong Hong, Quanxin Shou, Xiao Ma, Song Guo. ICLR 2026; Key: world model, VLA, GRPO, on-policy RL in imagination; ExpEnv: VLA robotic manipulation
Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea Finn. ICLR 2026; Key: world model, VLA, robot manipulation, policy evaluation, policy improvement; ExpEnv: robot manipulation
Yue Liao, et al. ICLR 2026; Key: world model, foundation model, embodied AI, robotic manipulation, video diffusion; ExpEnv: robotic manipulation
Siqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao, Mingsheng Long. ICLR 2026; Key: world models, video diffusion models, causal video generation; ExpEnv: sequential decision-making
Moo Jin Kim, et al. ICLR 2026; Key: world models, robotics, manipulation, model-based planning, video generation; ExpEnv: robotic manipulation
Julian Hector Quevedo, Ansh Kumar Sharma, Yixiang Sun, Varad Suryavanshi, Percy Liang, Sherry Yang. ICLR 2026; Key: world model, policy evaluation, generative simulator, autoregressive video generation; ExpEnv: VLA robotic policies, WorldGym
Zaid Khan, Archiki Prasad, Elias Stengel-Eskin, Jaemin Cho, Mohit Bansal. ICLR 2026; Key: symbolic world models, programmatic RL, probabilistic programs, stochastic environments; ExpEnv: stochastic complex environments
Yichao Liang, Thanh Dat Nguyen, Cambridge Yang, Tianyang Li, Joshua B. Tenenbaum, Carl Edward Rasmussen, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis. ICLR 2026; Key: abstract world models, symbolic state representations, learned dynamics, planning, neurosymbolic; ExpEnv: tabletop robotics
Keyan Miao, et al. ICLR 2026; Key: Koopman operator, latent dynamics, MPC, controllability, nonlinear dynamics; ExpEnv: nonlinear control benchmarks
Tyler Han, et al. ICLR 2026; Key: imitation learning, RL, model predictive control, adversarial IRL; ExpEnv: robot imitation benchmarks
Yi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou, Tianyu Cui, Le Chen, Dieter Büchler, Arno Solin, Juho Kannala, Joni Pajarinen. ICLR 2026; Key: world model, offline-to-online RL, non-curated data; ExpEnv: offline-to-online RL benchmarks
Misagh Soltani, Forest Agostinelli. NeurIPS 2025; Key: visual planning, aligned representations, discrete latent state, heuristic search; ExpEnv: Rubik's Cube, Sokoban
Mingsheng Long, et al. NeurIPS 2025; Key: world model training, decision-aware, verifiable rewards; ExpEnv: text games, robot manipulation
Microsoft Research et al. NeurIPS 2025; Key: structured world models, object-centric, physics modeling; ExpEnv: physical interaction, object manipulation
Anonymous et al. NeurIPS 2025; Key: exploration, diffusion model, synthetic experience, data augmentation; ExpEnv: mujoco, sparse reward tasks
Xiu Li, et al. NeurIPS 2025; Key: multi-agent MBRL, diffusion-inspired, sequence modeling, joint distribution; ExpEnv: SMAC, MPE
Yarden As, Chengrui Qu, Benjamin Unger, Dongho Kang, Max van der Hart, Laixi Shi, Stelian Coros, Adam Wierman, Andreas Krause. NeurIPS 2025; Key: safe MBRL, sim-to-real, ensemble uncertainty, robust control; ExpEnv: real-world robotics, safety gym
Shrinivas Ramasubramanian, Benjamin Freed, Alexandre Capone, Jeff Schneider. NeurIPS 2025; Key: model error, simulation lemma, model generalization,; ExpEnv: DMC, Atari100k, HumanoidBench
Antoine Dedieu, Joseph Ortiz, Xinghua Lou, Carter Wendelken, Wolfgang Lehrach, J Swaroop Guntupalli, Miguel Lazaro-Gredilla, Kevin Murphy; Key: dyna with warmup, patch nearestneighbor tokenization, block teacher forcing; OpenReview: 4, 4, 4, 3; ExpEnv: craftax-classic
Brett Barkley, David Fridovich-Keil; Key: Dyna-style algorithms significantly degrades performance across most DMC environments.; OpenReview: 4, 4, 3, 2; ExpEnv: gym, DeepMind Control Suite
Haotian Fu, Yixiang Sun, Michael L. Littman, George Konidaris; Key: synthetic experience rehearsal, regaining memories through exploration; OpenReview: 4, 3, 3, 3; ExpEnv: mini-grid, deepmind control suite
Anh N Nhu, Sanghyun Son, Ming Lin; Key: condition on the time-step size ∆t and and train over a diverse range of ∆t values; OpenReview: 4, 3, 3; ExpEnv: meta-world control tasks, PDE-control tasks
Minting Pan, Yitao Zheng, Jiajian Li, Yunbo Wang, Xiaokang Yang; Key: behavior abstraction network, hierarchical world model; OpenReview: 3, 3, 3, 2; ExpEnv: meta-world, carla, minedojo
Dongsu Lee, Minhae Kwon; Key: learn a latent abstraction that captures a temporal distance from both trajectory and transition levels of state space.; OpenReview: 4, 3, 3, 2; ExpEnv: D4RL, AntMaze, FrankaKitchen, CALVIN, pixel-based FrankaKitchen.
Dongchi Huang, Jiaqi WANG, Yang Li, Chunhe Xia, Tianle Zhang, Kaige Zhang; Key: leverage privileged information through privileged representation alignment and an asymmetric actor-critic structure; OpenReview: 3, 3, 3; ExpEnv: safety gymnasium benchmark, guard benchmark
Shangzhe Li, Zhiao Huang, Hao Su; Key: reward-free world model, inverse soft-Q learning objective; OpenReview: 4, 3, 3, 3; ExpEnv: DMControl, MyoSuite, ManiSkill2
Yucen Wang, Rui Yu, Shenghua Wan, Le Gan, De-Chuan Zhan; Key: ground FM representations into the WM state space, model-based goal-condition RL; OpenReview: 4, 3, 3, 3; ExpEnv: DMControl, Kitchen, minecraft
Zichen Liu, Guoji Fu, Chao Du, Wee Sun Lee, Min Lin; Key: plan with online world model, regret analysis; OpenReview: 4, 4, 4, 3; ExpEnv: ContinualBench
Tim Pearce*, Tabish Rashid*, David Bignell, Raluca Georgescu, Sam Devlin, Katja Hofmann; Key: scaling laws, embodied AI, behavior cloning, world modeling, tokenizer, architecture; ExpEnv: Bleeding Edge, RT-1 (robotics), Atari, NetHack
Gaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel Pinto; Key: world models, offline learning, zero-shot planning, pretrained visual features, task-agnostic reasoning; ExpEnv: Maze, Wall, Reach, Push-T, Rope Manipulation, Granular Manipulation
Jonathan Richens, Tom Everitt, David Abel; Key: world models, goal-directed behavior, model-free learning, policy analysis, regret bounds; ExpEnv: synthetic controlled Markov process (cMP) environments with varying sample trajectories and goal depths
Yushuai Li, Hengyu Liu, Torben Bach Pedersen, Yuqiang He, Kim Guldstrand Larsen, Lu Chen, Christian S. Jensen, Jiachen Xu, Tianyi Li; Key: MuZero, robustness, reinforcement learning, state perturbations, self-supervised learning, adaptive adjustment; ExpEnv: CartPole, Pendulum, IEEE 34-bus, IEEE…
Maxime Burchi, Radu Timofte; Key: model-based reinforcement learning, world models, MaskGIT, spatial latent space, Dreamer, Transformer, efficiency; ExpEnv: Crafter, Atari 100k
Shaofeng Yin, Jialong Wu, Siqiao Huang, Xingjian Su, Xu He, Jianye Hao, Mingsheng Long; Key: world models, heterogeneous environments, pre-training, in-context learning, model transfer, trajectory data; ExpEnv: UniTraj (80 diverse environments), D4RL (HalfCheetah, Hopper, Walker2D), Cart-2-Pole,…
Raanan Y. Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo, Vasudev Lal; Key: GPT, causal inference, attention mechanism, structural causal model, zero-shot causal discovery; ExpEnv: Othello, Chess
Maxime Burchi, Radu Timofte; Key: model-based reinforcement learning, transformer network, contrastive predictive coding; ExpEnv: Atari 100k benchmark
Yang Tian, Sizhe Yang, Jia Zeng, Ping Wang, Dahua Lin, Hao Dong, Jiangmiao Pang; Key: Robotic Manipulation, Pre-training, Visual Foresight, Inverse Dynamics, Large-scale robot dataset; ExpEnv: LIBERO-LONG benchmark, CALVIN ABC-D, real-world tasks
Po-Wei Huang, Pei-Chiun Peng, Hung Guei, Ti-Rong Wu; Key: Option, Semi-MDP, MuZero, MCTS, Planning, Reinforcement Learning; ExpEnv: Atari
Claas A Voelcker, Marcel Hussing, Eric Eaton, Amir-massoud Farahmand, Igor Gilitschenski; Key: reinforcement learning, model based reinforcement learning, data augmentation, high update ratios; ExpEnv: DeepMind Control Suite
Michael Matthews, Michael Beukman, Chris Lu, Jakob Nicolaus Foerster; Key: Reinforcement Learning, Open-Endedness, Unsupervised Environment Design, Automatic Curriculum Learning, Benchmark; ExpEnv: 2D Physics-Based Tasks, Robotic Locomotion, Grasping, Video Games, Classic RL Environments
Dixant Mittal, Liwei Kang, Wee Sun Lee; Key: Planning, Reasoning, Learning to Search, Reinforcement Learning, Large Language Model; ExpEnv: Game of 24, 2D Grid Navigation, Procgen Games
Jiajian Li, Qi Wang, Yunbo Wang, Xin Jin, Yang Li, Wenjun Zeng, Xiaokang Yang; Key: Reinforcement Learning, World Models, Visual Control; ExpEnv: MineDojo
Martin Klissarov, Mikael Henaff, Roberta Raileanu, Shagun Sodhani, Pascal Vincent, Amy Zhang, Pierre-Luc Bacon, Doina Precup, Marlos C. Machado, Pierluca D'Oro; Key: Hierarchical RL, Reinforcement Learning, LLMs; ExpEnv: NetHack Learning Environment (NLE)
Authors: Tai Hoang, Huy Le, Philipp Becker, Vien Anh Ngo, Gerhard Neumann; Key: Robotic Manipulation, Equivariance, Graph Neural Networks, Reinforcement Learning, Deformable Objects; ExpEnv: Rigid insertion, rope manipulation, cloth manipulation with multiple end-effectors
Kehan Wen, Yutong Hu, Yao Mu, Lei Ke; Key: Offline-to-Online Reinforcement Learning, Model-based Reinforcement Learning, Masked Autoencoding, Robot Learning; ExpEnv: D4RL, RoboMimic
Rong-Xi Tan, Ke Xue, Shen-Huan Lyu, Haopu Shang, yaowang, Yaoyuan Wang, Fu Sheng, Chao Qian; Key: Offline model-based optimization, black-box optimization, learning to rank, learning to optimize; ExpEnv: Diverse tasks across optimization scenarios
Zijing Shi, Meng Fang, Ling Chen; Key: Large language model, Monte Carlo tree search, Text-based games; ExpEnv: Jericho benchmark
Thomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso, David Krueger; Key: reinforcement learning, interpretability, planning, probes, model-free, mechanistic interpretability, sokoban; ExpEnv: Sokoban
Wenlong Wang, Ivana Dusparic, Yucheng Shi, Ke Zhang, Vinny Cahill; Key: Mamba-2, Model based reinforcement learning, Mamba, State space models; ExpEnv: Atari 100K
Abdelhakim Benechehab, Youssef Attia El Hili, Ambroise Odonnat, Oussama Zekri, Albert Thomas, Giuseppe Paolo, Maurizio Filippone, Ievgen Redko, Balázs Kégl; Key: Model-based Reinforcement Learning, Large language models, Zero-shot Learning, In-context Learning; ExpEnv: D4RL, Pendulum, HalfCheetah,…
Bernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian Trimpe; Key: Model-Based Reinforcement Learning, Model Rollouts, Uncertainty Quantification; ExpEnv: Gym MuJoCo
Haoxin Lin, Yu-Yan Xu, Yihao Sun, Zhilong Zhang, Yi-Chen Li, Chengxing Jia, Junyin Ye, Jiaji Zhang, Yang Yu; Key: model-based reinforcement learning, any-step dynamics model; ExpEnv: D4RL, NeoRL, Gym MuJoCo-v3
Aidan Scannell, Mohammadreza Nakhaeinezhadfard, Kalle Kujanpää, Yi Zhao, Kevin Sebastian Luck, Arno Solin, Joni Pajarinen; Key: reinforcement learning, world model, representation learning, self-supervised learning, model-based reinforcement learning, continuous control; ExpEnv: deepmind control…
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, Mingsheng Long; Key: world models, video generative models, autoregressive transformer, reinforcement learning, video prediction, visual planning; ExpEnv: Meta-world
ZiRui Wang, Yue Deng, Junfeng Long, Yin Zhang; Key: reinforcement learning, model-based reinforcement learning, parallelization, sequence length, world model, eligibility trace, sample efficiency; ExpEnv: Atari 100K, DMControl
Philip Amortila, Dylan J. Foster, Nan Jiang, Akshay Krishnamurthy, Zakaria Mhammedi; Key: reinforcement learning, latent dynamics, statistical modularity, algorithmic modularity, observable-to-latent reductions, self-predictive models; ExpEnv: None
Matthew V Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth, Alexandre Laterre; Key: reinforcement learning, rl, model-based reinforcement learning, sequential monte carlo, expectation maximisation, planning; ExpEnv: Brax, Boxoban, Rubik's Cube
Yangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen, Yonghong Tian; Key: multi-modal reinforcement learning, visual RL, dynamics modeling, modality consistency, modality inconsistency, DDM; ExpEnv: CARLA, DMControl
Moritz Schneider, Robert Krug, Narunas Vaskevicius, Luigi Palmieri, Joschka Boedecker; Key: reinforcement learning, rl, model-based reinforcement learning, representation learning, pvr, visual representations; ExpEnv: DMC, ManiSkill2, Miniworld
Hao Tang, Darren Key, Kevin Ellis; Key: learn world models as code, LLM; ExpEnv: sokoban, minigrid, alfworld
Anya Sims, Cong Lu, Jakob Foerster, Yee Whye Teh; Key: edge-of-reach problem, reach-aware value learning; ExpEnv: d4rl, v-r4rl
Abdullah Akgül, Manuel Haussmann, Melih Kandemir; Key: The paper argues that uncertainty-based reward penalization introduces excessive conservatism, potentially resulting in suboptimal policies through underestimation.; ExpEnv: d4rl
Haohong Lin, Wenhao Ding, Jian Chen, Laixi Shi, Jiacheng Zhu, Bo Li, DING ZHAO; Key: objective mismatch problem, capture causal representation for both states and actions; ExpEnv: list, unlock, crash
Jung-Hoon Cho, Vindula Jayawardana, Sirui Li, Cathy Wu; Key: bayesian optimization, contextual rl; ExpEnv: gaussian process, traffic signal, eco-driving, advisory autonomy, control tasks
Guhao Feng, Han Zhong; Key: rl representation complexity; ExpEnv: mujoco
Haoyu Ma, Jialong Wu, Ningya Feng, Chenjun Xiao, Dong Li, Jianye Hao, Jianmin Wang, Mingsheng Long; Key: observation modeling and reward modeling analysis in world models; ExpEnv: meta-world, rlbench, deepmind control suite, atari 100k
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, Chuang Gan; Key: unify 3D perception, reasoning, and action with a generative world model; create a large-scale 3D embodied instruction tuning dataset; ExpEnv: rlbench, calvin
Qinlin Zhao, Jindong Wang, Yixuan Zhang, Yiqiao Jin, Kaijie Zhu, Hao Chen, Xing Xie; Key: propose a competitive framework for LLM-based agents; build a simulated competitive environment; ExpEnv: a virtual town with only restaurants and customers
Renhao Zhang, Haotian Fu, Yilin Miao, George Konidaris; Key: discrete-continuous hybrid action space, dynamics model with parameterized actions, MPC with parameterized actions; ExpEnv: platform, goal, hard goal, catch point, hard move
Ruixiang Sun, Hongyu Zang, Xin Li, Riashat Islam; Key: modified Dreamer architecture, hybrid-recurrent state space model; ExpEnv: deepmind control suite, distracted deepmind control suite, mani-skill2
Yucen Wang, Shenghua Wan, Le Gan, Shuai Feng, De-Chuan Zhan; Key: implicit action generator, action-conditioned separated world models; ExpEnv: deepmind control suite
Paul Mattes, Rainer Schlosser, Ralf Herbrich; Key: state-space models, multilayered hierarchical imagination, S5 based world model; ExpEnv: atari 100k
Lior Cohen, Kaixin Wang, Bingyi Kang, Shie Mannor; Key: pixel-based mbrl, token-based world models, retentive environment model; ExpEnv: atari 100k
Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro, Pierre-Luc Bacon; Key: actions world model; ExpEnv: double-pendulum, Myriad
Hany Hamed, Subin Kim, Dongyeong Kim, Jaesik Yoon, Sungjin Ahn; Key: during strategeic dreaming, train three policies -- highway policy, explorer policy and achiever policy, and then achieve downstream tasks; ExpEnv: 2D Navigation, 3D-Maze Navigation, RoboKitchen
Chenlu Ye, Jiafan He, Quanquan Gu, Tong Zhang; Key: theoretical analysis of adversarial corruption for model-based rl, encompassing both online and offline settings; ExpEnv: None
Mao Hong, Zhengling Qi, Yanxun Xu; Key: model-based RL, POMDP; ExpEnv: None
Chengxing Jia, Chenxiao Gao, Hao Yin, Fuxiang Zhang, Xiong-Hui Chen, Tian Xu, Lei Yuan, Zongzhang Zhang, Zhi-Hua Zhou, Yang Yu; Key: Reinforcement Learning, Model-based Reinforcement Learning, Offline Reinforcement Learning; OpenReview: 8, 8, 8, 6; ExpEnv: d4rl
Arnab Kumar Mondal, Siba Smarak Panigrahi, Sai Rajeswar, Kaleem Siddiqi, Siamak Ravanbakhsh; Key: Koopman Theory, Reinforcement Learning, Dynamical System, Planning, Longe range dynamics prediction models, Efficient forward dynamics; OpenReview: 8, 6, 5, 3; ExpEnv: mujoco
Mingde Zhao, Safa Alver, Harm van Seijen, Romain Laroche, Doina Precup, Yoshua Bengio; Key: Reinforcement Learning, Planning, Neural Networks, Temporal Difference Learning, Generalization, Deep Reinforcement Learning; OpenReview: 6, 6, 6, 5; ExpEnv: MiniGrid-BabyAI framework
Mohammad Reza Samsami, Artem Zholus, Janarthanan Rajendran, Sarath Chandar; Key: recall to imagine module, based on DreamerV3; OpenReview: 10, 8, 6; ExpEnv: bsuite, popgym, atari, deepmind control suite, memory maze
Edward S. Hu, James Springer, Oleh Rybkin, Dinesh Jayaraman; Key: privileged information, based on DreamerV3; OpenReview: 10, 8, 8, 8; ExpEnv: gymnasium robotics
Nicklas Hansen, Hao Su, Xiaolong Wang; Key: implicit world model, model predictive control, generalist td-mpc2; OpenReview: 8, 8, 8, 8; ExpEnv: deepmind control suite, Meta-World, maniskill2, myosuite
Minjun Sung, Sambhu Harimanas Karumanchi, Aditya Gahlawat, Naira Hovakimyan; Key: L1 Adaptive Control; OpenReview: 8, 6, 6, 6; ExpEnv: mujoco
Christian Gumbsch, Noor Sajid, Georg Martius, Martin V. Butz; Key: Context-specific Recurrent State Space Model, hierarchical world model; OpenReview: 8, 6, 6; ExpEnv: MiniHack, VisualPinPad, MultiWorld
Lunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas, Rui Hu, Raquel Urtasun; Key: discrete diffusion; world model; autonomous driving; OpenReview: 10, 8, 6, 6, 6; ExpEnv: NuScenes, KITTI Odometry, Argoverse2 Lidar
Xiyao Wang, Ruijie Zheng, Yanchao Sun, Ruonan Jia, Wichayaporn Wongkamjan, Huazhe Xu, Furong Huang; Key: conservative model rollouts, optimistic environment exploration; OpenReview: 6, 6, 6; ExpEnv: mujoco, deepmind control suite
Qihan Liu, Jianing Ye, Xiaoteng Ma, Jun Yang, Bin Liang, Chongjie Zhang; Key: mcts, optimistic search lambda, advantage-weighted policy optimization; OpenReview: 8, 6, 6, 6; ExpEnv: smac
Weikang Wan, Yufei Wang, Zackory Erickson, David Held; Key: differentiable trajectory optimization; OpenReview: 10, 8, 8, 5; ExpEnv: deepmind control suite, robomimic, maniskill
Zhihe YANG, Yunjian Xu; Key: conditional diffusion, offline RL; OpenReview: 8, 8, 6, 6; ExpEnv: d4rl
Zohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler, Aviv Tamar; Key: context-based meta-RL, based on dreamer; OpenReview: 6, 6, 6, 6; ExpEnv: Point Robot Navigation, Escape Room, Reacher Sparse
Fan-Ming Luo, Tian Xu, Xingchen Cao, Yang Yu; Key: reward learning, offline RL; OpenReview: 8, 6, 6, 6; ExpEnv: d4rl, NeoRL
Vint Lee, Pieter Abbeel, Youngwoon Lee; Key: learn to predict a temporally-smoothed reward rather than the exact reward at each timestep; OpenReview: 6, 6, 6, 5; ExpEnv: robodesk, hand, earthmoving
Gaspard Lambrechts, Adrien Bolland, Damien Ernst; Key: informed world model, based on DreamerV3; OpenReview: 6, 6, 6, 5; ExpEnv: varying mountain hike, deepmind control suite, pop gym, flickering atari and flickering control
Zirui Zhao, Wee Sun Lee, David Hsu; Key: LLM-MCTS; ExpEnv: VirtualHome
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian (Shawn) Ma, Yitao Liang; Key: interactive planning approach based on LLM; ExpEnv: minecraft
Fei Deng, Junyeong Park, Sungjin Ahn; Key: world model backbones; ExpEnv: MiniGrid, memory maze
Jialong Wu, Haoyu Ma, Chaoyi Deng, Mingsheng Long; Key: Contextualized World Models; ExpEnv: CARLA, deepmind control suite
Jiankai Sun, Yiqi Jiang, Jianing Qiu, Parth Nobel, Mykel J Kochenderfer, Mac Schwager; Key: Diffusion Dynamics Model; ExpEnv: d4rl, Maze2D
Yazhe Niu, Yuan Pu, Zhenjie Yang, Xueyan Li, Tong Zhou, Jiyuan Ren, Shuai Hu, Hongsheng Li, Yu Liu; Key: MCTS-style benchmark; ExpEnv: board games, atari, mujoco, gobigger
Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, Xuelong Li; Key: GPT-based diffusion model for planning and data synthesizing; ExpEnv: Meta-World, Maze2D
Sizhe Yang, Yanjie Ze, Huazhe Xu; Key: view generalization, spatial adaptive encoder; ExpEnv: deepmind control suite, adroit, xArm
Shenao Zhang, Boyi Liu, Zhaoran Wang, Tuo Zhao; Key: model-based reparameterization policy gradient method, smoothness regularization; ExpEnv: mujoco
Lin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao Kambhampati; Key: construct an explicit world (domain) model in planning domain definition language; ExpEnv: household-robot domain, tyreworld and logistics
Chuning Zhu, Max Simchowitz, Siri Gadipudi, Abhishek Gupta; Key: representation resilience for visual RL; ExpEnv: deepmind control suite, maniskill
Ziang Liu, Jeff He, Genggeng Zhou, Tobia Marcucci, Fei-Fei Li, Jiajun Wu, Yunzhu Li; Key: network sparsification, mixed-integer formulation of ReLU neural dynamics; ExpEnv: gym, cartpole, reacher
Andrew Wagenmaker, Guanya Shi, Kevin Jamieson; Key: optimal sample complexity for nonlinear dynamical systems; ExpEnv: affine dynamics system
Devleena Das, Sonia Chernova, Been Kim; Key: a joint embedding model between state-action pairs and concept-based explanations; ExpEnv: connect4, lunar lander
Lenart Treven, Jonas Hübotter, Bhavya, Florian Dorfler, Andreas Krause; Key: nonlinear ordinary differential equations, regret bound, measurement selection strategies; ExpEnv: system’s tasks
Xingyuan Zhang, Philip Becker-Ehmck, Patrick van der Smagt, Maximilian Karl; Key: pretrained world models, imitation learning from observation only; ExpEnv: deepmind control suite
Weipu Zhang, Gang Wang, Jian Sun, Yetian Yuan, Gao Huang; Key: categorical-VAE, transformer structure, DreamerV3; ExpEnv: atari
Sai Rajeswar Mudumba, Pietro Mazzaglia, Tim Verbelen, Alexandre Piche, Bart Dhoedt, Aaron Courville, Alexandre Lacoste; Key: unsupervised pretrain, task-aware finetune, dyna-mpc; ExpEnv: URLB benchmark, RWRL suite
Zhiao Huang, Litian Liang, Zhan Ling, Xuanlin Li, Chuang Gan, Hao Su; Key: multimodal policy learning, reparameterized policy gradient; ExpEnv: Meta-World, mujoco
Xiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong Huang; Key: policy-adapted model learning, weight design; ExpEnv: mujoco
Seohong Park, Sergey Levine; Key: predictable MDP abstraction, tackle model exploitation; ExpEnv: mujoco
Jacob C Walker, Eszter Vértes, Yazhe Li, Gabriel Dulac-Arnold, Ankesh Anand, Jessica Hamrick, Theophane Weber; Key Insights: (1) Is there an advantage to an agent being model-based during unsupervised exploration and/or fine-tuning? (2) What are the contributions of each component of a model-based…
Anirudh Vemula, Yuda Song, Aarti Singh, J. Bagnell, Sanjiban Choudhury; Key: objective mismatch, mbrl framework; ExpEnv: Helicopter, WideTree, Linear Dynamical System, Maze, mujoco
Kenny Young, Aditya Ramesh, Louis Kirsch, Jürgen Schmidhuber; Key: experience replay, when and how learned model generalization; ExpEnv: ProcMaze, ButtonGrid, PanFlute
Souradip Chakraborty, Amrit Bedi, Alec Koppel, Mengdi Wang, Furong Huang, Dinesh Manocha; Key: information directed sampling, kernelized Stein discrepancy; ExpEnv: DeepSea
Paavo Parmas, Takuma Seno, Yuma Aoki; Key: extension of Dreamer, total propagation computation graph; ExpEnv: deepmind control suite
Guy Tennenholtz, Nadav Merlis, Lior Shani, Martin Mladenov, Craig Boutilier; Key: non-Markov context dynamics, logistic DCMDPs, theoretical analysis, extension of MuZero; ExpEnv: MovieLens dataset
Yihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin, Junyin Ye, Yang Yu; Key: pessimistic value estimation, theoretical analysis; ExpEnv: d4rl, NeoRL
Yi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala, Joni Pajarinen; Key: representation learning, temporal consistency; ExpEnv: deepmind control suite
Isaac Kauvar, Chris Doyle, Linqi Zhou, Nick Haber; Key: extension of DreamerV3, curious replay, count-based replay, adversarial replay; ExpEnv: Crafter, deepmind control suite
Michal Nauman, Marek Cygan; Key: bias and variance, theoretical analysis; ExpEnv: deepmind control suite
Remo Sasso, Michelangelo Conserva, Paulo Rauber; Key: posterior sampling, continual value network; ExpEnv: atari
Byeongchan Kim, Min-hwan Oh; Key: count estimation, theoretical analysis; ExpEnv: d4rl
Vincent Micheli, Eloi Alonso, François Fleuret; Key: discrete autoencoder, transformer based world model; OpenReview: 8, 8, 8, 8; ExpEnv: atari
Jihwan Jeong, Xiaoyu Wang, Michael Gimelfarb, Hyunwoo Kim, Baher Abdulhai, Scott Sanner; Key: model-based offline, bayesian posterior value estimate; OpenReview: 8, 8, 6, 6; ExpEnv: d4rl
Phillip Swazinna, Steffen Udluft, Thomas Runkler; Key: let the user adapt the policy behavior after training is finished; OpenReview: 10, 8, 6, 3; ExpEnv: 2d-world, industrial benchmark
Sheng Yue, Guanbo Wang, Wei Shao, Zhaofeng Zhang, Sen Lin, Ju Ren, Junshan Zhang; Key: offline IRL, reward extrapolation error; OpenReview: 8, 8, 6, 6; ExpEnv: d4rl
Zichen Liu, Siyi Li, Wee Sun Lee, Shuicheng YAN, Zhongwen Xu; Key: offline rl, analysis of MuZero Unplugged, one-step look-ahead policy improvement; OpenReview: 8, 6, 5; ExpEnv: atari dataset
zhengyao jiang, Tianjun Zhang, Michael Janner, Yueying Li, Tim Rocktäschel, Edward Grefenstette, Yuandong Tian; Key: planning with VQ-VAE; OpenReview: 6, 6, 6, 6; ExpEnv: d4rl dataset
Ruijie Zheng, Xiyao Wang, Huazhe Xu, Furong Huang; Key: lipschitz regularization; OpenReview: 8, 8, 6, 6; ExpEnv: mujoco
Nicklas Hansen, Yixin Lin, Hao Su, Xiaolong Wang, Vikash Kumar, Aravind Rajeswaran; Key: three phases -- policy pretraining, targeted exploration, interactive learning; OpenReview: 8, 6, 6, 6; ExpEnv: adroit, meta-world, deepmind control suite
Raj Ghugare, Homanga Bharadhwaj, Benjamin Eysenbach, Sergey Levine, Ruslan Salakhutdinov; Key: Aligned Latent Models; OpenReview: 8, 6, 6, 6, 6; ExpEnv: mujoco
Daniel Palenicek, Michael Lutter, Joao Carvalho, Jan Peters; Key: longer horizons yield diminishing returns in terms of sample efficiency; OpenReview: 8, 6, 6, 6; ExpEnv: brax
Edward S. Hu, Richard Chang, Oleh Rybkin, Dinesh Jayaraman; Key: sampling-based planning, set goals for each training episode to directly optimize an intrinsic exploration reward; OpenReview: 8, 8, 8, 8, 6; ExpEnv: point maze, walker, ant maze, 3-block stack
Jinhua Zhu, Yue Wang, Lijun Wu, Tao Qin, Wengang Zhou, Tie-Yan Liu, Houqiang Li; Key: deep differentiable dynamic programming planner; OpenReview: 8, 8, 8, 6; ExpEnv: mujoco
Tongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li, Zhaoran Wang, sujay sanghavi, Dale Schuurmans, Bo Dai; Key: variational learning, representation learning; OpenReview: 8, 6, 6, 3; ExpEnv: mujoco, deepmind control suite
Yixuan Mei, Jiaxuan Gao, Weirui Ye, Shaohuai Liu, Yang Gao, Yi Wu; Key: distributed model-based rl, speed up EfficientZero; OpenReview: 6, 6, 5; ExpEnv: atari 100k
Jan Robine, Marc Höftmann, Tobias Uelwer, Stefan Harmeling; Key: autoregressive world model, Transformer-XL, balanced cross-entropy loss, balanced dataset sampling; OpenReview: 8, 6, 6, 6; ExpEnv: atari 100k
Yifan Xu, Nicklas Hansen, Zirui Wang, Yung-Chieh Chan, Hao Su, Zhuowen Tu; Key: offline multi-task pretraining, online finetuning; OpenReview: 6, 6, 6, 6; ExpEnv: atari 100k
Weirui Ye, Yunsheng Zhang, Pieter Abbeel, Yang Gao; Key: unsupervised pre-training, finetune with down-stream tasks; OpenReview: 8, 6, 6, 5; ExpEnv: atari 100k
Yifu Yuan, Jianye HAO, Fei Ni, Yao Mu, YAN ZHENG, Yujing Hu, Jinyi Liu, Yingfeng Chen, Changjie Fan; Key: jointly pretrain the multi-headed dynamics model and unsupervised exploration policy, finetune to downstream tasks; OpenReview: 6, 6, 6, 6; ExpEnv: URLB benchmark
Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Alexandre Lacoste, Sai Rajeswar; Key: world model, skill discovery, skill learning, Skill adaptation; OpenReview: 8, 8, 6, 6; ExpEnv: deepmind control suite, Meta-World
Can Chen, Yingxue Zhang, Jie Fu, Xue Liu, Mark Coates; Key: model-based, offline; OpenReview: 7, 6, 5; ExpEnv: design-bench
Shentao Yang, Shujian Zhang, Yihao Feng, Mingyuan Zhou; Key: model-based, offline, marginal importance weight; OpenReview: 7, 6, 6, 5; ExpEnv: d4rl dataset
Kaiyang Guo, Shao Yunfeng, Yanhui Geng; Key: model-based, offline; OpenReview: 8, 8, 7, 7; ExpEnv: d4rl dataset
Jiafei Lyu, Xiu Li, Zongqing Lu; Key: double check mechanism, bidirectional modeling, offline RL; OpenReview: 7, 6, 6; ExpEnv: d4rl dataset
XiaoPeng Yu, Jiechuan Jiang, Wanpeng Zhang, Haobin Jiang, Zongqing Lu; Key: multi-agent, model-based; OpenReview: 7, 6, 4, 3; ExpEnv: mpe, google research football
Zhiwei Xu, Dapeng Li, Bin Zhang, Yuan Zhan, Yunpeng Bai, Guoliang Fan; Key: multi-agent, model-based; OpenReview: 6, 5; ExpEnv: StarCraft II, Google Research Football, Multi-Agent Discrete MuJoCo
Silviu Pitis, Elliot Creager, Ajay Mandlekar, Animesh Garg; Key: data augmentation framework, offline RL; OpenReview: 7, 7, 7, 6; ExpEnv: 2D Navigation, Hook-Sweep
Tianying Ji, Yu Luo, Fuchun Sun, Mingxuan Jing, Fengxiang He, Wenbing Huang; Key: event-triggered mechanism, constrained model-shift lower-bound optimization; OpenReview: 6, 6, 5, 5; ExpEnv: mujoco
Ashish Jayant, Shalabh Bhatnagar; Key: constrained RL, model-based; OpenReview: 7, 6, 5, 5; ExpEnv: safety gym
Henger Li, Xiaolin Sun, Zizhan Zheng; Key: attack & defense, federated learning, model-based; OpenReview: 6, 6, 6, 5; ExpEnv: MNIST, FashionMNIST, EMNIST, CIFAR-10 and synthetic dataset
Anthony Hu, Gianluca Corrado, Nicolas Griffiths, Zachary Murez, Corina Gurau, Hudson Yeo, Alex Kendall, Roberto Cipolla, Jamie Shotton; Key: model-based, imitation learning, autonomous driving; OpenReview: 7, 6, 6; ExpEnv: CARLA
Han Qi, Yi Su, Aviral Kumar, Sergey Levine; Key: domain adaptation, invariant objective models, representation learning (no about model-based RL); OpenReview: 7, 6, 6, 5, 5; ExpEnv: design-bench
Haotian Fu, Shangqun Yu, Michael Littman, George Konidaris; Key: lifelong RL, variational bayesian; OpenReview: 7, 6, 6; ExpEnv: mujoco, meta-world
Zifan Wu, Chao Yu, Chen Chen, Jianye Hao, Hankz Hankui Zhuo; Key: treat the model rollout process as a sequential decision making problem; OpenReview: 7, 7, 6, 6; ExpEnv: mujoco, d4rl
Benjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Russ Salakhutdinov; Key: unified objective for model-based RL; OpenReview: 8, 8, 7, 6; ExpEnv: gridworld, mujoco, ROBEL manipulation
Marc Rigter, Bruno Lacerda, Nick Hawes; Key: offline rl, model-based rl, two-player game, adversarial model training; OpenReview: 6, 6, 6, 4; ExpEnv: d4rl
Shenao Zhang; Key: posterior sampling RL, referential update, constrained conservative update; OpenReview: 7, 7, 5, 5; ExpEnv: mujoco, N-Chain MDPs
Chenyang Wu, Tianci Li, Zongzhang Zhang, Yang Yu; Key: optimism in the face of uncertainty(OFU), BOO Regret; OpenReview: 6, 6, 5; ExpEnv: RiverSwim, Chain, Random MDPs
Alekh Agarwal, Tong Zhang; Key: posterior sampling RL, Bellman error decoupling framework; OpenReview: 7, 7, 7, 6; ExpEnv: None
Gene Li, Junbo Li, Nathan Srebro, Zhaoran Wang, Zhuoran Yang; Key: optimistic model-based, score matching; OpenReview: 7, 7, 6; ExpEnv: None
Danijar Hafner, Kuang-Huei Lee, Ian Fischer, Pieter Abbeel; Key: hierarchical RL, long-horizon and sparse reward tasks; OpenReview: 6, 6, 5; ExpEnv: atari, deepmind control suite, deepmind lab, crafter
Sahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger, Doina Precup; Key: Homomorphic Policy Gradient, Continuous MDP Homomorphisms, Lax Bisimulation Loss; OpenReview: 7, 7, 7; ExpEnv: deepmind control suite
Fei Deng, Ingook Jang, Sungjin Ahn; Key: dreamer, prototypes; ExpEnv: deepmind control suite
Tongzhou Wang, Simon Du, Antonio Torralba, Phillip Isola, Amy Zhang, Yuandong Tian; Key: representation learning, denoised model; ExpEnv: deepmind control suite, RoboDesk
Qi Wang, Herke van Hoof; Key: graph structured surrogate model, meta training; ExpEnv: atari, mujoco
Yi Wan, Ali Rahimi-Kalahroudi, Janarthanan Rajendran, Ida Momennejad, Sarath Chandar, Harm van Seijen; Key: local change adaptation; ExpEnv: GridWorldLoCA, ReacherLoCA, MountaincarLoCA
Pier Giuseppe Sessa, Maryam Kamgarpour, Andreas Krause; Key: model-based multi-agent, confidence bound; ExpEnv: SMART
Shentao Yang, Yihao Feng, Shujian Zhang, Mingyuan Zhou; Key: offline rl, model-based rl, stationary distribution regularization; ExpEnv: d4rl
Brandon Trabucco, Xinyang Geng, Aviral Kumar, Sergey Levine; Key: benchmark, offline MBO; ExpEnv: Design-Bench Benchmark Tasks
Nicklas Hansen, Hao Su, Xiaolong Wang; Key: td-learning, MPC; ExpEnv: deepmind control suite, Meta-World
Cong Lu, Philip Ball, Jack Parker-Holder, Michael Osborne, Stephen J. Roberts; Key: model-based offline, uncertainty quantification; OpenReview: 8, 8, 6, 6, 6; ExpEnv: d4rl dataset
Claas A Voelcker, Victor Liao, Animesh Garg, Amir-massoud Farahmand; Key: Value-Gradient weighted Model loss; OpenReview: 8, 8, 6, 6; ExpEnv: mujoco
Ioannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K Hubert, David Silver; Key: MCTS, stochastic MuZero; OpenReview: 10, 8, 8, 5; ExpEnv: 2048 game, Backgammon, Go
Ivo Danihelka, Arthur Guez, Julian Schrittwieser, David Silver; Key: Gumbel AlphaZero, Gumbel MuZero; OpenReview: 8, 8, 8, 6; ExpEnv: go, chess, atari
Sen Lin, Jialin Wan, Tengyu Xu, Yingbin Liang, Junshan Zhang; Key: model-based offline Meta-RL; OpenReview: 8, 6, 6, 6; ExpEnv: d4rl dataset
Lukas Froehlich, Maksym Lefarov, Melanie Zeilinger, Felix Berkenkamp; Key: model errors, on-policy corrections; OpenReview: 8, 6, 6, 5; ExpEnv: mujoco, pybullet
Jiaxian Guo, Mingming Gong, Dacheng Tao; Key: relational intervention, dynamics generalization; OpenReview: 8, 8, 6, 6; ExpEnv: Pendulum, mujoco
Homanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, Sergey Levine; Key: mutual information, visual model-based RL; OpenReview: 8, 8, 8, 6; ExpEnv: deepmind control suite, Kinetics dataset
Yanchao Sun, Ruijie Zheng, Xiyao Wang, Andrew E Cohen, Furong Huang; Key: latent dynamics model, transfer RL; OpenReview: 8, 6, 5, 5; ExpEnv: CartPole, Acrobot and Cheetah-Run, mujoco, 3DBall
Changmin Yu, Dong Li, Jianye HAO, Jun Wang, Neil Burgess; Key: representation learning, learning via retracing; OpenReview: 8, 6, 5, 3; ExpEnv: deepmind control suite
Youngmin Oh, Jinwoo Shin, Eunho Yang, Sung Ju Hwang; Key: prioritized experience replay, mbrl; OpenReview: 8, 8, 6, 5; ExpEnv: pybullet
Arunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza, Alessandro Davide Ialongo, Yuval Tassa, Jost Tobias Springenberg, Abbas Abdolmaleki, Nicolas Heess, Josh Merel, Martin Riedmiller; Key: model predictive control; OpenReview: 8, 6, 6, 6; ExpEnv: mujoco
Chongchong Li, Yue Wang, Wei Chen, Yuting Liu, Zhi-Ming Ma, Tie-Yan Liu; Key: two-model-based method, analyze model error and policy gradient; OpenReview: 8, 8, 6, 6; ExpEnv: mujoco
Yijun Yang, Jing Jiang, Tianyi Zhou, Jie Ma, Yuhui Shi; Key: model-based offline, model return-uncertainty trade-off; OpenReview: 8, 8, 6, 5; ExpEnv: d4rl dataset
Masatoshi Uehara, Wen Sun; Key: model-based offline theory, PAC bounds; OpenReview: 8, 6, 6, 5; ExpEnv: None
Edward S. Hu, Kun Huang, Oleh Rybkin, Dinesh Jayaraman; Key: world models that transfer to new robots; OpenReview: 8, 6, 6, 5; ExpEnv: mujoco, WidowX and Franka Panda robot
Hang Lai, Jian Shen, Weinan Zhang, Yimin Huang, Xing Zhang, Ruiming Tang, Yong Yu, Zhenguo Li; Key: extension of mbpo, hyper-controller learning; OpenReview: 8, 6, 6; ExpEnv: mujoco, pybullet
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, Chelsea Finn; Key: offline reinforcement learning, model-based reinforcement learning, deep reinforcement learning; OpenReview: 6, 7, 6, 8; ExpEnv: d4rl dataset
Garrett Thomas, Yuping Luo, Tengyu Ma; Key: safe rl, reward penalty, theory about model-based rollouts; OpenReview: 8, 6, 6; ExpEnv: mujoco
Yao Mu, Yuzheng Zhuang, Bin Wang, Guangxiang Zhu, Wulong Liu, Jianyu Chen, Ping Luo, Shengbo Eben Li, Chongjie Zhang, Jianye HAO; Key: extension of dreamer, prediction-reliability weight; OpenReview: 6, 6, 6, 6; ExpEnv: deepmind control suite
Rahul Kidambi, Jonathan Chang, Wen Sun; Key: imitation learning from observations alone, mbrl; OpenReview: 6, 6, 6, 4; ExpEnv: cartpole, mujoco
Hung Le, Thommen Karimpanal George, Majid Abdolshah, Truyen Tran, Svetha Venkatesh; Key: model-based, episodic control; OpenReview: 7, 7, 6, 6; ExpEnv: 2D maze navigation, cartpole, mountainCar and lunarlander, atari, 3D navigation: gym-miniworld
Mingde Zhao, Zhen Liu, Sitao Luan, Shuyuan Zhang, Doina Precup, Yoshua Bengio; Key: mbrl, set representation; OpenReview: 7, 7, 7, 6; ExpEnv: MiniGrid-BabyAI framework
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, Yang Gao; Key: muzero, self-supervised consistency loss; OpenReview: 7, 7, 7, 5; ExpEnv: atrai 100k, deepmind control suite
Julian Schrittwieser, Thomas K Hubert, Amol Mandhane, Mohammadamin Barekatain, Ioannis Antonoglou, David Silver; Key: muzero, reanalyse, offline; OpenReview: 8, 8, 7, 6; ExpEnv: atrai dataset, deepmind control suite dataset
Gregory Farquhar, Kate Baumli, Zita Marinho, Angelos Filos, Matteo Hessel, Hado van Hasselt, David Silver; Key: new model learning way; OpenReview: 7, 7, 7, 6; ExpEnv: tabular MDP, Sokoban, atari
Christopher Grimm, Andre Barreto, Gregory Farquhar, David Silver, Satinder Singh; Key: value equivalence, value-based planning, muzero; OpenReview: 8, 7, 7, 6; ExpEnv: four rooms, atari
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, Tengyu Ma; Key: model-based, offline; OpenReview: None; ExpEnv: d4rl dataset, halfcheetah-jump and ant-angle
Sihyun Yu, Sungsoo Ahn, Le Song, Jinwoo Shin; Key: model-based, offline; OpenReview: 7, 6, 6; ExpEnv: design-bench
Jianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu, Siyuan Li, Chongjie Zhang; Key: model-based, offline; OpenReview: 7, 6, 6, 5; ExpEnv: d4rl dataset
Xiong-Hui Chen, Yang Yu, Qingyang Li, Fan-Ming Luo, Zhiwei Tony Qin, Shang Wenjie, Jieping Ye; Key: model-based, offline; OpenReview: 6, 6, 6, 4; ExpEnv: d4rl dataset
Toru Hishinuma, Kei Senda; Key: model-based, offline, off-policy evaluation; OpenReview: 7, 6, 6, 6; ExpEnv: pendulum, d4rl dataset
Weitong Zhang, Dongruo Zhou, Quanquan Gu; Key: learning theory, model-based reward-free RL, linear function approximation; OpenReview: 6, 6, 5, 5; ExpEnv: None
Kefan Dong, Jiaqi Yang, Tengyu Ma; Key: learning theory, model-based bandit RL, nonlinear function approximation; OpenReview: 7, 7, 7, 6; ExpEnv: None
Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, Deepak Pathak; Key: unsupervised goal reaching, goal-conditioned RL; OpenReview: 6, 6, 6, 6, 6; ExpEnv: walker, quadruped, bins, kitchen
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, Shixiang Gu; Key: model-based, behavior cloning (warmup), trpo; OpenReview: 8, 7, 7, 5; ExpEnv: d4rl dataset
Brandon Cui, Yinlam Chow, Mohammad Ghavamzadeh; Key: representation learning, model-based soft actor-critic; OpenReview: 6, 6, 6; ExpEnv: planar system, inverted pendulum – swingup, cartpole, 3-link manipulator — swingUp & balance
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, Jimmy Ba; Key: DreamerV2, many tricks(multiple categorical variables, KL balancing, etc); OpenReview: 9, 8, 5, 4; ExpEnv: atari
Stephen Tian, Suraj Nair, Frederik Ebert, Sudeep Dasari, Benjamin Eysenbach, Chelsea Finn, Sergey Levine; Key: goal-reaching task, dynamics learning, distance learning (goal-conditioned Q-function); OpenReview: 7, 7, 7, 7; ExpEnv: sawyer, door sliding
Arthur Argenson, Gabriel Dulac-Arnold; Key: model-based, offline; OpenReview: 8, 7, 5, 5; ExpEnv: RL Unplugged(RLU), d4rl dataset
Justin Fu, Sergey Levine; Key: model-based, offline; OpenReview: 8, 6, 6; ExpEnv: design-bench
Jessica B. Hamrick, Abram L. Friesen, Feryal Behbahani, Arthur Guez, Fabio Viola, Sims Witherspoon, Thomas Anthony, Lars Buesing, Petar Veličković, Théophane Weber; Key: discussion about planning in MuZero; OpenReview: 7, 7, 6, 5; ExpEnv: atari, go, deepmind control suite
Byung-Jun Lee, Jongmin Lee, Kee-Eung Kim; Key: Representation Balancing MDP, model-based, offline; OpenReview: 7, 7, 7, 6; ExpEnv: d4rl dataset
Balázs Kégl, Gabriel Hurtado, Albert Thomas; Key: mixture density nets, heteroscedasticity; OpenReview: 7, 7, 7, 6, 5; ExpEnv: acrobot system
Brandon Trabucco, Aviral Kumar, Xinyang Geng, Sergey Levine; Key: conservative objective model, offline mbrl; ExpEnv: design-bench
Çağatay Yıldız, Markus Heinonen, Harri Lähdesmäki; Key: continuous-time; ExpEnv: pendulum, cartPole and acrobot
Oleh Rybkin, Chuning Zhu, Anusha Nagabandi, Kostas Daniilidis, Igor Mordatch, Sergey Levine; Key: latent space collocation; ExpEnv: sparse metaworld tasks
David A Bruns-Smith; Key: worst-case bounds; ExpEnv: ope-tools
Matteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez, Simon Schmitt, Laurent Sifre, Theophane Weber, David Silver, Hado van Hasselt; Key: value equivalence; ExpEnv: atari
Sherjil Ozair, Yazhe Li, Ali Razavi, Ioannis Antonoglou, Aäron van den Oord, Oriol Vinyals; Key: VQVAE, MCTS; ExpEnv: chess datasets, DeepMind Lab
Yuda Song, Wen Sun; Key: sample complexity, kernelized nonlinear regulators, linear MDPs; ExpEnv: mountain car, antmaze, mujoco
Tung Nguyen, Rui Shu, Tuan Pham, Hung Bui, Stefano Ermon; Key: temporal predictive coding with a RSSM, latent space; ExpEnv: deepmind control suite
Ying Fan, Yifei Ming; Key: regret bound of psrl, mpc; ExpEnv: continuous cartpole, pendulum swingup, mujoco
Qinghua Liu, Tiancheng Yu, Yu Bai, Chi Jin; Key: learning theory, multi-agent, model-based self play, two-player zero-sum Markov games; ExpEnv: None
Yuan Pu, Yazhe Niu, Zhenjie Yang, Jiyuan Ren, Hongsheng Li, Yu Liu TMLR2025; Key: world model, MCTS, model-based reinforcement learning, transformer, latent planning, multitask learning; ExpEnv: Atari, DMControl, VisualMatch
Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, Zhaoxiang Zhang CVPR 2024; Key: AutoDrive world modeling; ExpEnv: nuScenes
Chen Min, Dawei Zhao, Liang Xiao, Jian Zhao, Xinli Xu, Zheng Zhu, Lei Jin, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai CVPR 2024; Key: AutoDrive world modeling; ExpEnv: nuScenes, OpenScene
Marc Rigter, Jun Yamada, Ingmar Posner Arxiv 2023; Key: Diffusion model, world model; ExpEnv: deepmind control suite, gridworld
Carlos E. Luis, Alessandro G. Bottero, Julia Vinogradska, Felix Berkenkamp, Jan Peters Arxiv 2023; Key: cumulative rewards uncertainty estimation in MBRL; ExpEnv: mujoco
Thomas Bi, Raffaello D'Andrea. Arxiv 2023; Key: Data-Augmented, DreamerV3; ExpEnv: Real-World Labyrinth Game
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap. Arxiv 2023; Key: DreamerV3, scaling property to world model; ExpEnv: deepmind control suite, atari, DMLab, minecraft
Chuming Li, Ruonan Jia, Jiawei Yao, Jie Liu, Yinmin Zhang, Yazhe Niu, Yaodong Yang, Yu Liu, Wanli Ouyang. IJCAI Workshop 2023; Key: extended policy improvement, model regularization, planning theorem; ExpEnv: mujoco
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
matiassingers/awesome-readme
A curated list of awesome READMEs