Video-game gameplay datasets for agent training
Recent Hub datasets of video game gameplay (frames/video + actions, replays, trajectories) for imitation learning and game-playing agents.
Updated • 793 • 22Note CS:GO deathmatch video with frame-aligned action labels + metadata. A canonical behavioral-cloning / imitation-learning FPS benchmark (basis of the 'Counter-Strike from pixels' work). ~690 downloads, 17 likes.
RekaAI/CS2-10k
Preview • Updated • 125k • 34Note Large egocentric Counter-Strike 2 video dataset from professional matches, 600+ hours of POV footage. Recent (2026-06) and heavily used (27k downloads) for FPS agent/world-model training.
zizhaotong/CrossFPS-train
Updated • 622 • 9Note CrossFPS: multi-game first-person-shooter dataset, ~69k five-second 20fps clips paired with frame-aligned actions. Good for cross-game FPS imitation learning / generalization.
nyu-visionx/solaris-training-dataset
Updated • 14.5k • 3Note SolarisEngine multi-agent Minecraft world-modeling corpus: 12.64M frames at 20 FPS with actions. Large-scale (3.2k downloads) for action-conditioned world models and Minecraft agents.
hlillemark/mc_combined_sa_ma_dataset
Viewer • Updated • 2.23k • 7Note 73 Minecraft trajectories from single- and multi-agent scenarios (collected via GPT-4o / LLM agents). Compact set of action-labeled Minecraft gameplay for agent training.
TESS-Computer/minecraft-vla-stage1
Viewer • Updated • 15.2M • 902 • 3Note Frame-action pairs from 17,886 early-game Minecraft videos, processed with Lumine's 5 FPS method. Formatted for vision-language-action (VLA) game agent pretraining.
p-doom/atari-pong-dataset
Updated • 29 • 1Note 10M frames + actions from the Atari Pong environment (Bellemare et al. ALE). Part of the p-doom per-game Atari suite (Alien, Breakout, Battle Zone, Demon Attack, etc.) for RL / world-model / IL research.
TESS-Computer/atari-vla-stage1-15hz
Viewer • Updated • 1.34M • 227Note 1.3M human gameplay demonstrations across Atari games (sourced from Zenodo), formatted for vision-language-action training. Human demos rather than RL-agent rollouts.
DylanRiden/smb-worldmodel-data
Updated • 19Note 118,166 frames from 8 Super Mario Bros TAS playthroughs as NumPy .npz with action labels. Small, clean action-conditioned dataset for platformer world models / imitation learning.
Karajan42/dreamer4-flappybird
Updated • 66Note Dreamer4-formatted Flappy Bird gameplay: frame shards paired with action and reward signals. Ready-to-use for model-based RL / world-model (Dreamer) experiments.
dasgringuen/assettoCorsaGym
Preview • Updated • 586 • 6Note Assetto Corsa racing sim: 64M steps from human and Soft Actor-Critic policies in an autonomous-racing Gym. Mixes human demonstrations with RL rollouts for driving-agent training.
Ethosoft/ck3-gameplay-mouse-keyboard-dataset
Preview • Updated • 1.07k • 1Note 56 self-recorded Crusader Kings III sessions capturing raw mouse/keyboard events + screen telemetry. Grand-strategy gameplay for behavior cloning / imitation learning (recent, 2026-07).
arelius/nxml-pokemon-legends-za
Viewer • Updated • 234 • 7Note 234 synchronized video episodes of Pokémon Legends: Z-A gameplay with 26-dimensional control signals. Action-conditioned console-game footage for imitation learning.
xxxTEMPESTxxx/PokeDreamer
Updated • 131Note PokeDreamer v2: 200,000 high-resolution RGB trajectories of Pokémon Red played in a PyBoy emulator. Emulator-generated trajectories for world-model / game-agent training.
obaydata/world-model-gameplay-recording
Preview • Updated • 24Note Synchronized, action-conditioned gameplay videos with frame-accurate input logs. General-purpose gameplay recording aimed at action-conditioned world-model training.
wolframko/betty-dota2
Viewer • Updated • 2.26B • 5.42k • 1Note 9,385 professional Dota 2 matches parsed from .dem replays into per-second decision context (hero states, cooldowns, combat events, objectives). For Transformer/RL models of MOBA game state; replay-derived rather than pixel gameplay.
strakammm/generals_io_replays
Viewer • Updated • 18.8k • 77 • 4Note 1v1 Generals.io replays featuring at least one 70+ star player, formatted for reinforcement-learning agent training on this real-time strategy game.
EpicPinkPenguin/procgen
Viewer • Updated • 160M • 452 • 6Note Expert PPO trajectories across all 16 Procgen environments.