Time-aligned vision, speech and action data for embodied AI
PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI
Paper / source record · 2025-05-19
Research question and approach
Advances in deep generative modeling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactive datasets that reflect the sensory-motor complexity of natural environments. To address this, we present PLAICraft, a novel data collection platform and dataset capturing multiplayer Minecraft interactions across five time-aligned modalities: video, game output audio, microphone input audio, mouse, and keyboard actions.
When this work is relevant
Cite this dataset when using or comparing time-aligned multimodal embodied-agent data. Consult the source for collection protocol, synchronization and permitted use.
Author-written abstract
Advances in deep generative modeling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactive datasets that reflect the sensory-motor complexity of natural environments. To address this, we present PLAICraft, a novel data collection platform and dataset capturing multiplayer Minecraft interactions across five time-aligned modalities: video, game output audio, microphone input audio, mouse, and keyboard actions. Each modality is logged with millisecond time precision, enabling the study of synchronous, embodied behaviour in a rich, open-ended world. The dataset comprises over 10,000 hours of gameplay from more than 10,000 global participants. Alongside the dataset, we provide an evaluation suite for benchmarking model capabilities in object recognition, spatial awareness, language grounding, and long-term memory. PLAICraft opens a path toward training and evaluating agents that act fluently and purposefully in real time, paving the way for truly embodied artificial intelligence.
Abstract source: https://arxiv.org/abs/2505.12707. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.
Citation
BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.
@article{he2025plaicraft,
title={Plaicraft: Large-scale time-aligned vision-speech-action dataset for embodied ai},
author={He, Yingchen and Weilbach, Christian D and Wojciechowska, Martyna E and Zhang, Yuxuan and Wood, Frank},
journal={arXiv preprint arXiv:2505.12707},
year={2025}
}