arXiv · 2608.28884
MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft
Abstract
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.
Explore related subjects
Keep this discovery
Sewoong Lee, Risham Sidhu, Julia Hockenmaier, Yoonhwa Jung. 2026-08-28. MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft. https://arxiv.org/abs/2608.28884
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.