arXiv · 2309.01933
Provably safe systems: the only path to controllable AGI
Abstract
We describe a path to humanity safely thriving with powerful Artificial General Intelligences (AGIs) by building them to provably satisfy human-specified requirements. We argue that this will soon be technically feasible using advanced AI for formal verification and mechanistic interpretability. We further argue that it is the only path which guarantees safe controlled AGI. We end with a list of challenge problems whose solution would contribute to this positive outcome and invite readers to join in this work.
Explore related subjects
Keep this discovery
Max Tegmark, Steve Omohundro. 2023-09-05. Provably safe systems: the only path to controllable AGI. https://arxiv.org/abs/2309.01933
Cite the original work for its findings. Save a collection to share your selection of sources.