arXiv ScienceSearch

arXiv subjects

Adrian Collister

Publications and source records attributed to Adrian Collister.

8 recordsLinked to original sources

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a multi-embodiment Vision-Language-Action (VLA) model, and Gemini Robotics-ER 1.5, a state-of-the-art Embodied Reasoning (ER) model. We are bringing together three major innovations. First, Gemini Robotics 1.5 features a novel architecture and a Motion Transfer (MT) mechanism, which enables it to learn from heterogeneous, multi-embodiment robot data and makes the VLA more general. Second, Gemini Robotics 1.5 interleaves actions with a multi-level internal reasoning process in natural language. This enables the robot to "think before acting" and notably improves its ability to decompose and execute complex, multi-step tasks, and also makes the robot's behavior more interpretable to the user. Third, Gemini Robotics-ER 1.5 establishes a new state-of-the-art for embodied reasoning, i.e., for reasoning capabilities that are critical for robots, such as visual and spatial understanding, task planning, and progress estimation. Together, this family of models takes us a step towards an era of physical agents-enabling robots to perceive, think and then act so they can solve complex multi-step tasks.

cs.RO

Proc4Gem: Foundation models for physical agency through procedural generation

In robot learning, it is common to either ignore the environment semantics, focusing on tasks like whole-body control which only require reasoning about robot-environment contacts, or conversely to ignore contact dynamics, focusing on grounding high-level movement in vision and language. In this work, we show that advances in generative modeling, photorealistic rendering, and procedural generation allow us to tackle tasks requiring both. By generating contact-rich trajectories with accurate physics in semantically-diverse simulations, we can distill behaviors into large multimodal models that directly transfer to the real world: a system we call Proc4Gem. Specifically, we show that a foundation model, Gemini, fine-tuned on only simulation data, can be instructed in language to control a quadruped robot to push an object with its body to unseen targets in unseen real-world environments. Our real-world results demonstrate the promise of using simulation to imbue foundation models with physical agency. Videos can be found at our website: https://sites.google.com/view/proc4gem

cs.RO

Scaling Instructable Agents Across Many Simulated Worlds

Building embodied AI systems that can follow arbitrary language instructions in any 3D environment is a key challenge for creating general AI. Accomplishing this goal requires learning to ground language in perception and embodied actions, in order to accomplish complex tasks. The Scalable, Instructable, Multiworld Agent (SIMA) project tackles this by training agents to follow free-form instructions across a diverse range of virtual 3D environments, including curated research environments as well as open-ended, commercial video games. Our goal is to develop an instructable agent that can accomplish anything a human can do in any simulated 3D environment. Our approach focuses on language-driven generality while imposing minimal assumptions. Our agents interact with environments in real-time using a generic, human-like interface: the inputs are image observations and language instructions and the outputs are keyboard-and-mouse actions. This general approach is challenging, but it allows agents to ground language across many visually complex and semantically rich environments while also allowing us to readily run agents in new environments. In this paper we describe our motivation and goal, the initial progress we have made, and promising preliminary results on several diverse research environments and a variety of commercial video games.

cs.RO

Human-Timescale Adaptation in an Open-Ended Task Space

Foundation models have shown impressive adaptation and scalability in supervised and self-supervised learning problems, but so far these successes have not fully translated to reinforcement learning (RL). In this work, we demonstrate that training an RL agent at scale leads to a general in-context learning algorithm that can adapt to open-ended novel embodied 3D problems as quickly as humans. In a vast space of held-out environment dynamics, our adaptive agent (AdA) displays on-the-fly hypothesis-driven exploration, efficient exploitation of acquired knowledge, and can successfully be prompted with first-person demonstrations. Adaptation emerges from three ingredients: (1) meta-reinforcement learning across a vast, smooth and diverse task distribution, (2) a policy parameterised as a large-scale attention-based memory architecture, and (3) an effective automated curriculum that prioritises tasks at the frontier of an agent's capabilities. We demonstrate characteristic scaling laws with respect to network size, memory length, and richness of the training task distribution. We believe our results lay the foundation for increasingly general and adaptive RL agents that perform well across ever-larger open-ended domains.

cs.LG

Learning Robust Real-Time Cultural Transmission without Human Data

Cultural transmission is the domain-general social skill that allows agents to acquire and use information from each other in real-time with high fidelity and recall. In humans, it is the inheritance process that powers cumulative cultural evolution, expanding our skills, tools and knowledge across generations. We provide a method for generating zero-shot, high recall cultural transmission in artificially intelligent agents. Our agents succeed at real-time cultural transmission from humans in novel contexts without using any pre-collected human data. We identify a surprisingly simple set of ingredients sufficient for generating cultural transmission and develop an evaluation methodology for rigorously assessing it. This paves the way for cultural evolution as an algorithm for developing artificial general intelligence.

cs.LG

Halo-model signatures from 380,000 SDSS Luminous Red Galaxies with photometric redshifts

We analyze the small-scale clustering in "MegaZ-LRG", a large photometric-redshift catalogue of Luminous Red Galaxies extracted from the imaging dataset of the Sloan Digital Sky Survey. MegaZ-LRG, presented in a companion paper, spans the redshift range 0.4 < z < 0.7 with an r.m.s. redshift error dz ~ 0.03(1+z), covering 5,914 deg^2 to map out a total cosmic volume 2.5 h^-3 Gpc^3. In this study we use 380,000 photometric redshifts to measure significant deviations from the canonical power-law fit to the angular correlation function in a series of narrow redshift slices, in which we construct volume-limited samples. These deviations are direct signatures of the manner in which these galaxies populate the underlying network of dark matter haloes. We cleanly delineate the separate contributions of the "1-halo" and "2-halo" clustering terms and fit our measurements by parameterizing the halo occupation distribution N(M) of the galaxies. Our results are successfully fit by a "central" galaxy contribution with a "soft" transition from zero to one galaxies, combined with a power-law "satellite" galaxy component, the slope of which is a strong function of galaxy luminosity. The large majority of galaxies are classified as central objects of their host dark matter haloes rather than satellites in more massive systems. The effective halo mass of MegaZ-LRG galaxies lies in the range log_10 (M_eff / h^-1 M_sol) = 13.61 - 13.8 (increasing with redshift, assuming large-scale normalization sigma_8 = 0.8) for corresponding number densities in the range n_g = 5.03 - 0.56 x 10^-4 h^3 Mpc^-3. Our results confirm the usefulness of the halo model for gaining physical insight into the patterns of galaxy clustering.

astro-ph

MegaZ-LRG: A photometric redshift catalogue of one million SDSS Luminous Red Galaxies

We describe the construction of MegaZ-LRG, a photometric redshift catalogue of over one million luminous red galaxies (LRGs) in the redshift range 0.4 < z < 0.7 with limiting magnitude i < 20. The catalogue is selected from the imaging data of the Sloan Digital Sky Survey Data Release 4. The 2dF-SDSS LRG and Quasar (2SLAQ) spectroscopic redshift catalogue of 13,000 intermediate-redshift LRGs provides a photometric redshift training set, allowing use of ANNz, a neural network-based photometric-redshift estimator. The rms photometric redshift accuracy obtained for an evaluation set selected from the 2SLAQ sample is sigma_z = 0.049 averaged over all galaxies, and sigma_z = 0.040 for a brighter subsample (i < 19.0). The catalogue is expected to contain ~5 per cent stellar contamination. The ANNz code is used to compute a refined star/galaxy probability based on a range of photometric parameters; this allows the contamination fraction to be reduced to 2 per cent with negligible loss of genuine galaxies. The MegaZ-LRG catalogue is publicly available on the World Wide Web from http://www.2slaq.info .

astro-ph

Cosmological baryonic and matter densities from 600,000 SDSS Luminous Red Galaxies with photometric redshifts

We analyze MegaZ-LRG, a photometric-redshift catalogue of Luminous Red Galaxies (LRGs) based on the imaging data of the Sloan Digital Sky Survey (SDSS) 4th Data Release. MegaZ-LRG, presented in a companion paper, contains 10^6 photometric redshifts derived with ANNz, an Artificial Neural Network method, constrained by a spectroscopic sub-sample of 13,000 galaxies obtained by the 2dF-SDSS LRG and Quasar (2SLAQ) survey. The catalogue spans the redshift range 0.4 < z < 0.7 with an r.m.s. redshift error ~ 0.03(1+z), covering 5,914 deg^2 to map out a total cosmic volume 2.5 h^-3 Gpc^3. In this study we use the most reliable 600,000 photometric redshifts to present the first cosmological parameter fits to galaxy angular power spectra from a photometric redshift survey. Combining the redshift slices with appropriate covariances, we determine best-fitting values for the matter and baryon densities of Omega_m h = 0.195 +/- 0.023 and Omega_b/Omega_m = 0.16 +/- 0.036 (with the Hubble parameter h = 0.75 and scalar index of primordial fluctuations n = 1 held fixed). These results are in agreement with and independent of the latest studies of the Cosmic Microwave Background radiation, and their precision is comparable to analyses of contemporary spectroscopic-redshift surveys. We perform an extensive series of tests which conclude that our power spectrum measurements are robust against potential systematic photometric errors in the catalogue. We conclude that photometric-redshift surveys are competitive with spectroscopic surveys for measuring cosmological parameters in the simplest vanilla models. Future deep imaging surveys have great potential for further improvement, provided that systematic errors can be controlled.

astro-ph