arXiv Science⌕ Search

arXiv subjects

Alexander T Toshev

Publications and source records attributed to Alexander T Toshev.

2 recordsLinked to original sources

Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models

Training agents that generalize to unseen, stateful environments requires a massive dataset of state-changing trajectories covering a vast and diverse set of APIs. However, scaling this broad supervision is severely bottlenecked by the immense effort required to implement and populate fully-executable environments across a broad spectrum of domains. To bypass this barrier, we introduce a data generation pipeline that decouples data synthesis from environment construction by leveraging LLMs as digital world models. Starting from only a list of broad domain names, our automated pipeline synthesizes diverse APIs and tasks. To produce trajectories, a teacher agent iteratively solves these tasks while an LLM simulator dynamically tracks state and provides coherent API responses on-the-fly. Finally, an automated judge filters the trajectories for quality. Fine-tuning on our broad synthetic dataset yields significant performance gains on AppWorld and OfficeBench, two challenging stateful benchmarks featuring environments completely unseen during training. These results establish our LLM world model-based synthesis approach as a highly scalable path for training generalizable, stateful API-calling agents.

cs.AI↗

Multimodal Autoregressive Pre-training of Large Vision Encoders

We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.

cs.CV↗