arXiv Science⌕ Search

arXiv · 2609.35639

GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation

Abstract

Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks cover fluids, deformable solids, and granular materials, from individual simulation operators to complete simulators. Agents write, compile, test, and optimize GPU code with access to a NVIDIA GPU under fixed time budgets. We report pass rates and runtime performance relative to expert-optimized reference implementations. In a single-attempt evaluation of six frontier model-harness pairs, the two strongest pass all 50 tasks, but even the fastest reaches at least 0.9 the reference speed on only 22% of them, and no submission is more than 5% faster than the reference. The largest gaps arise in collision detection, constraint solving, and iterative solvers. GPUPhysBench brings physical simulation workloads to coding-agent evaluation, testing both the ability to implement numerical methods correctly and the ability to make them run efficiently.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuchen Sun, Jinjin He, Sinan Wang, Bo Zhu. 2026-09-28. GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation. https://arxiv.org/abs/2609.35639

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications

Large Language Models (LLMs) are increasingly deployed in complex multi-agent applications that rely on external function calls. This workload creates severe performance challenges for the KV Cache: spatial contention leads to the eviction of critical agents' caches and temporal underutilization leaves the cache of agents stalled on long-running function calls idling in GPU memory. We present TokenCake, a KV-Cache-centric serving framework that bridges this gap by co-optimizing scheduling and memory management through an agent-aware design. TokenCake's Temporal Scheduler employs an event-driven, opportunistic policy to proactively offload idle KV Caches during function calls and uses predictive uploading to hide data transfer latency. TokenCake's Spatial Scheduler uses dynamic memory partitioning, guided by a hybrid priority metric combining graph structure and runtime state, to reserve GPU memory for critical-path agents. Our evaluation on representative multi-agent benchmarks shows that TokenCake reduces end-to-end latency by up to 47.06% in memory-constrained settings and improves effective GPU utilization by 16.9 percentage points compared to vLLM. TokenCake is publicly available at https://github.com/pkulemonade/TokenCake.

cs.DC↗

SpSYRK: Half the Work in Distributed Sparse Matrix Multiplication

The symmetric rank-$k$ update (SYRK), $C = AA^\top$, computes the dot product of each pair of rows of $A$, producing the Gram matrix $C$. Its sparse variant underpins similarity search in machine learning, graph analytics, and genomics, including Jaccard similarity on datasets too large for a single node. Despite the symmetry in its inputs and outputs, existing distributed sparse matrix multiplication algorithms such as Sparse SUMMA treat sparse SYRK as generic multiplication, computing the full output and materializing the explicit transpose even when the calling application uses only one triangle. Prior distributed $AA^\top$ computations in similarity search and genome assembly inherit this overhead from the underlying SpGEMM. This paper presents SpSYRK and CommSpSYRK, two distributed sparse SYRK algorithms that exploit symmetry. The first, SpSYRK, partitions the off-diagonal blocks of the output between the upper and lower triangular regions of the process grid and computes only the lower-triangular part of each diagonal block, halving per-process computation compared with state-of-the-art distributed SpGEMM. The second, CommSpSYRK, further reorders communication to avoid forming $A^\top$, which reduces per-process communication volume. On 32 nodes of the Perlmutter supercomputer, SpSYRK achieves a 2$\times$ speedup over an optimized Sparse SUMMA on matrices where local multiplication dominates the runtime; the advantage narrows on communication-bound inputs, a dependence that the cost model predicts from the arithmetic intensity. CommSpSYRK fixes this and consistently achieves superior scaling at high process counts. The approach is a drop-in replacement for any application computing $C = AA^\top$ via a distributed SpGEMM routine, and its triangular output can be consumed directly by subsequent operations, reducing both computation and memory footprint.

cs.DC↗

Brain API: An Intent-Aware Control Plane for Policy-Governed Agentic Systems

Contemporary cloud and distributed systems expose control through resource-centric abstractions: services, deployments, network flows, execution graphs. Agentic and tool-augmented systems have meanwhile shifted application logic toward intent-driven, adaptive execution. Existing control planes, workflow engines and service meshes lack abstractions for intent-level decision governance: they cannot represent high-level goals as first-class control objects, cannot enforce policy over the mapping from intent to execution plan, and cannot produce auditable records of why one execution path was chosen over its alternatives. Control logic is therefore embedded in application code, leaving systems brittle, opaque and hard to govern. We propose Brain API, an intent-aware control plane for policy-governed agentic systems. Its central contribution is the decision artifact: a durable, versioned, auditable record of how an intent became an executable plan, capturing which policies applied, which capabilities were evaluated, which alternatives were rejected, and why. A motivating use case is agentic datasets: datasets participating as policy-governed capabilities under residency, compliance and cost constraints. We evaluate a prototype of the decision layer against two external policy corpora we did not author. On the OPA Gatekeeper constraint library it agrees with the library's own published verdicts on 42 of 42 encodable cases, 19 admit and 23 deny. On Cedar example policies, labeled by differential testing against its reference implementation, a deliberately dissimilar domain exposed three defects in our model, including a default-allow assumption that would have inverted every authorization policy. The evaluation covers policy filtering and selection; candidate generation, context signals, ranking and plan synthesis are not measured, nor is decision latency under load.

cs.DC↗