arXiv Science⌕ Search

arXiv · 2610.07941

UltraDiff: Differentiable Ray Tracing in Ultrasound for Shape Optimization

Abstract

Physically-based differentiable rendering enables gradient-based optimization of scene parameters by matching rendered images to measurements, but has so far mainly focused on light transport. We extend this paradigm to medical ultrasound, where image formation resembles transient rendering: echoes are binned by time-of-flight rather than projected onto an image plane. We present UltraDiff, a modular framework for differentiable ultrasound ray tracing. UltraDiff formulates ultrasound image formation as a path-space integral, gated by travel time between the transducer and tissue interfaces, and derives a Monte Carlo estimator of both the forward model and its gradients with respect to scene parameters. We demonstrate this on an inverse geometry estimation: starting from a sphere, an SDF is optimized until simulated echoes match measured ones, recovering vertebral surfaces from simulated B-mode sweeps and from a real robotic acquisition of a spine phantom. Unlike state-of-the-art ultrasound shape reconstruction methods, which rely on pre-segmented images, our approach operates unsupervised on B-mode images through analysis-by-synthesis, while achieving competitive geometric accuracy. Implemented on top of Mitsuba 3, UltraDiff brings differentiable path tracing to a new sensing modality and provides a foundation for inverse problems in acoustic imaging.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Felix Duelmer, Magdalena Wysocki, Nassir Navab, Mohammad Farid Azampour. 2026-10-06. UltraDiff: Differentiable Ray Tracing in Ultrasound for Shape Optimization. https://doi.org/10.1145/3829339.3847840

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Neuroll: Real-Time Neural Strand-Based Hair Simulation via Simulator-in-the-Loop Unrolling

Time integration has been the cornerstone of physics-based animation that enables the simulation of complex interactions between rigid and deformable objects, including the motion of hair. Despite recent advances with optimized time integration that enabled thousands of hair strands to be simulated in real time, achieving the same performance on commodity hardware remains infeasible due to the computational demands of resolving complex dynamics and interactions between thousands of individual strands. With the rise of learning-based techniques, the offload of time integration to neural networks helps to achieve significant performance gains, making these approaches suitable for real-time applications such as gaming and virtual avatars. However, state-of-the-art neural techniques tend to produce less physically plausible motion and oftentimes fail to generalize to out-of-distribution scenarios. Inspired by classical time integrators, we design a neural counterpart that mirrors their input-output formulation -- taking previous hair states, material stiffness, and collision geometry as the inputs for the neural time integrator, which is then trained via a self-supervised, simulator-in-the-loop method with randomized unrolling horizons. By formulating training in each strand's local coordinate frame, we obtain a network that generalizes across multiple dimensions, including hairstyle, material property, body motion, and body type. Our method inherits the benefits of a strand-based neural simulator, and hence is density-independent, lightweight, memory-efficient, and performant. Our neural hair integrator produces stable long-horizon rollouts and can be naturally extended to support quasi-static simulation simply by resetting hair states.

cs.GR↗

PrimitiveCAD: An LLM-Based Point-to-CAD Reconstruction with Primitive-Aware Tokenization and Operation Alignment

Large-model-based point-to-CAD generation holds immense potential for advancing industrial design and enhancing 3D modeling efficiency. However, most existing methods approach the problem as a general point-cloud encoding and token prediction task, neglecting the tokenization and supervision specifically for CAD-related primitives. As a result, these methods often struggle to accurately reconstruct the intricate primitive structures. To address this limitation, we propose PrimitiveCAD, a novel multi-stage paradigm for point-to-CAD reconstruction that enhances the geometric accuracy of generated CAD models while better preserving critical geometric features. First, we introduce a primitive-aware point cloud tokenization model, enabling the system to learn more robust geometric representations from CAD point clouds. Next, we perform supervised finetuning on a large language model (LLM) and introduce an operation alignment loss to align key CAD operation frequencies, thereby improving the preservation of global shape features. Finally, we incorporate reinforcement learning (RL) and introduce a feature-line alignment reward to further reduce stochasticity and enhance the fine-grained preservation of geometric features. Experiments on the DeepCAD and Fusion360 datasets show that our method achieves state-of-the-art performance in code validity, geometric accuracy, and geometric feature preservation.

cs.GR↗

Local Content-Style Control for Diffusion-based Image Stylization

Image stylization with latent-diffusion models entangles two independently refined axes: what a region depicts and how it is depicted. Such pipelines expose only global controls, yet professional retouching demands deliberate, region-specific control. We lift two conditioning weights already present in a ControlNet + IP-Adapter stylization pipeline from global scalars to per-location spatial maps, yielding local, per-axis control of content and style in a single generative pass. Because the two weights act on disjoint pathways, adjusting them independently spans a 2x2 retouching vocabulary, from free regeneration to identity preservation. We validate that edits stay confined to the retouched region and that each weight predominantly steers its own axis. Our approach requires no retraining and drops unchanged into any such pipeline.

cs.GR↗