arXiv Science⌕ Search

arXiv · 2609.37875

Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design

Abstract

Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder-decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based-encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer-based decoders on $\sim80{,}000$ EBSD-derived microstructure dataset to learn a minimal bottleneck, $z$. The ViT-FMDiT model ($z$=$768$) reconstructs high-fidelity microstructure images (FID $27.86$, MS-SSIM $0.178$), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce MERIDIAN, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, failure-aware feasibility prediction, manifold-aware trust regions, and target-aware acquisition. Within a budget of $160$ simulations, the ViT-FMDiT and MERIDIAN combination yields the best target-driven objective score, reducing the relative target error by $3$--$22\%$ against seven baselines (DANTE, TuRBO, BAxUS, CMA-ES, DDOM, SEIKO, DDPO) on the same decoder.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mahish K. Guru, Mayank Nagar, Ayush vyas, Jan Bohlen, Roland Aydin, Noomane Ben Khalifa. 2026-09-29. Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design. https://arxiv.org/abs/2609.37875

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Hierarchical Reasoning Model

Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language models (LLMs) primarily employ Chain-of-Thought (CoT) techniques, which suffer from brittle task decomposition, extensive data requirements, and high latency. Inspired by the hierarchical and multi-timescale processing in the human brain, we propose the Hierarchical Reasoning Model (HRM), a novel recurrent architecture that attains significant computational depth while maintaining both training stability and efficiency. HRM executes sequential reasoning tasks in a single forward pass without explicit supervision of the intermediate process, through two interdependent recurrent modules: a high-level module responsible for slow, abstract planning, and a low-level module handling rapid, detailed computations. With only 27 million parameters, HRM achieves exceptional performance on complex reasoning tasks using only 1000 training samples. The model operates without pre-training or CoT data, yet achieves nearly perfect performance on challenging tasks including complex Sudoku puzzles and optimal path finding in large mazes. Furthermore, HRM outperforms much larger models with significantly longer context windows on the Abstraction and Reasoning Corpus (ARC), a key benchmark for measuring artificial general intelligence capabilities. These results underscore HRM's potential as a transformative advancement toward universal computation and general-purpose reasoning systems.

cs.AI↗

A memory-based active inference model of DishBrain-like adaptive behaviour

Recent and rapid advances in artificial intelligence (AI) make it increasingly important to understand the foundations of adaptive behaviour in autonomous agents, especially for building safe and efficient systems. While artificial neural networks have dominated the development of AI, recent work has begun to explore living biological neuronal networks as an alternative substrate for computation. These systems promise remarkable data and sample efficiency and rich dynamics, and may also inspire explainable and biologically plausible models. Here, we develop an experiment-informed active inference framework to model decision-making in closed-loop agents that mirror experimental setups using biological neurons. Using a generative model whose dimensions are matched to an experiment protocol, we systematically compare three decision-making schemes within this common generative model. Under matched episode counts (i.e. total data available for learning) to the in-vitro experiment, our simulations show that agents with short memory horizons reach a level of performance close to that of mouse and human cortical cultures (DishBrain platform), whereas longer memory horizons depart from it substantially. Increasing the planning horizon, by contrast, confers no comparable benefit. Because all model parameters are explicit, we can also track the quantities in our generative model that accompany this improvement, such as the risk term and the entropy of the transition and state-action mappings. Together, these results illustrate how active inference offers a formal language for comparing decision-making schemes in similar closed-loop control environments.

cs.AI↗

A Quantitative Study of Sustained Focus in Large Language Models via Repetitive Deterministic Prediction Tasks

We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rate (SAR) scales with output length. Each such task involves the repetition of the same operation $N$ times. Examples of such tasks include letter replacement in letter strings following a given rule, integer addition, and multiplication of string operators in many-body quantum mechanics. If the LLM performs the task by a simple repetition algorithm, the success rate would follow an exponential decay with sequence length. In contrast, our experiments on leading LLMs reveal a crossover that is sharper than exponential: $-\log\mathrm{SAR}$ grows super-linearly with $N$, and accuracy collapses around a characteristic length $N_*$, the accuracy cliff that separates reliable from unreliable generation. The hypothesis of independent per-step errors is rejected for every model and task we studied. The crossover is well described by a double-exponential accumulation law, $\mathrm{SAR}=\exp(-β_0 Nα^{N-1})$, whose crossover scale $N_*$ does not depend on the functional form chosen to fit it. To interpret this behaviour we introduce a minimal effective model in which step-correctness variables interact through dense random couplings and compete with an external field set by the prompt. Solved by direct enumeration, the model reproduces the super-linear error accumulation and the accuracy cliff qualitatively, and it assigns to each model--task pair two interpretable parameters, an intrinsic error rate and an error-accumulation factor.

cs.AI↗