arXiv ScienceSearch

arXiv · 2609.02746

HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design

Abstract

Polymeric materials are central to modern technologies, with applications ranging from energy to health and transportation. Although AI has made significant advances in materials discovery, the hierarchical structure of polymers across multiple length scales makes them inherently difficult to represent in a unified and physically meaningful way. Here we introduce HiPoly, a polymer-native AI framework that processes complete polymer descriptions through a three-level hierarchical graph architecture built on the G2RINS representation. HiPoly encodes stochastic inter-monomer connectivity, composition, and molecular weight directly within its architecture, using physically motivated design principles that mirror the multi-scale nature of polymeric systems. The framework establishes an end-to-end AI-driven workflow from experimental formulation data to property prediction, generative molecular design, and physics-based validation through molecular simulations, all unified by a single polymer representation. We demonstrate state-of-the-art prediction accuracy for thermophysical properties of multi-component polymer systems, with ablation studies confirming that each hierarchical design choice contributes independently to model performance. As an example, the generative design pathway is applied here to the discovery of sustainable alternatives to persistent fluorinated polymers, where it is possible to identify and independently validate PFAS-free candidates with target surface-energy properties. This work demonstrates how polymer-native AI can accelerate discovery by linking representation, prediction, and design across complex polymer chemistries.

Explore related subjects

Keep this discovery

BibTeXRIS

Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Daryna Safarian, Ming Han, Juan J. de Pablo. 2026-09-02. HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design. https://arxiv.org/abs/2609.02746

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Unraveling the PFAS helix: A statistical approach

The extreme persistence of per- and polyfluoroalkyl substances (PFAS) in the environment is rooted in their three-dimensional molecular conformation. The helical twist adopted by perfluoroalkyl chains due to hyperconjugation involving their recalcitrant C-F bonds governs their resistance to degradation, yet a quantitative, continuous metric for helicity has remained absent. Here we introduce a data-driven framework that quantifies backbone helicity using three statistical descriptors: void-state (binary occupancy) spatial autocorrelations, local backbone principal component analysis, and persistent homology. We further validate each statistical descriptor against geometric dihedral benchmarks and DFT-calculated Vibrational Circular Dichroism (VCD) spectra. The framework is initially established on a homologous series of perfluorocarboxylic acids (FC2 through FC16) and their hydrogenated analogues, then extended across perfluorosulfonic acids, fluorotelomer alcohols, polyfluoroalkyl hexanoic acid analogues, and longer-chain PFOA and PFOS analogues. Agreement between statistical, geometric, and spectroscopic definitions of helicity is established for PFCAs and extended to structurally distinct subfamilies, including PFOA analogues of varying fluorine content and perfluorosulfonic acids, demonstrating that the descriptors are robust to changes in headgroup chemistry and fluorine substitution pattern. The framework provides a chemistry-agnostic toolset for incorporating backbone conformation into predictive models of PFAS environmental fate and degradation reactivity.

physics.chem-ph

Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the frozen cameras and re-solving cross-view correspondence recovers exactly the answer the images support. We prove the ceiling fails only through an enumerable set of projection coincidences and certify that set empty on 2,160 rendered crystal structures, so every point of a model's deficit belongs to the model. Across fourteen vision-language models, supplying exact geometry as text lifts every model yet closes under half the gap for thirteen, while a supervised vision model with no language component reads the same images at 0.8952, above every vision-language model. The instrument exposes extraction-stage fabrication that downstream accuracy would misattribute to reasoning, yields camera-placement rules for benchmark builders, and transfers to any benchmark with an invertible forward rendering.

cs.CV

Polarizable atomic multipoles for learning long-range electrostatics

Long-range electrostatics and polarization remain central obstacles to extending machine learning interatomic potentials (MLIPs) to ionic, polar, and interfacial systems. Here we introduce a semi-local framework for learning electrostatics from energies and forces using polarizable atomic multipoles. Local equivariant descriptors predict environment-dependent latent monopoles, dipoles, and quadrupoles, while residual non-local charge transfer and polarization are captured by non-self-consistent linear response in induced charges and dipoles. Across four diverse benchmarks and four short-range MLIP architectures, the multipole hierarchy and response terms systematically improve potential energy surface accuracy, with the largest gains in systems where long-range effects are essential. More importantly, physically meaningful electrical responses emerge without direct supervision. The learned latent multipoles yield accurate Born effective charge tensors and infrared spectra in close agreement with experiments. The induced-dipole extension introduces new capabilities: it predicts polarizabilities and thereby enables semi-quantitative Raman spectra for bulk water and hybrid MAPbI$_3$ perovskite, as well as the essential features of the surface-specific vibrational sum-frequency generation spectrum at the water-air interface. In ferroelectric HfO$_2$, the predicted electrical response also captures LO-TO splitting and polarization switching. This systematically improvable, physically transparent framework enables MLIPs trained on standard energy and force labels to predict polarization-sensitive observables.

cond-mat.mtrl-sci