arXiv ScienceSearch

arXiv subjects

Chao Tang

Publications and source records attributed to Chao Tang.

At least 19 recordsLinked to original sources

Cell size and confinement drive asymmetric cell division through a cortical instability

Asymmetric cell division -- in which a mother cell divides into two daughter cells of unequal size -- is a fundamental problem in biology. It is believed that the asymmetry originates from the prior polarization of the mother cell. Here we show that division asymmetry can occur spontaneously even in unpolarized mother cells. Specifically, curvature-dependent active stresses in the cell cortex can lead to this symmetry breaking without any molecular polarity cue if the mother cell is confined within a restricted space. Either reducing the cell size or tightening mechanical confinement triggers the same spontaneous symmetry-breaking instability, in which the contractile ring slips off the equator to yield daughters of unequal volume. In the presence of a polarity cue, this instability cooperates with the cue to program the division asymmetry. The model prediction is compared with the imaging data of C. elegans embryogenesis, in which successive cell divisions in a confined eggshell lead to smaller and smaller cell sizes. The measured division asymmetry indeed increases as the cells shrink, and is further amplified when the embryo is mechanically compressed, both in agreement with the model prediction.

cond-mat.soft

Phonon-Localization-Driven Decoupling of Dual-Channel Transport for Record-Low Intrinsic Lattice Thermal Conductivity

A fundamental bottleneck in pushing the intrinsic lattice thermal conductivity of inorganic crystalline solids to its lowest limit arises from the inherent competition between the particle-like propagation (\(\kappa_{\mathrm{L}}^{\mathrm{P}}\)) and wave-like tunneling (\(\kappa_{\mathrm{L}}^{\mathrm{C}}\)) channels. Herein, we demonstrate that phonon localization provides a robust pathway to decouple the dual-channel transport, achieving record-low \(\kappa_{\mathrm{L}}\) in quasi-1D ternary helical crystals. Despite the structural complexity leading to densely populated phonon branches and thus inducing abundant coherent phonons, the weak interchain interactions and heavy elements compress numerous branches into highly localized, nearly dispersionless flat bands. Such strong localization simultaneously suppresses both the diagonal and off-diagonal components of the group velocity, thereby synergistically suppressing \(\kappa_{\mathrm{L}}^{\mathrm{P}}\) and \(\kappa_{\mathrm{L}}^{\mathrm{C}}\). Taking InSeI as an example, the interchain room-temperature \(\kappa_{\mathrm{L}}^{\mathrm{P}}\) and \(\kappa_{\mathrm{L}}^{\mathrm{C}}\) are 0.145 and 0.053 W/mK, respectively, yielding an ultralow total \(\kappa_{\mathrm{L}}\) of 0.198 W/mK. Weaker interchain interactions further drive the room-temperature \(\kappa_{\mathrm{L}}\) of GaSeI and AlSeI to record lows of 0.086 and 0.089 W/mK, respectively; these values even drop to 0.058 and 0.059 W/mK at 900 K. These findings provide useful insights into exploring the thermal conductivity limit in crystals.

cond-mat.mtrl-sci

CARE: Context-Aware Ranking Evolution with Executable Scoring Programs for Budgeted Reaction Optimization

High-throughput experimentation can evaluate many reaction conditions, yet combinatorial condition spaces still exceed the available experiment budget. This makes experiment selection a sequential decision problem: each new condition must be chosen from limited observations before its outcome is known. LLMs can express task-specific selection logic. A direct recommendation, however, is neither a persistent executable object that can be validated and revised nor an independently auditable decision rule. We introduce CARE, a reference-conditioned controller that separates program synthesis from experiment selection. An LLM writes an executable scoring program that ranks the remaining conditions, while a non-LLM reference policy supplies a numerical candidate and support summary. CARE forms an optional alternative from the program, applies a reference-conditioned intervention gate to compare it with the reference, and records the decision before the selected outcome is revealed. Each new outcome updates controller state and can trigger retention, revision, or regeneration of the active program. This outcome-guided program evolution changes the scoring logic without updating the LLM parameters. In matched offline replay with 30 seeds on eight reaction-optimization tasks, CARE attains the lowest normalized regret, the highest normalized best-so-far AUC, and the highest Top-1% Success@15 among the evaluated methods. These results support using a scoring program written by an LLM as one component of a reference-conditioned optimizer rather than as a standalone experiment selector.

cs.LG

High-Throughput Discovery of Semimetallic Borophenes with Diverse Dirac States Via Transferable Tight-Binding Approach

Borophene has attracted extensive interest due to its structural flexibility and emergent topological electronic states. However, semimetallic borophenes hosting robust Dirac states remain rare among the large number of predicted allotropes. Here, we develop a transferable tight-binding framework for planar borophenes and combine it with a graph- and group-theory-based random generation strategy to perform high-throughput screening of 522 borophene candidates. Eight previously unreported semimetallic borophenes are identified, hosting diverse topological band crossings, including type-I and type-III Dirac cones, Dirac nodal lines, and quadratic nodal points. Notably, quadratic nodal-point semimetals are predicted in borophene for the first time. Symmetry analysis reveals crystalline-symmetry-protected Dirac states, while first-principles calculations confirm their dynamical and thermal stability. These findings establish borophene as a versatile platform for engineering emergent Dirac physics in two dimensions.

cond-mat.mtrl-sci

Towards Customized Multimodal Role-Play

Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while maintaining output consistency across modalities remains largely unexplored. To mitigate this gap, we introduce a new task, Customized Multimodal Role-Play (CMRP). We construct the RoleScape-20 dataset comprising 20 characters, including training and evaluation data that cover persona, stylistic descriptions, visual/expressive cues, and text-image interactions. Building on a unified model, we devise UniCharacter, a two-stage training framework containing Unified Supervised Finetuning (Unified-SFT) and character-specific group relative policy optimization (Character-GRPO). Given only 10 images plus corresponding interaction examples, the model acquires the target character and exhibits coherent persona, style, and visual identity in both generated text and images. This process takes about 100 GPU hours. Experiments on the RoleScape-20 dataset show that the proposed method substantially outperforms prior approaches. Ablation studies further validate the effectiveness of our cross-modal consistency design and few-shot customization strategy. We argue that CMRP, coupled with unified modeling, provides a basis for next-generation characterful and immersive interactive agents.

cs.LG

Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations

Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constrained to narrow object/task sets or rely on prohibitively large-scale data collection to capture real-world variability. In this work, we present an alternative approach, GraspDreamer, a method that leverages human demonstrations synthesized by visual generative models (VGMs) (e.g., video generation models) to enable zero-shot functional grasping without labor-intensive data collection. The key idea is that VGMs pre-trained on internet-scale human data implicitly encode generalized priors about how humans interact with the physical world, which can be combined with embodiment-specific action optimization to enable functional grasping with minimal effort. Extensive experiments on the public benchmarks with different robot hands demonstrate the superior data efficiency and generalization performance of GraspDreamer compared to previous methods. Real-world evaluations further validate the effectiveness on real robots. Additionally, we showcase that GraspDreamer can (1) be naturally extended to downstream manipulation tasks, and (2) can generate data to support visuomotor policy learning.

cs.RO

Easy-IIL: Reducing Human Operational Burden in Interactive Imitation Learning via Assistant Experts

Interactive Imitation Learning (IIL) typically relies on extensive human involvement for both offline demonstration and online interaction. Prior work primarily focuses on reducing human effort in passive monitoring rather than active operation. Interestingly, structured model-based imitation approaches achieve comparable performance with significantly fewer demonstrations than end-to-end imitation learning policies in the low-data regime. However, these methods are typically surpassed by end-to-end policies as the data increases. Leveraging this insight, we propose Easy-IIL, a framework that utilizes off-the-shelf model-based imitation methods as an assistant expert to replace active human operation for the majority of data collection. The human expert only provides a single demonstration to initialize the assistant expert and intervenes in critical states where the task is approaching failure. Furthermore, Easy-IIL can maintain IIL performance by preserving both offline and online data quality. Extensive simulation and real-world experiments demonstrate that Easy-IIL significantly reduces human operational burden while maintaining performance comparable to mainstream IIL baselines. User studies further confirm that Easy-IIL reduces subjective workload on the human expert. Project page: https://sites.google.com/view/easy-iil

cs.RO

Micrometer-scale displacement and thickness sensing using a single terahertz resonant-tunneling diode

Resonant tunneling diodes (RTDs) support room-temperature terahertz (THz) oscillation and simultaneous THz-band detection, enabling compact monostatic THz sensors for practical and cost-effective sensing applications. In this paper, we present a highly integrated 280 GHz-band radar system based on a single RTD that exploits the self-mixing effect to generate a low-frequency interferometric signal. The resulting self-mixing signal is further analyzed from a radar perspective and processed to extract micrometer-scale displacement and thin-film thickness variations. Experimentally, the proposed system demonstrates a minimum detectable displacement of approximately 5 um and quantitatively resolves polymer film thicknesses of 12.5, 25, and 50 um.

physics.optics

Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent

Stochastic gradient descent (SGD) is central to deep learning, yet the dynamical origin of its preference for flatter, more generalizable solutions remains unclear. Here, by analyzing SGD learning dynamics, we identify a nonequilibrium mechanism that governs solution selection during training. Numerical experiments reveal a transient exploratory phase in which SGD trajectories repeatedly escape sharp valleys and migrate toward flatter regions of the loss landscape before becoming confined to a final basin. Using a tractable physical model, we show that SGD noise reshapes the loss landscape into an effective potential that preferentially stabilizes flat solutions. We further uncover a transient freezing mechanism: as training progresses, the flattening landscape suppresses transitions between competing valleys. Stronger SGD noise delays this freezing transition, prolonging the exploratory phase and thereby increasing the probability of convergence to flatter minima. Together, these results provide a unified physical framework connecting learning dynamics, loss-landscape geometry, and generalization, and suggest guiding principles for the design of more effective optimization algorithms.

cs.LG

CTransformer: Deep-transformer-based 3D cell membrane tracking with subcellular-resolved molecular quantification

Deep learning segmentation and fluorescence imaging techniques allow the cellular morphology of living embryos to be constructed spatiotemporally. These development processes involve numerous molecules distributed at the subcellular scale, such as cell adhesion (E-cadherin), which accumulate at cell-cell interfaces to regulate intercellular connection. However, quantifying molecular distributions within specific subcellular regions across the entire embryo, where cell movement and molecular redistribution occur rapidly, is challenging due to the need for simultaneous cell morphology reconstruction and lineage tracing due to photobleaching and phototoxicity. We report a transformer-based pipeline, CTransformer, that establishes a 4D cellular morphology map before the 550-cell (late) stage. CTransformer constructed 4D cellular morphology atlases, reaching 80% accuracy at the 550-cell stage. Through this advanced architecture, we use only one channel to reconstruct cell morphology and achieve cell tracing. With each cell's morphology as a reference, the distribution of specific molecules throughout the cell body and at cell interfaces can be quantitatively measured in another fluorescence channel. We apply this methodology to track E-cadherin during embryonic development of the worm Caenorhabditis elegans, from fertilization to gastrulation. Our results reveal that E-cadherin is tightly regulated across individual embryos, both within single cells and at cell-cell interfaces, displaying an anterior-posterior gradient and cell- and lineage-specific patterns. Furthermore, its spatiotemporal heterogeneity influences cell mechanics and embryonic morphogenesis, helping explain how C. elegans achieves stereotypical developmental patterns at cellular resolution.

physics.bio-ph

HyperAIRI: a plug-and-play algorithm for precise hyperspectral image reconstruction in radio interferometry

The next-generation radio-interferometric (RI) telescopes require imaging algorithms capable of forming high-resolution high-dynamic-range images from large data volumes spanning wide frequency bands. Recently, AIRI, a plug-and-play (PnP) approach taking the forward-backward algorithmic structure (FB), has demonstrated state-of-the-art performance in monochromatic RI imaging by alternating a data-fidelity step with a regularization step via learned denoisers. In this work, we introduce HyperAIRI, its hyperspectral extension, underpinned by learned hyperspectral denoisers enforcing a power-law spectral model. For each spectral channel, the HyperAIRI denoiser takes as input its current image estimate, alongside estimates of its two immediate neighboring channels and the spectral index map, and provides as output its associated denoised image. To ensure convergence of HyperAIRI, the denoisers are trained with a Jacobian regularization enforcing non-expansiveness. To accommodate varying dynamic ranges, we assemble a shelf of pre-trained denoisers, each tailored to a specific dynamic range. At each HyperAIRI iteration, the spectral channels of the target image cube are updated in parallel using dynamic-range-matched denoisers from the pre-trained shelf. The denoisers are also endowed with a spatial image faceting functionality, enabling scalability to varied image sizes. Additionally, we formally introduce Hyper-uSARA, a variant of the optimization-based algorithm HyperSARA, promoting joint sparsity across spectral channels via the $\ell_{2,1}$-norm, also adopting FB. We evaluate HyperAIRI's performance on simulated and real observations. We showcase its superior performance compared to its optimization-based counterpart Hyper-uSARA, CLEAN's hyperspectral variant in WSClean, and the monochromatic imaging algorithms AIRI and uSARA.

astro-ph.IM

Unique Hierarchical Rotational Dynamics Induces Ultralow Lattice Thermal Conductivity in Cyanide-bridged Framework Materials

The pursuit of materials combining light constituent elements with ultralow lattice thermal conductivity ($\kappa_{\mathrm{L}}$) is crucial to advancing technologies like thermoelectrics and thermal barrier coatings, yet it remains a formidable challenge to date. Herein, we achieve ultralow $\kappa_{\mathrm{L}}$ in lightweight cyanide-bridged framework materials (CFMs) through the rational integration of properties such as the hierarchical vibrations exhibited in superatomic structures and rotational dynamics exhibited in perovskites. Unique hierarchical rotation behavior leads to multiple negative peaks in Gr\"uneisen parameters across a wide frequency range, thereby inducing pronounced negative thermal expansion and strong cubic anharmonicity in CFMs. Meanwhile, the synergistic effect between large four-phonon scattering phase space (induced by phonon quasi-flat bands and wide bandgaps) and strong quartic anharmonicity (associated with rotation modes) leads to giant quartic anharmonic scattering rates in these materials. Consequently, the $\kappa_{\mathrm{L}}$ of these CFMs decreases by one to two orders of magnitude compared to the known perovskites or perovskite-like materials with equivalent average atomic masses. For instance, the Cd(CN)$_{2}$, NaB(CN)$_{4}$, LiIn(CN)$_{4}$, and AgX(CN)$_{4}$ (X = B, Al, Ga, In) exhibit ultralow room-temperature $\kappa_{\mathrm{L}}$ values ranging from 0.35 to 0.81 W/mK. This work not only establishes CFMs as a novel and rich platform for studying extreme phonon anharmonicity, but also provides a new paradigm for achieving ultralow thermal conductivity in lightweight materials via the conscious integration of hierarchical and rotational dynamics.

cond-mat.mtrl-sci

MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence

Imitating tool manipulation from human videos offers an intuitive approach to teaching robots, while also providing a promising and scalable alternative to labor-intensive teleoperation data collection for visuomotor policy learning. While humans can mimic tool manipulation behavior by observing others perform a task just once and effortlessly transfer the skill to diverse tools for functionally equivalent tasks, current robots struggle to achieve this level of generalization. A key challenge lies in establishing function-level correspondences, considering the significant geometric variations among functionally similar tools, referred to as intra-function variations. To address this challenge, we propose MimicFunc, a framework that establishes functional correspondences with function frame, a function-centric local coordinate frame constructed with keypoint-based abstraction, for imitating tool manipulation skills. Experiments demonstrate that MimicFunc effectively enables the robot to generalize the skill from a single RGB-D human video to manipulating novel tools for functionally equivalent tasks. Furthermore, leveraging MimicFunc's one-shot generalization capability, the generated rollouts can be used to train visuomotor policies without requiring labor-intensive teleoperation data collection for novel objects. Our code and video are available at https://sites.google.com/view/mimicfunc.

cs.RO

CLASP: General-Purpose Clothes Manipulation with Semantic Keypoints

Clothes manipulation, such as folding or hanging, is a critical capability for home service robots. Despite recent advances, most existing methods remain limited to specific clothes types and tasks, due to the complex, high-dimensional geometry of clothes. This paper presents CLothes mAnipulation with Semantic keyPoints (CLASP), which aims at general-purpose clothes manipulation over diverse clothes types, T-shirts, shorts, skirts, long dresses, ..., as well as different tasks, folding, flattening, hanging, .... The core idea of CLASP is semantic keypoints-e.g., ''left sleeve'' and ''right shoulder''-a sparse spatial-semantic representation, salient for both perception and action. Semantic keypoints of clothes can be reliably extracted from RGB-D images and provide an effective representation for a wide range of clothes manipulation policies. CLASP uses semantic keypoints as an intermediate representation to connect high-level task planning and low-level action execution. At the high level, it exploits vision language models (VLMs) to predict task plans over the semantic keypoints. At the low level, it executes the plans with the help of a set of pre-built manipulation skills conditioned on the keypoints. Extensive simulation experiments show that CLASP outperforms state-of-the-art baseline methods on multiple tasks across diverse clothes types, demonstrating strong performance and generalization. Further experiments with a Franka dual-arm system on four distinct tasks-folding, flattening, hanging, and placing-confirm CLASP's performance on real-life clothes manipulation.

cs.RO

Photometric redshift estimation for emission line galaxies of DESI Legacy Imaging Surveys by CNN-MLP

Emission Line Galaxies (ELGs) are crucial for cosmological studies, particularly in understanding the large-scale structure of the Universe and the role of dark energy. ELGs form an essential component of the target catalogue for the Dark Energy Spectroscopic Instrument (DESI), a major astronomical survey. However, the accurate selection of ELGs for such surveys is challenging due to the inherent uncertainties in determining their redshifts with photometric data. In order to improve the accuracy of photometric redshift estimation for ELGs, we propose a novel approach CNN-MLP that combines Convolutional Neural Networks (CNNs) with Multilayer Perceptrons (MLPs). This approach integrates both images and photometric data derived from the DESI Legacy Imaging Surveys Data Release 10. By leveraging the complementary strengths of CNNs (for image data processing) and MLPs (for photometric feature integration), the CNN-MLP model achieves a $\sigma_{\mathrm{NMAD}}$ (normalised median absolute deviation) of 0.0140 and an outlier fraction of 2.57%. Compared to other models, CNN-MLP demonstrates a significant improvement in the accuracy of ELG photometric redshift estimation, which directly benefits the target selection process for DESI. In addition, we explore the photometric redshifts of different galaxy types (Starforming, Starburst, AGN, Broadline). Furthermore, this approach will contribute to more reliable photometric redshift estimation in ongoing and future large-scale sky surveys (e.g. LSST, CSST, Euclid), enhancing the overall efficiency of cosmological research and galaxy surveys.

astro-ph.IM

An Empirical Study of GPT-4o Image Generation Capabilities

The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to bridge understanding and generation tasks. Recent advances, especially the GPT-4o, have demonstrated the feasibility of high-fidelity multimodal generation, their architectural design remains mysterious and unpublished. This prompts the question of whether image and text generation have already been successfully integrated into a unified framework for those methods. In this work, we conduct an empirical study of GPT-4o's image generation capabilities, benchmarking it against leading open-source and commercial models. Our evaluation covers four main categories, including text-to-image, image-to-image, image-to-3D, and image-to-X generation, with more than 20 tasks. Our analysis highlights the strengths and limitations of GPT-4o under various settings, and situates it within the broader evolution of generative modeling. Through this investigation, we identify promising directions for future unified generative models, emphasizing the role of architectural design and data scaling. For a high-definition version of the PDF, please refer to the link on GitHub: \href{https://github.com/Ephemeral182/Empirical-Study-of-GPT-4o-Image-Gen}{https://github.com/Ephemeral182/Empirical-Study-of-GPT-4o-Image-Gen}.

cs.CV

Dexterous Manipulation through Imitation Learning: A Survey

Dexterous manipulation, which refers to the ability of a robotic hand or multi-fingered end-effector to skillfully control, reorient, and manipulate objects through precise, coordinated finger movements and adaptive force modulation, enables complex interactions similar to human hand dexterity. With recent advances in robotics and machine learning, there is a growing demand for these systems to operate in complex and unstructured environments. Traditional model-based approaches struggle to generalize across tasks and object variations due to the high dimensionality and complex contact dynamics of dexterous manipulation. Although model-free methods such as reinforcement learning (RL) show promise, they require extensive training, large-scale interaction data, and carefully designed rewards for stability and effectiveness. Imitation learning (IL) offers an alternative by allowing robots to acquire dexterous manipulation skills directly from expert demonstrations, capturing fine-grained coordination and contact dynamics while bypassing the need for explicit modeling and large-scale trial-and-error. This survey provides an overview of dexterous manipulation methods based on imitation learning, details recent advances, and addresses key challenges in the field. Additionally, it explores potential research directions to enhance IL-driven dexterous manipulation. Our goal is to offer researchers and practitioners a comprehensive introduction to this rapidly evolving domain.

cs.RO

First-principles predictions of the diversity in atomic structures and electronic properties of the reconstructed Si(111)-7x7 surface

The 7x7 reconstruction of Si(111) surface is widely understood by the dimer-adatom-stacking-fault model (DAS), but the predicted metallicity of DAS contradicts experimental signs of insulation. It is still challenge to predict DAS-like reconstructions by traditional method to solve such a puzzle. Here, we show that low-energy reconstructions of Si(111)-7x7 surface with (DAS-d8-T12, DAS-d8-T9H3-A, DAS-d8-T9H3-B and DAS-d8-T6H6) and without (AB-d10-T12, AB-d10-T9H3, AA-d10-T12 and AA-d10-T9H3) stacking-fault can be quickly discovered by graph theory as implemented in RG2 code for crystal structure prediction. They exhibit comparable stability to the DAS (DAS-d8-T12) model and similar STM patterns, offering a plausible explanation for the observed Si(111)-7x7 reconstruction. All these reconstructions exhibit metallic behavior in the nonmagnetic (NM) state with isolated narrow bands crossing the Fermi level in varying occupancy. And they are further confirmed as ferromagnetic (FM) metals (DAS-d8-T9H3-B), half-metals (DAS-d8-T12, AB-d10-T9H3, AA-d10-T12 and AA-d10-T9H3), half-semimetals (DAS-d8-T9H3-A and DAS-d8-T6H6) and even insulators (AB-d10-T12), depending their occupancies of the NM band structures. These findings not only demonstrate the rich electromagnetic phases of reconstructed Si(111) surfaces and their potential for spintronic applications, but also provide a plausible physical explanation for the metal-insulator transition observed on the Si(111) surface.

cond-mat.mes-hall