arXiv ScienceSearch

arXiv subjects

Huang

Publications and source records attributed to Huang.

At least 19 recordsLinked to original sources

OptiClear: Differentiable Curvilinear Design Rule Legalization for Inverse-Designed Photonic Devices

Photonic inverse design enables ultra-compact, high-performance devices with highly curvilinear and non-intuitive geometries, but the resulting layouts often violate fabrication design rules and hinder foundry manufacturing. Legalization methods designed for rectilinear Manhattan electrical layouts are not directly applicable to curvilinear inverse-designed photonic devices. Meanwhile, existing fabrication-aware inverse-design methods apply soft penalties on small features and sharp curvatures, but still cannot guarantee design-rule-compliant final layouts. In this work, we present OptiClear, a curvilinear design rule legalization framework for inverse-designed photonic devices. OptiClear provides two complementary legalization engines: OptiClear-R, a rule-based morphological legalizer that efficiently resolves violation regions through iterative morphology-guided mask processing, and OptiClear-D, a differentiable legalizer that formulates legalization as a minimum-distortion mask optimization problem under morphological stationary-point constraints, explicitly seeking a rule-compliant layout with minimal geometric deviation from the original design. We further develop customized differentiable morphological GPU operators that significantly improve the scalability of high-resolution mask legalization. Comprehensive evaluation across diverse inverse-designed photonic devices and a wide range of design-rule settings shows that OptiClear reduces design-rule violations from thousands to zero. The rule-based legalizer offers high runtime efficiency, while the differentiable legalizer more faithfully preserves the original optical functionality. This work establishes curvilinear design rule legalization as a practical post-design electronic-photonic design automation (EPDA) stage for translating high-performance inverse-designed photonic layouts into manufacturable tape-out-ready devices.

physics.optics

Regression Accumulation in Multi-Turn LLM Programming Conversations

In LLM-assisted software development, coding is often iterative. We study regression accumulation in multi-turn LLM programming conversations, where later code suggestions may break requirements introduced in earlier turns. Reliability therefore depends not only on satisfying the current request, but also on preserving previously satisfied behavior. We construct 542 tasks from HumanEval+ and MBPP+ and extend each task into an 8-turn requirement-evolution chain. We evaluate six LLMs on 26,016 turn instances (542 x 6 x 8). At each turn, we test whether the current code still passes earlier benchmark tests. We also analyze 384 failure cases from the failure population and build a taxonomy of multi-turn regression bugs through independent four-annotator labeling. Our results show that regression accumulation appears across all six models: 40% to 73% of tasks lose previously correct behavior over the full conversation. Final-turn quality is lower than initial-turn quality across models, especially when later turns add input validation or broader input types. Manual analysis shows that Cross-Turn Conflict, where later code conflicts with earlier requirements, is the main failure class. We further find that Verification Gate, which checks new code against prior tests and triggers rollback and retry, is the only strategy that consistently improves all models, raising final-turn quality from 75.8% to 87.9% on DeepSeek-V3 and from 31.6% to 47.3% on Llama-3.1-8B. These findings suggest that strong single-turn performance can overestimate reliability in multi-turn coding conversations. Future evaluation and tool design should test whether later code suggestions preserve earlier requirements and should include Verification Gate mechanisms.

cs.SE

Efficient Financial Language Understanding via Distillation with Synthetic Data

Large instruction-following models are powerful but costly to deploy, particularly in finance, where labelled data are limited by confidentiality and expert annotation cost. We present an efficient framework for financial sentiment analysis through distillation with synthetic data, transferring knowledge from a large instruction-tuned teacher to compact student models. The framework is designed for low-resource conditions, where a small set of real examples are collected and labelled by hand. The framework then clusters the examples and uses the clusters to select seeds for generating synthetic examples via structured few-shot prompting. Experiments show that clustering-based seed selection yields more representative synthetic data than random sampling, enabling compact models to achieve strong performance with minimal supervision. Notably, on a more complex and noisy text domain, the compact model trained on the complete synthetic-seed corpus even outperforms the teacher model, while remaining competitive on formal text. The framework provides a practical route toward resource-efficient domain adaptation in financial NLP with minimal human labelling effort.

cs.CL

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling. There the default validity test, correlating a metric to human judgment, has no stable anchor: inter-rater agreement is low, structured by annotator identity, barely reproducible, and length-biased. So we cannot answer the question that matters: does capability that scales on objective benchmarks transfer to subjective behavior, and would our instruments even tell us if it did not? We build an instrument for this regime and report what it reveals at the frontier. We contribute, first, a self-evolving instrument that selects and then authors its own behavioral dimensions under a multiplicative anti-gaming fitness, self-halting when it stops improving; second, a trust-by-construction paradigm that earns belief through three certificates established without a human gold standard, where human raters saturate (rho ~ 0.45); and third, the finding it makes visible -- capability transfer is dissociable. Across 49 models, 8 families, and 24 months, subjective behaviors are where objective-benchmark scaling fails to carry over: the sharpest case, advice-restraint (knowing when not to give advice), is the frontier's universal-lowest dimension, and at gpt-4.1->gpt-5 it ran backwards while the aggregate score hid it -- a regression one instruction recovers. Warm restraint is moved by model generation, not by raw scale, MoE width, inference budget, or reasoning mode; the open-weight Pareto frontier matches closed flagships at ~10-80x lower per-call cost; and four judge families replicate the rubric on held-out human ESConv conversations. Data, code, the locked rubric, and judge prompts will be released upon publication.

cs.CL

Replicating Real-World 23-Hz Oscillations Caused by Large Electronic Loads

In 2024, Texas operators observed 23-Hz oscillations in real power measurements close to a large electronic load (LEL). Oscillations emerged when the load's power consumption reached approximately 320 MW level and subsided as the active power demand decreased. The paper aims to analyze the event and reproduce the oscillations using electromagnetic transient (EMT) simulations. In the first stage, a representative feedback system is developed, and frequency-domain analysis is conducted to examine the phenomenon and identify its key influencing factors. Next, detailed EMT simulations are performed to further validate the proposed analytical approach. The results show that the feedback system effectively captures and characterizes the critical features of the 23-Hz oscillation incident. In addition, the EMT simulations successfully reproduce the real-world event, with the simulated results closely matching the fault recorder data.

eess.SY

Urban to Rural Migration in Eastern Europe: Unpacking digital ruralities through TikTok video analysis

Urban to rural migration is a less-researched phenomenon compared to its counterpart: rural to urban migration. In parts of Europe, an increasing number of people living in big urban centers within the country, or moving from other countries decide to relocate to rural areas. In this paper, we examine this phenomenon by analysing content posted on TikTok that documents this transition. We collected a corpus of 901 videos posted until late 2025, documenting urban to rural migration in Romania, under three hashtags, which have collectively been played a total of 24 million times at the time when we gathered the dataset. We analyse this corpus both quantitatively and qualitatively and discuss our findings through the lens of digital rurality - a theory based on Harvey's and Soja's spatial triad, applied to rural spaces, and based on the role of digital technologies as (re-)mediators of everyday lived experience. Specifically, we analyze the corpus as: (a) digital rural localities, (b) formal representations of the digital rural, and (c) everyday lives of the digital rural. We find that (a) Social media platforms enable new forms of paid labor that sometimes involve the commodification of the self in rural areas, although many of the creators we analyze do not explicitly acknowledge this with their audiences. (b) The digital rural gains new forms of representation, and rural areas in remote Romania are highly data-rich across TikTok. (c) The everyday lives represented through the digital rural are sometimes idealized or romanticised. However, they serve as promoters for tourism and are used as sites to document and discuss a variety of topics including giving ample health advice, typically by non-specialists and sometimes criticizing Western medicine, expressing and promoting religious and political views but also acting as forms of general self-expression.

cs.HC

Margin-Consistent Deep Subtyping of Invasive Lung Adenocarcinoma via Perturbation Fidelity in Whole-Slide Image Analysis

Whole-slide image classification for invasive lung adenocarcinoma subtyping remains vulnerable to real-world imaging perturbations that undermine model reliability at the decision boundary. We propose a margin consistency framework evaluated on 203,226 patches from 143 whole-slide images spanning five adenocarcinoma subtypes in the BMIRDS-LUAD dataset. By combining attention-weighted patch aggregation with margin-aware training, our approach achieves robust feature-logit space alignment measured by Kendall correlations of 0.88 during training and 0.64 during validation. Contrastive regularization, while effective at improving class separation, tends to over-cluster features and suppress fine-grained morphological variation; to counteract this, we introduce Perturbation Fidelity (PF) scoring, which imposes structured perturbations through Bayesian-optimized parameters. Vision Transformer-Large achieves 95.20 +/- 4.65% accuracy, representing a 40% error reduction from the 92.00 +/- 5.36% baseline, while ResNet101 with an attention mechanism reaches 95.89 +/- 5.37% from 91.73 +/- 9.23%, a 50% error reduction. All five subtypes exceed an area under the receiver operating characteristic curve (AUC) of 0.99. On the WSSS4LUAD external benchmark, ResNet50 with an attention mechanism attains 80.1% accuracy, demonstrating cross-institutional generalizability despite approximately 15-20% domain-shift-related degradation and identifying opportunities for future adaptation research.

cs.CV

AI Deception: Risks, Dynamics, and Controls

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides a comprehensive and up-to-date overview of the AI deception field, covering its core concepts, methodologies, genesis, and potential mitigations. First, we identify a formal definition of AI deception, grounded in signaling theory from studies of animal deception. We then review existing empirical studies and associated risks, highlighting deception as a sociotechnical safety challenge. We organize the landscape of AI deception research as a deception cycle, consisting of two key components: deception emergence and deception treatment. Deception emergence reveals the mechanisms underlying AI deception: systems with sufficient capability and incentive potential inevitably engage in deceptive behaviors when triggered by external conditions. Deception treatment, in turn, focuses on detecting and addressing such behaviors. On deception emergence, we analyze incentive foundations across three hierarchical levels and identify three essential capability preconditions required for deception. We further examine contextual triggers, including supervision gaps, distributional shifts, and environmental pressures. On deception treatment, we conclude detection methods covering benchmarks and evaluation protocols in static and interactive settings. Building on the three core factors of deception emergence, we outline potential mitigation strategies and propose auditing approaches that integrate technical, community, and governance efforts to address sociotechnical challenges and future AI risks. To support ongoing work in this area, we release a living resource at www.deceptionsurvey.com.

cs.AI

Nanoconfinement Effects on Intermolecular Forces Observed via Dewetting

Although wettability is a macroscopic manifestation of molecular-level forces, such as van der Waals (vdW) forces, the impact of nanoconfinement on material properties in reduced film thickness remains unexplored in predicting film stability. In this work, we investigate how nanoconfinement influences intermolecular interactions using a model trilayer system composed of a thick polystyrene (PS) base, a poly(methyl methacrylate) (PMMA) middle layer with tunable thickness (15-95 nm), and a 10 nm top PS film. We find that the dewetting behavior of the top PS layer is highly sensitive to middle PMMA thickness, deviating from classical vdW-based predictions that assume bulk material properties. By incorporating nanoconfinement-induced changes in PMMA refractive index into the calculation of the Hamaker constant, we present a modified theoretical framework that successfully captures the observed behavior. This study links dewetting behavior and material property change as a function of underlayer thickness, providing direct evidence that nanoconfinement in soft matter systems significantly influences long-range intermolecular interactions. We show that film stability can be tuned solely by adjusting underlying layer thickness, while preserving both chemistry and thickness of top functional film. This finding carries broad implications for thin-film technologies across scientific and engineering disciplines by enabling performance-targeted interface design.

cond-mat.soft

Semiconductor SEM Image Defect Classification Using Supervised and Semi-Supervised Learning with Vision Transformers

Controlling defects in semiconductor processes is important for maintaining yield, improving production cost, and preventing time-dependent critical component failures. Electron beam-based imaging has been used as a tool to survey wafers in the line and inspect for defects. However, manual classification of images for these nano-scale defects is limited by time, labor constraints, and human biases. In recent years, deep learning computer vision algorithms have shown to be effective solutions for image-based inspection applications in industry. This work proposes application of vision transformer (ViT) neural networks for automatic defect classification (ADC) of scanning electron microscope (SEM) images of wafer defects. We evaluated our proposed methods on 300mm wafer semiconductor defect data from our fab in IBM Albany. We studied 11 defect types from over 7400 total images and investigated the potential of transfer learning of DinoV2 and semi-supervised learning for improved classification accuracy and efficient computation. We were able to achieve classification accuracies of over 90% with less than 15 images per defect class. Our work demonstrates the potential to apply the proposed framework for a platform agnostic in-house classification tool with faster turnaround time and flexibility.

cs.CV

LiDAR 2.0: Hierarchical Curvy Waveguide Detailed Routing for Large-Scale Photonic Integrated Circuits

Driven by innovations in photonic computing and interconnects, photonic integrated circuit (PIC) designs advance and grow in complexity. Traditional manual physical design processes have become increasingly cumbersome. Available PIC layout tools are mostly schematic-driven, which has not alleviated the burden of manual waveguide planning and layout drawing. Previous research in PIC automated routing is largely adapted from electronic design, focusing on high-level planning and overlooking photonic-specific constraints such as curvy waveguides, bending, and port alignment. As a result, they fail to scale and cannot generate DRV-free layouts, highlighting the need for dedicated electronic-photonic design automation tools to streamline PIC physical design. In this work, we present LiDAR, the first automated PIC detailed router for large-scale designs. It features a grid-based, curvy-aware A* engine with adaptive crossing insertion, congestion-aware net ordering, and insertion-loss optimization. To enable routing in more compact and complex designs, we further extend our router to hierarchical routing as LiDAR 2.0. It introduces redundant-bend elimination, crossing space preservation, and routing order refinement for improved conflict resilience. We also develop and open-source a YAML-based PIC intermediate representation and diverse benchmarks, including TeMPO, GWOR, and Bennes, which feature hierarchical structures and high crossing densities. Evaluations across various benchmarks show that LiDAR 2.0 consistently produces DRV-free layouts, achieving up to 16% lower insertion loss and 7.69x speedup over prior methods on spacious cases, and 9% lower insertion loss with 6.95x speedup over LiDAR 1.0 on compact cases. Our codes are open-sourced at https://github.com/ScopeX-ASU/LiDAR.

cs.ET

Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection

Weakly Supervised Sound Event Detection (WSSED), which relies on audio tags without precise onset and offset times, has become prevalent due to the scarcity of strongly labeled data that includes exact temporal boundaries for events. This study introduces Frame-level Pseudo Strong Labeling (FPSL) to overcome the lack of temporal information in WSSED by generating pseudo strong labels from frame-level predictions. This enhances temporal localization during training and addresses the limitations of clip-wise weak supervision. We validate our approach across three benchmark datasets (DCASE2017 Task 4, DCASE2018 Task 4, and UrbanSED) and demonstrate significant improvements in key metrics such as the Polyphonic Sound Detection Scores (PSDS), event-based F1 scores, and intersection-based F1 scores. For example, Convolutional Recurrent Neural Networks (CRNNs) trained with FPSL outperform baseline models by 4.9% in PSDS1 on DCASE2017, 7.6% on DCASE2018, and 1.8% on UrbanSED, confirming the effectiveness of our method in enhancing model performance.

eess.AS

Ultra low-cost fabrication of homogeneous alginate hydrogel microspheres in symmetry designed microfluidic device

In this study, we present a two-stage method for fabricating monodisperse alginate hydrogel microspheres using a symmetrically designed flow-focusing microfluidic device. One of the flow-focusing junctions generates alginate hydrogel droplets without the addition of surfactants, while the other junction introduces corn oil with acetic acid, which facilitates the solidification of the homogeneous alginate hydrogel droplets and prevents coalescence. These hydrogel microspheres can be easily separated from the oil phase using an oscillation state, eliminating the need for a demulsifier. This microfluidic system for hydrogel microsphere formation is notable for its simplicity, ease of fabrication, and user-friendliness.

physics.chem-ph

On the Price of Decentralization in Decentralized Detection

Fundamental limits on the error probabilities of a family of decentralized detection algorithms (eg., the social learning rule proposed by Lalitha et al. over directed graphs are investigated. In decentralized detection, a network of nodes locally exchanging information about the samples they observe with their neighbors to collectively infer the underlying unknown hypothesis. Each node in the network weighs the messages received from its neighbors to form its private belief and only requires knowledge of the data generating distribution of its observation. In this work, it is first shown that while the original social learning rule of Lalitha et al. achieves asymptotically vanishing error probabilities as the number of samples tends to infinity, it suffers a gap in the achievable error exponent compared to the centralized case. The gap is due to the network imbalance caused by the local weights that each node chooses to weigh the messages received from its neighbors. To close this gap, a modified learning rule is proposed and shown to achieve error exponents as large as those in the centralized setup. This implies that there is essentially no first-order penalty caused by decentralization in the exponentially decaying rate of error probabilities.

cs.IT

Impact of Noisy Labels on Sound Event Detection: Deletion Errors Are More Detrimental Than Insertion Errors

This study explores the critical but underexamined impact of label noise on Sound Event Detection (SED), which requires both sound identification and precise temporal localization. We categorize label noise into deletion, insertion, substitution, and subjective types and systematically evaluate their effects on SED using synthetic and real-life datasets. Our analysis shows that deletion noise significantly degrades performance, while insertion noise is relatively benign. Moreover, loss functions effective against classification noise do not perform well for SED due to intra-class imbalance between foreground sound events and background sounds. We demonstrate that loss functions designed to address data imbalance in SED can effectively reduce the impact of noisy labels on system performance. For instance, halving the weight of background sounds in a synthetic dataset improved macro-F1 and micro-F1 scores by approximately $9\%$ with minimal Error Rate increase, with consistent results in real-life datasets. This research highlights the nuanced effects of noisy labels on SED systems and provides practical strategies to enhance model robustness, which are pivotal for both constructing new SED datasets and improving model performance, including efficient utilization of soft and crowdsourced labels.

eess.AS

Envisioning Possibilities and Challenges of AI for Personalized Cancer Care

The use of Artificial Intelligence (AI) in healthcare, including in caring for cancer survivors, has gained significant interest. However, gaps remain in our understanding of how such AI systems can provide care, especially for ethnic and racial minority groups who continue to face care disparities. Through interviews with six cancer survivors, we identify critical gaps in current healthcare systems such as a lack of personalized care and insufficient cultural and linguistic accommodation. AI, when applied to care, was seen as a way to address these issues by enabling real-time, culturally aligned, and linguistically appropriate interactions. We also uncovered concerns about the implications of AI-driven personalization, such as data privacy, loss of human touch in caregiving, and the risk of echo chambers that limit exposure to diverse information. We conclude by discussing the trade-offs between AI-enhanced personalization and the need for structural changes in healthcare that go beyond technological solutions, leading us to argue that we should begin by asking, ``Why personalization?''

cs.HC

The Clever Hans Mirage: A Comprehensive Survey on Spurious Correlations in Machine Learning

Back in the early 20th century, a horse named Hans appeared to perform arithmetic and other intellectual tasks during exhibitions in Germany, while it actually relied solely on involuntary cues in the body language from the human trainer. Modern machine learning models are no different. These models are known to be sensitive to spurious correlations between non-essential features of the inputs (e.g., background, texture, and secondary objects) and the corresponding labels. Such features and their correlations with the labels are known as "spurious" because they tend to change with shifts in real-world data distributions, which can negatively impact the model's generalization and robustness. In this paper, we provide a comprehensive survey of this emerging issue, along with a fine-grained taxonomy of existing state-of-the-art methods for addressing spurious correlations in machine learning models. Additionally, we summarize existing datasets, benchmarks, and metrics to facilitate future research. The paper concludes with a discussion of the broader impacts, the recent advancements, and future challenges in the era of generative AI, aiming to provide valuable insights for researchers in the related domains of the machine learning community.

cs.LG

Bernstein-Gelfand-Gelfand meets geometric complexity theory: resolving the 2 x 2 permanents of a 2 x n matrix

We describe the minimal free resolution of the ideal of $2 \times 2$ subpermanents of a $2 \times n$ generic matrix $M$. In contrast to the case of $2 \times 2$ determinants, the $2 \times 2$ permanents define an ideal which is neither prime nor Cohen-Macaulay. We combine work of Laubenbacher-Swanson on the Gr\"obner basis of an ideal of $2 \times 2$ permanents of a generic matrix with our previous work connecting the initial ideal of $2 \times 2$ permanents to a simplicial complex. The main technical tool is a spectral sequence arising from the Bernstein-Gelfand-Gelfand correspondence.

math.AC