arXiv ScienceSearch

arXiv subjects

Kevin Yang

Publications and source records attributed to Kevin Yang.

At least 19 recordsLinked to original sources

Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games

We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three contributions: First, we introduce methods of creating datasets of policies for a given game. Second, we propose methods to learn policy representations. Third, we introduce downstream tasks to evaluate the effectiveness of such representations. We evaluate each dataset method, embedding method, and downstream task on Kuhn and Leduc Poker. Although our methods are very basic, we demonstrate that useful behavioral representations are present in the learned embeddings. To our knowledge, this work is among the first to systematically compare self-supervised learning techniques for learning policy representations in games. Our code is available at https://github.com/VitamintK/ssl-project for others to extend.

cs.LG

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation

We introduce Lucid-XR, a generative data engine for creating diverse and realistic-looking multi-modal data to train real-world robotic systems. At the core of Lucid-XR is vuer, a web-based physics simulation environment that runs directly on the XR headset, enabling internet-scale access to immersive, latency-free virtual interactions without requiring specialized equipment. The complete system integrates on-device physics simulation with human-to-robot pose retargeting. Data collected is further amplified by a physics-guided video generation pipeline steerable via natural language specifications. We demonstrate zero-shot transfer of robot visual policies to unseen, cluttered, and badly lit evaluation environments, after training entirely on Lucid-XR's synthetic data. We include examples across dexterous manipulation tasks that involve soft materials, loosely bound particles, and rigid body contact. Project website: https://lucidxr.github.io

cs.RO

WildSci: Advancing Scientific Reasoning from In-the-Wild Literature

Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are readily available. In contrast, progress in LLM reasoning models remains limited in scientific domains such as medicine and materials science due to limited dataset coverage and the inherent complexity of open-ended scientific questions. To address these challenges, we introduce WildSci, a new dataset of domain-specific science questions automatically synthesized from peer-reviewed literature, covering 9 scientific disciplines and 26 subdomains. By framing complex scientific reasoning tasks in a multiple-choice format, we enable scalable training with well-defined reward signals. We further apply reinforcement learning to finetune models on these data and analyze the resulting training dynamics, including domain-specific performance changes, response behaviors, and generalization trends. Experiments on a suite of scientific benchmarks demonstrate the effectiveness of our dataset and approach. We release WildSci to enable scalable and sustainable research in scientific reasoning, available at https://huggingface.co/datasets/JustinTX/WildSci.

cs.AI

Thermodynamic stability and kinetic control of capsid morphologies in hepatitis B virus

Polymorphism has been observed in viral capsid assembly, demonstrating the ability of identical protein dimers to adopt multiple geometries under the same solution conditions. A well-studied example is the hepatitis B virus (HBV), which forms two stable capsid morphologies both in vivo and in vitro. These capsids differ in diameter, containing either 90 or 120 protein dimers. Experiments have shown that their relative prevalence depends on the ionic conditions of the solution during assembly. We developed a model that incorporates salt effects by altering the intermolecular binding free energy between capsid proteins, thereby shifting the relative thermodynamic stability of the two morphologies. This model reproduces experimental results on the prevalence ratios of the large and small HBV capsids. We also constructed a kinetic model that captures the time-dependent ratio of the two morphologies under subcritical capsid concentrations, consistent with experimental data.

physics.bio-ph

A Machine Learning-Based Framework to Shorten the Questionnaire for Assessing Autism Intervention

Caregivers of individuals with autism spectrum disorder (ASD) often find the 77-item Autism Treatment Evaluation Checklist (ATEC) burdensome, limiting its use for routine monitoring. This study introduces a generalizable machine learning framework that seeks to shorten assessments while maintaining evaluative accuracy. Using longitudinal ATEC data from 60 autistic children receiving therapy, we applied feature selection and cross-validation techniques to identify the most predictive items across two assessment goals: longitudinal therapy tracking and point-in-time severity estimation. For progress monitoring, the framework identified 16 items (21% of the original questionnaire) that retained strong correlation with total score change and full subdomain coverage. We also generated smaller subsets (1-7 items) for efficient approximations. For point-in-time severity assessment, our model achieved over 80% classification accuracy using just 13 items (17% of the original set). While demonstrated on ATEC, the methodology-based on subset optimization, model interpretability, and statistical rigor-is broadly applicable to other high-dimensional psychometric tools. The resulting framework could potentially enable more accessible, frequent, and scalable assessments and offer a data-driven approach for AI-supported interventions across neurodevelopmental and psychiatric contexts.

stat.AP

Understanding Mode Switching in Human-AI Collaboration: Behavioral Insights and Predictive Modeling

Human-AI collaboration is typically offered in one of two of user control levels: guidance, where the AI provides suggestions and the human makes the final decision, and delegation, where the AI acts autonomously within user-defined constraints. Systems that integrate both modes, common in robotic surgery or driving assistance, often overlook shifts in user preferences within a task in response to factors like evolving trust, decision complexity, and perceived control. In this work, we investigate how users dynamically switch between higher and lower levels of control during a sequential decision-making task. Using a hand-and-brain chess setup, participants either selected a piece and the AI decided how it moved (brain mode), or the AI selected a piece and the participant decided how it moved (hand mode). We collected over 400 mode-switching decisions from eight participants, along with gaze, emotional state, and subtask difficulty data. Statistical analysis revealed significant differences in gaze patterns and subtask complexity prior to a switch and in the quality of the subsequent move. Based on these results, we engineered behavioral and task-specific features to train a lightweight model that predicted control level switches ($F1 = 0.65$). The model performance suggests that real-time behavioral signals can serve as a complementary input alongside system-driven mode-switching mechanisms currently used. We complement our quantitative results with qualitative factors that influence switching including perceived AI ability, decision complexity, and level of control, identified from post-game interview analysis. The combined behavioral and modeling insights can help inform the design of shared autonomy systems that need dynamic, subtask-level control switches aligned with user intent and evolving task demands.

cs.HC

KPZ equation from open ASEP with general boundary asymmetry

We consider generalizations of open ASEP in the interval and half-space, where the speed of the reservoir dynamics can depend on the local particle configuration. We show that their height functions have a continuum limit given by the open KPZ equation. This removes the assumption of Liggett's condition in Corwin-Shen '18 and Parekh '19, thus answering a question of Corwin '22 and Himwich '25, and it removes the assumption of product invariant measures in Goncalves-Perkowski-Simon '20. In the case of the interval, we also show convergence of the stationary measure for the height function increment process to that of the increment process for the open KPZ equation.

math.PR

KPZ equation from a class of nonlinear SPDEs in infinite volume

We study a general class of nonlinear Ginzburg-Landau SPDEs in infinite volume under weak nonlinearity scaling and with non-equilibrium initial data. We derive the KPZ equation as a continuum limit of these equations. This makes rigorous the original derivation of the KPZ equation from physics in the full-space setting, which was a problem posed by Hairer-Quastel '18. Our analysis is based on a stochastic heat kernel for a linearization of said SPDEs.

math.PR

Delocalization of Two-Dimensional Random Band Matrices

We study a random band matrix $H=(H_{xy})_{x,y}$ of dimension $N\times N$ with mean-zero complex Gaussian entries, where $x,y$ belong to the discrete torus $(\mathbb{Z}/\sqrt{N}\mathbb{Z})^{2}$. The variance profile $\mathbb{E}|H_{xy}|^{2}=S_{xy}$ vanishes when the distance between $x,y$ is larger than some band-width parameter $W$ depending on $N$. We show that if the band-width satisfies $W\geq N^{\mathfrak{c}}$ for some $\mathfrak{c}>0$, then in the large-$N$ limit, we have the following results. The first result is a local semicircle law in the bulk down to scales $N^{-1+\varepsilon}$. The second is delocalization of bulk eigenvectors. The third is a quantum unique ergodicity for bulk eigenvectors. The fourth is universality of local bulk eigenvalue statistics. The fifth is a quantum diffusion profile for the associated $T$ matrix. Our method is based on embedding $H$ inside a matrix Brownian motion $H_{t}$ as done in [Dubova-Yang '24] and [Yau-Yin '25] for band matrices on the one-dimensional torus. In this paper, the key additional ingredient in our analysis of $H_{t}$ is a new CLT-type estimate for polynomials in the entries of the resolvent of $H_{t}$.

math.PR

Colorful Helly via induced matchings

We establish a theorem regarding the maximum size of an {\it{induced}} matching in the bipartite complement of the incidence graph of a set system $(X,\mathcal{F})$. We show that this quantity plus one provides an upper bound on the colorful Helly number of this set system, i.e. the minimum positive integer $N$ for which the following statement holds: if finite subfamilies $\mathcal{F}_1,\ldots, \mathcal{F}_{N} \subset \mathcal{F}$ are such that $\cap_{F \in \mathcal{F}_{i}} F = 0$ for every $i=1,\ldots,N$, then there exists $F_i \in \mathcal{F}_i$ such that $F_1 \cap \ldots \cap F_{N} = \emptyset$. We will also discuss some natural refinements of this result and applications.

math.CO

Quantum diffusion and delocalization in one-dimensional band matrices via the flow method

We study a class of Gaussian random band matrices of dimension $N \times N$ and band-width $W$. We show that delocalization holds for bulk eigenvectors and that quantum diffusion holds for the resolvent, all under the assumption that $W \gg N^{8/11}$. Our analysis is based on a flow method, and a refinement of it may lead to an improvement on the condition $W \gg N^{8/11}$.

math.PR

KPZ equation from ASEP plus general speed-change drift

We derive the KPZ equation as a continuum limit of height functions in asymmetric simple exclusion processes with drift that depends on the local particle configuration. To our knowledge, it is a first such result for a class of particle systems without duality or explicit invariant measures. The tools developed in this paper consist of estimates on the corresponding Kolmogorov equations, giving a more robust proof of the Boltzmann-Gibbs principle. These tools are not exclusive to KPZ.

math.PR

FACTTRACK: Time-Aware World State Tracking in Story Outlines

While accurately detecting and correcting factual contradictions in language model outputs has become increasingly important as their capabilities improve, doing so is highly challenging. We propose a novel method, FACTTRACK, for tracking atomic facts and addressing factual contradictions. Crucially, FACTTRACK also maintains time-aware validity intervals for each fact, allowing for change over time. At a high level, FACTTRACK consists of a four-step pipeline to update a world state data structure for each new event: (1) decompose the event into directional atomic facts; (2) determine the validity interval of each atomic fact using the world state; (3) detect contradictions with existing facts in the world state; and finally (4) add new facts to the world state and update existing atomic facts. When we apply FACTTRACK to contradiction detection on structured story outlines, we find that FACTTRACK using LLaMA2-7B-Chat substantially outperforms a fair baseline using LLaMA2-7B-Chat, and achieves performance comparable to a GPT4 baseline. Moreover, when using GPT4, FACTTRACK significantly outperforms the GPT4 baseline.

cs.CL

THOUGHTSCULPT: Reasoning with Intermediate Revision and Search

We present THOUGHTSCULPT, a general reasoning and search method for tasks with outputs that can be decomposed into components. THOUGHTSCULPT explores a search tree of potential solutions using Monte Carlo Tree Search (MCTS), building solutions one action at a time and evaluating according to any domain-specific heuristic, which in practice is often simply an LLM evaluator. Critically, our action space includes revision actions: THOUGHTSCULPT may choose to revise part of its previous output rather than continuing to build the rest of its output. Empirically, THOUGHTSCULPT outperforms state-of-the-art reasoning methods across three challenging tasks: Story Outline Improvement (up to +30% interestingness), Mini-Crosswords Solving (up to +16% word success rate), and Constrained Generation (up to +10% concept coverage).

cs.CL

KPZ-type equation from growth driven by a non-Markovian diffusion

We study a stochastic PDE model for an evolving set $\mathbb{M}(t)\subseteq\mathbb{R}^{\mathrm{d}+1}$ that resembles a continuum version of origin-excited or reinforced random walk. We show that long-time fluctuations of an associated height function are given by a regularized Kardar-Parisi-Zhang (KPZ)-type PDE on a hypersurface in $\mathbb{R}^{\mathrm{d}+1}$, modulated by a Dirichlet-to-Neumann operator. We also show that for $\mathrm{d}+1=2$, the regularization in this KPZ-type equation can be removed after renormalization. To our knowledge, this gives the first instance of KPZ-type behavior in Laplacian growth, which was asked about (for somewhat different models) in Parisi-Zhang '84 and Ramirez-Sidoravicius '04.

math.PR

Improving Pacing in Long-Form Story Planning

Existing LLM-based systems for writing long-form stories or story outlines frequently suffer from unnatural pacing, whether glossing over important events or over-elaborating on insignificant details, resulting in a jarring experience for the reader. We propose a CONCrete Outline ConTrol (CONCOCT) system to improve pacing when automatically generating story outlines. We first train a concreteness evaluator to judge which of two events is more concrete (low-level-detailed). This evaluator can then be used to control pacing in hierarchical outline generation; in this work, we explore a vaguest-first expansion procedure that aims for uniform pacing. We further use the evaluator to filter new outline items based on predicted concreteness. Compared to a baseline hierarchical outline generator, humans judge CONCOCT's pacing to be more consistent over 57% of the time across multiple outline lengths; the gains also translate to downstream stories. All code, data, and models are open-sourced.

cs.CL