arXiv ScienceSearch

arXiv subjects

Tianyi Yu

Publications and source records attributed to Tianyi Yu.

At least 19 recordsLinked to original sources

Testing Interchangeability in LLM Agent Teams

Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team's history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.

cs.AI

Closed-loop AI achieves certifiable engineering design

Agentic AI has automated parts of scientific discovery, including paper generation, expert-level coding, therapeutic proposal, and autonomous experimentation. Complex physical engineering design remains a gap, because candidates must satisfy simultaneous constraints in fluid dynamics, solid mechanics, and structural stability. We introduce The AI Engineer, an agentic framework that couples large language models (LLMs) to deterministic engineering backends in a closed loop: natural-language requirements are converted into design-domain geometry and mesh; topology is optimized with bi-directional evolutionary structural optimization (BESO) coupled to the CalculiX solver; and member sizes are refined with particle swarm optimization (PSO) coupled to Zwind under offshore aero-hydro-servo-elastic load cases. To explore many designs without per-candidate certification cost, an Automated Reviewer scores each candidate on five dimensions (capacity, steel intensity, unit cost, constructability, and fatigue life) using piecewise-linear functions calibrated on 11 real floating-wind projects. Search terminates only when a candidate reaches a composite score $S \ge 85$ (grade A) with no subscore below 60. We validated this gate by submitting the top-scoring design to the China Classification Society (CCS) for Approval in Principle (AIP), which it passed; AIP is thus an external check that the reviewer tracks professional judgment, not the daily objective. The certified design outperforms the human-optimized TuQiang baseline, reducing steel mass and unit capital cost by 8.1% each while meeting all AIP criteria. This verification-closed regime, in which every proposal is judged by deterministic physics and codified limit states, distinguishes The AI Engineer from open-ended generative systems. Remaining limits include detailed design and fabrication-hard constraints.

cs.AI

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks evaluate them under fixed game mechanics. High performance in these settings does not show whether an agent can continue making progress when familiar recipes, drops, and other rules change. In this paper, we introduce MirrorCraft, a paired benchmark for evaluating agents under hidden rule changes in Minecraft. Each Mirror world is a copy of its paired Vanilla world, with selected server-side rules modified by the corresponding datapack. Terrain, spawn, resource placement, objective, interface, and action budget remain matched within every Vanilla-Mirror pair. MirrorCraft includes five controlled biomes, six rule suites, three progression objectives, two model families, and six agent configurations under a shared Mineflayer interface. We evaluate task progress with deterministic advancement milestones and success rate and use the Rule Intervention Effect (RIE) to measure the performance change between matched Vanilla and Mirror worlds. The experiments show that hidden rule changes have strongly different effects across suites. Among the configurations evaluated without rule descriptions, ReAct achieves the highest pooled Mirror score. Providing the exact rules yields modest gains in average progress and completion across all three objectives. MirrorCraft extends Minecraft evaluation beyond fixed mechanics and provides a controlled setting for studying how agents use gameplay outcomes when the rules of the current world differ from familiar ones.

cs.AI

A positive combinatorial formula for the double Edelman--Greene coefficients

Lam, Lee, and Shimozono introduced the double Stanley symmetric functions in their study of the equivariant geometry of the affine Grassmannian. They proved that the associated double Edelman--Greene coefficients, the double Schur expansion coefficients of these functions, are positive, a result later refined by Anderson. They further asked for a combinatorial proof of this positivity. In this paper, we provide the first such proof, together with a combinatorial formula that manifests the finer positivity established by Anderson. Our formula is built from two combinatorial models: bumpless pipedreams and increasing chains in the Bruhat order. The proof relies on three key ingredients: a correspondence between these two models, a natural subdivision of bumpless pipedreams, and a symmetry property of increasing chains.

math.CO

Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing

Recent alignment methods based on Direct Preference Optimization (DPO) reformulate preference learning as supervised optimization over pairwise comparisons, offering improved efficiency and stability over reinforcement learning from human feedback (RLHF). However, existing DPO-style methods implicitly assume a single fixed preference objective, which limits their ability to model the structured and sometimes conflicting nature of real-world human judgments that span multiple preference dimensions. In this work, we propose Listwise Direct Preference Optimization ($\lambda$-DPO), a unified framework that simultaneously improves supervision granularity and preference flexibility. Instead of collapsing multi-dimensional preference signals into a single ranking, $\lambda$-DPO constructs a mixture of listwise preference distributions weighted by a preference vector $\lambda$ on the probability simplex, enabling a single model to internalize a continuous spectrum of preference trade-offs. To further improve robustness, we introduce a performance-driven stochastic $\lambda$ scheduler that adaptively samples preference weights based on empirical downstream performance, explicitly mitigating the risks of misspecification inherent to static weighting schemes. We evaluate our method across multiple model families and scales on six widely used benchmarks. Experimental results show the consistent improvement against baselines.

cs.LG

Grothendieck positivity for normal square root crystals

Normal crystals (also known as Stembridge crystals) are commonly used to establish the Schur positivity of symmetric functions, as their characters are sums of Schur polynomials. In this paper, we develop a combinatorial framework for a novel family of objects called normal square root crystals, which are closely related to symmetric Grothendieck functions, the $K$-theoretic analogue of Schur functions. Among other applications, this tool leads to a new proof of Buch's combinatorial rule for the multiplication of symmetric Grothendieck functions. The definition of a normal square root crystal, originally formulated by the first two authors, largely mirrors that of normal crystals. Our main result is to show that the character of such a crystal is always a sum of symmetric Grothendieck polynomials. The proof relies on an unexpected connection between the raising operators for our crystals and the Hecke insertion algorithm developed by Buch, Kresch, Shimozono, Tamvakis, and Yong.

math.CO

Tableau formula for vexillary double Edelman--Greene coefficients

Lam, Lee and Shimozono recently introduced backstable double Grothendieck polynomials to represent $K$-theory classes of the infinite flag variety. They used them to define double $\beta$-Stanley symmetric functions, which expand into double stable Grothendieck functions with polynomial coefficients called double $\beta$-Edelman--Greene coefficients. Anderson proved these coefficients are $\beta$-Graham positive. For vexillary permutations, this is equivalent to a statement for skew flagged double $\beta$-Grothendieck functions. Working in this setting, we give a tableau formula for vexillary double $\beta$-Edelman--Greene coefficients that is manifestly $\beta$-Graham positive. Our formula demonstrates a finer notion of positivity than was previously known.

math.CO

Marked Bumpless Pipedreams and Compatible Pairs

We construct a bijection between marked bumpless pipedreams with reverse compatible pairs, which are in bijection with not-necessarily-reduced pipedreams. This directly unifies various formulas for Grothendieck polynomials in the literature. Our bijection is a generalization of a variant of the bijection of Gao and Huang in the unmarked, reduced case.

math.CO

Embedding bumpless pipedreams as Bruhat chains

Schubert polynomials are distinguished representatives of Schubert cycles in the cohomology of the flag variety. In the spirit of Bergeron and Sottile, we use the Bruhat order to give $(n-1)!$ different combinatorial formulas for the Schubert polynomial of a permutation in $S_n$. By work of Lenart and Sottile, one extreme of the formulas recover the classical Pipedream (PD) formula. We prove the other extreme corresponds to Bumpless pipedreams (BPDs). We give two applications of this perspective to view BPDs: Using the Fomin-Kirrilov algebra, we solve the problem of finding a BPD analogue of Fomin and Stanley's algebraic construction on PDs; We also establish a bijection between PDs and BPDs using Lenart's growth diagram, which conjectually agrees with the existing bijection of Gao and Huang.

math.CO

Grothendieck polynomials of inverse fireworks permutations

Pipedreams are combinatorial objects that compute Grothendieck polynomials. We introduce a new combinatorial object that naturally recast the pipedream formula. From this, we obtain the first direct combinatorial formula for the top degree components of Grothendieck polynomials, also known as the Castelnuovo-Mumford polynomials. We also prove the inverse fireworks case of a conjecture of M\'esz\'aros, Setiabrata, and St. Dizier on the support of Grothendieck polynomials.

math.CO

Lascoux expansion of the product of a Lascoux and a stable Grothendieck

This paper gives a tableau formula for expanding the product of a Lascoux polynomial and a stable Grothendieck polynomial into Lascoux polynomials. Lascoux and stable Grothendieck polynomials are inhomogeneous analogues of key polynomials and Stanley symmetric functions, respectively. Our formula refines the K-theoretic Littlewood-Richardson rule of Buch and extends the key expansion of key times Schur established by Haglund, Luoto, Mason, and van Willigenburg. Our proof is combinatorial, relying heavily on a novel row insertion algorithm of Huang, Shimozono and Yu.

math.CO

Constructing maximal pipedreams of double Grothendieck polynomials

Pechenik, Speyer and Weigandt defined a statistic $\mathsf{rajcode}(\cdot)$ on permutations which characterizes the leading monomial in top degree components of double Grothendieck polynomials. Their proof is combinatorial: They showed there exists a unique pipedream of a permutation $w$ with row weight $\mathsf{rajcode}(w)$ and column weight $\mathsf{rajcode}(w^{-1})$. They proposed the problem of finding a ``direct recipe'' for this pipedream. We solve this problem by providing an algorithm that constructs this pipedream via ladder moves.

math.CO

Constructing a Gr\"obner basis of Griffin's ideal

In his Ph.D. thesis, Sean Griffin introduced a family of ideals and found monomial bases for their quotient rings. These rings simultaneously generalize the Delta Conjecture coinvariant rings of Haglund-Rhoades-Shimozono and the cohomology rings of Springer fibers studied by Tanisaki and Garsia-Procesi. We recursively construct a Gr\"{o}bner basis of Griffin's ideals with respect to the graded reverse lexicographical order. Consequently, Griffin's monomial basis is the standard monomial basis. Coefficients of polynomials in our Gr\"{o}bner basis are integers and leading coefficients are one.

math.CO

Connection between Schubert polynomials and top Lascoux polynomials

Schubert polynomials form a basis of the polynomial ring. This basis and its structure constants have received extensive study. Recently, Pan and Yu initiated the study of top Lascoux polynomials. These polynomials form a basis of a subalgebra of the polynomial ring where each graded piece has finite dimension. This paper connects Schubert polynomials and top Lascoux polynomials via a simple operator. We use this connection to show these two bases share the same structure constants. We also translate several results on Schubert polynomials to top Lascoux polynomials, including combinatorial formulas for their monomial expansions and supports.

math.CO

Top-degree components of Grothendieck and Lascoux polynomials

The Castelnuovo-Mumford polynomial $\widehat{\mathfrak{G}}_w$ with $w \in S_n$ is the highest homogeneous component of the Grothendieck polynomial $\mathfrak{G}_w$. Pechenik, Speyer and Weigandt define a statistic $\mathsf{rajcode}(\cdot)$ on $S_n$ that gives the leading monomial of $\widehat{\mathfrak{G}}_w$. We introduce a statistic $\mathsf{rajcode}(\cdot)$ on any diagram $D$ through a combinatorial construction ``snow diagram'' that augments and decorates $D$. When $D$ is the Rothe diagram of a permutation $w$, $\mathsf{rajcode}(D)$ agrees with the aforementioned $\mathsf{rajcode}(w)$. When $D$ is the key diagram of a weak composition $\alpha$, $\mathsf{rajcode}(D)$ yields the leading monomial of $\widehat{\mathfrak{L}}_\alpha$, the highest homogeneous component of the Lascoux polynomials $\mathfrak{L}_\alpha$. We use $\widehat{\mathfrak{L}}_\alpha$ to construct a basis of $\widehat{V}_n$, the span of $\widehat{\mathfrak{G}}_w$ with $w \in S_n$. Then we show $\widehat{V}_n$ gives a natural algebraic interpretation of a classical $q$-analogue of Bell numbers.

math.CO

A row analogue of Hecke column insertion

We introduce a new row insertion algorithm on decreasing tableaux and increasing tableaux, generalizing Edelman-Greene (EG) row insertion. Our row insertion algorithm is a nontrivial variation of Hecke column insertion which generalizes EG column insertion. Similar to Hecke column insertion, our row insertion is bijective and respects Hecke equivalence, and therefore recovers the expansions of stable Grothendieck functions into Grassmannian stable Grothendieck functions.

math.CO

Federated Graph-based Networks with Shared Embedding

Nowadays, user privacy is becoming an issue that cannot be bypassed for system developers, especially for that of web applications where data can be easily transferred through internet. Thankfully, federated learning proposes an innovative method to train models with distributed devices while data are kept in local storage. However, unlike general neural networks, although graph-based networks have achieved great success in classification tasks and advanced recommendation system, its high performance relies on the rich context provided by a graph structure, which is vulnerable when data attributes are incomplete. Therefore, the latter becomes a realistic problem when implementing federated learning for graph-based networks. Knowing that data embedding is a representation in a different space, we propose our Federated Graph-based Networks with Shared Embedding (Feras), which uses shared embedding data to train the network and avoids the direct sharing of original data. A solid theoretical proof of the convergence of Feras is given in this work. Experiments on different datasets (PPI, Flickr, Reddit) are conducted to show the efficiency of Feras for centralized learning. Finally, Feras enables the training of current graph-based models in the federated learning framework for privacy concern.

cs.LG

A bijection between $K$-Kohnert diagrams and reverse set-valued tableaux

Lascoux polynomials are $K$-theoretic analogues of the key polynomials. They both have combinatorial formulas involving tableaux: reverse set-valued tableaux ($\mathsf{RSVT}$) rule for Lascoux polynomials and reverse semistandard Young tableaux ($\mathsf{RSSYT}$) rule for key polynomials. Furthermore, key polynomials have a simple algorithmic model in terms of Kohnert diagrams, which are in bijection with $\mathsf{RSSYT}$. Ross and Yong introduced $K$-Kohnert diagrams, which are analogues of Kohnert diagrams. They conjectured a $K$-Kohnert diagram rule for Lascoux polynomials. We establish this conjecture by constructing a weight-preserving bijection between $\mathsf{RSVT}$ and $K$-Kohnert diagrams.

math.CO