arXiv ScienceSearch

arXiv subjects

Ruixi Lin

Publications and source records attributed to Ruixi Lin.

At least 19 recordsLinked to original sources

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy. JuryProbe estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift; when flagged high-risk, reference-free majority accepts are routed to the same judges with trusted references. On audited FEVER corruptions, reference-free panels show correlated false negatives (FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x), while unanimous false consensus drops to zero under a trusted-reference best-case diagnostic on both minimal-pair and non-minimal-pair evidence. In flagged settings, the routed policy is by construction equivalent to grounding every reference-free majority accept (verified in 34/34 splits): improvement comes from accept-conditioned grounding, while the diagnostic determines whether to activate it. A fixed, pre-specified rule flags 8-10 of 10 splits across synthetic, benchmark-authored, and scientific families and 0 of 10 on a negative control, where standing down avoids 28% of reference acquisitions at a 0.004 increase in false accepts. False-accept reduction persists under weak BM25 retrieval at substantial coverage cost, while stale stand-down labels require periodic recalibration. JuryProbe provides no formal risk guarantee and does not establish reliable stand-down on natural panels; its supported contribution is an empirical diagnostic of high-risk panel error dependence.

cs.CL

When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

Whether input-dependent ("dynamic") combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment. Can a small labeled target-domain probe tell us when reallocating trust across regions of the input space will pay off? We answer this with $\widehat{D}_{\mathrm{CF5}}$, which estimates from the probe the cross-fitted gain of the regionwise convex combination over the best static convex blend: the realizable value of deciding, region by region, whom to trust. Across a frozen suite of 12 dataset-shift pairs (spatial, temporal, domain, feature-cluster), $\widehat{D}_{\mathrm{CF5}}$ predicts realized regionwise test gains with dataset-level Spearman $+0.98$ (95% CI $[+0.83, +1.00]$; $p=5\times10^{-5}$), including two cases overturning preregistered expectations. The relationship holds in a 16-pair sensitivity analysis (Spearman $+0.83$), whereas alternative probe diagnostics reach at most $+0.66$. This contrast isolates regional trust reallocation: correlation is $+0.98$ for regionwise-convex gain, but $+0.01$ for smooth covariate-dependent stacking after affine correction. A controlled generator shows dynamic gains arise from the interaction of shift heterogeneity and local competence, increase with shift severity, and become realizable between 128 and 256 probe labels in the tested grid. The Probe-Validated Ensemble Selector chooses among a static affine stacker and dynamic realizers, deploying a candidate only when a held-out lower confidence bound clears the static-convex floor. In a preregistered prospective batch, it matched or improved the floor in all 12 runs; two deployments reduced test risk by 11% and 16%, while the gate rejected a candidate whose un-gated deployment incurred $>30\times$ the static loss. We release OpenRegShift, a reproducible evaluation harness for regression ensembles under distribution shift.

cs.LG

RSDM: The Consensus Honest Money in the AI Era

The medium of exchange of the traditional economy is mainly the fiat currency of each country or region, and when cross-border transactions occur, they need to be settled according to the exchange rate. In the AI world, however, the medium of exchange tends to be a globally recognized currency. Especially when AI acts as an agent for cross-border capital pool and cross cyclical asset allocation, it needs a sound money that can resist the depreciation of fiat currency and store long-term value. Therefore, we propose a globally consensus and universally accepted monetary rule framework for the AI era. The devaluation of money runs through almost the whole process of history, from the weight reduction and purity decrease of metallic coin to the unanchored over-issuance of paper currency. Whether it is the periodic compulsory recoinage in medieval Europe or Gesell's stamp scrip, both are essentially mechanisms for taxing money holdings. Unlike Gesell's stamp scrip, Redeemable Self-Decaying/Devaluing Money (RSDM) is a tokenized commodity money. Its essential innovation is to fill the hole in the storage fee of metal coins through the self-devaluing of metal weight recorded on the deposit certificate (warehouse receipt) of metal coins. In a sense, RSDM is an innovative version of Jiaozi (a deposit receipt for base metal coin that emerged in Sichuan, China, about a thousand years ago). In this paper, we propose five forms of online and offline issuance of RSDM, providing a prototype for creating a globally recognized modern honest money.

econ.GN

Discovering the Hidden Role of Gini Index In Prompt-based Classification

In classification tasks, the long-tailed minority classes usually offer the predictions that are most important. Yet these classes consistently exhibit low accuracies, whereas a few high-performing classes dominate the game. We pursue a foundational understanding of the hidden role of Gini Index as a tool for detecting and optimizing (debiasing) disparities in class accuracy, focusing on the case of prompt-based classification. We introduce the intuitions, benchmark Gini scores in real-world LLMs and vision models, and thoroughly discuss the insights of Gini not only as a measure of relative accuracy dominance but also as a direct optimization metric. Through rigorous case analyses, we first show that weak to strong relative accuracy imbalance exists in both prompt-based, text and image classification results and regardless of whether the classification is high-dimensional or low-dimensional. Then, we harness the Gini metric to propose a post-hoc model-agnostic bias mitigation method. Experimental results across few-shot news, biomedical, and zero-shot image classification show that our method significantly reduces both relative and absolute accuracy imbalances, minimizing top class relative dominance while elevating weakest classes.

cs.LG

A parallel monetary system based on the redeemable self-decaying money -- The ultimate hedge and safe haven of private wealth in the rising wave of over issuance of fiat and token money/stablecoin

A currency with stable purchasing power can always provide a psychological haven for people around the world. However, since the collapse of the Bretton Woods system, issuing more cheap currencies has become a common trend in the international community, and the legalization and over issuance of stablecoins will strengthen this trend. In this context, our study focused on a parallel monetary system based on a redeemable self-decay/devalued money(RSDM). Firstly, we point out the idea of redeeming gold at a fixed denomination with gold certificates is similar to an impossible perpetual motion machine. Only when the face value of a gold token self-decays or self-depreciates and the weight of the reduced value can compensate for the storage cost of physical gold, can it be convertible or redeemable. Secondly, we pointed out that as a modern "good money" under the Internet environment, it must have two basic functions: long-term value storage and zero logistics cost of money circulation. Thirdly, we found that a single type of money is difficult to shoulder the responsibility of modern "good money". Only a parallel monetary system, including RSDM, such as a triple-monetary system consisting of RSDM, domestic fiat and major international reserve currencies, can form the ultimate safe haven of wealth and safeguard the reverse Gresham law. Based on this analysis, we build an integer programming model for currency optimization selection in a multi-monetary pool. Fourthly, several potential application scenarios of RSDM in the real world were discussed, including a new approach to activate dormant gold assets in India based on RSDM, and the gold monetization scheme in the United States. Finally, the demand for RSDM with precious metals as collateral was analyzed, providing theoretical support for establishing a sound parallel monetary system based on RSDM.

q-fin.GN

Ensemble Debiasing Across Class and Sample Levels for Fairer Prompting Accuracy

Language models are strong few-shot learners and achieve good overall accuracy in text classification tasks, masking the fact that their results suffer from great class accuracy imbalance. We believe that the pursuit of overall accuracy should not come from enriching the strong classes, but from raising up the weak ones. To address the imbalance, we propose a Heaviside step function based ensemble debiasing method, which enables flexible rectifications of in-context learned class probabilities at both class and sample levels. Evaluations with Llama-2-13B on seven text classification benchmarks show that our approach achieves state-of-the-art overall accuracy gains with balanced class accuracies. More importantly, we perform analyses on the resulted probability correction scheme, showing that sample-level corrections are necessary to elevate weak classes. Due to effectively correcting weak classes, our method also brings significant performance gains to a larger model variant, Llama-2-70B, especially on a biomedical domain task, further demonstrating the necessity of ensemble debiasing at both levels. Our source code is available at https://github.com/NUS-HPC-AI-Lab/DCS.

cs.CL

Let the Fuzzy Rule Speak: Enhancing In-context Learning Debiasing with Interpretability

Large language models (LLMs) often struggle with balanced class accuracy in text classification tasks using in-context learning (ICL), hindering some practical uses due to user dissatisfaction or safety risks caused by misclassifications. Retraining LLMs to address root causes in data or model priors is neither easy nor cost-effective. This paper delves deeper into the class accuracy imbalance issue, identifying that it arises because certain classes consistently receive disproportionately high ICL probabilities, causing under-prediction and lower accuracy for others. More importantly, probability ranges affect the imbalance differently, allowing for precise, range-specific corrections. We introduce FuRud (Fuzzy Rule Optimization-based Debiasing), a method for sample-level class probability correction. FuRud tackles interpretability challenges by determining why certain classes need corrections and tailoring adjustments for each instance's class probabilities which is powered by fuzzy sets with triangular membership functions, transforming a class probability based on the range it belongs to. By solving a nonlinear integer programming problem with a labeled set of ICL class probabilities to minimize class accuracy bias (COBias) and maximize overall accuracy, each class selects an optimal correction function from 19 triangular membership functions without updating an LLM, and the selected functions correct test instances at inference. Across seven benchmark datasets, FuRud reduces COBias by over half (56%) and improves overall accuracy by 21% relatively, outperforming state-of-the-art debiasing methods.

cs.CL

Can a large language model be a gaslighter?

Large language models (LLMs) have gained human trust due to their capabilities and helpfulness. However, this in turn may allow LLMs to affect users' mindsets by manipulating language. It is termed as gaslighting, a psychological effect. In this work, we aim to investigate the vulnerability of LLMs under prompt-based and fine-tuning-based gaslighting attacks. Therefore, we propose a two-stage framework DeepCoG designed to: 1) elicit gaslighting plans from LLMs with the proposed DeepGaslighting prompting template, and 2) acquire gaslighting conversations from LLMs through our Chain-of-Gaslighting method. The gaslighting conversation dataset along with a corresponding safe dataset is applied to fine-tuning-based attacks on open-source LLMs and anti-gaslighting safety alignment on these LLMs. Experiments demonstrate that both prompt-based and fine-tuning-based attacks transform three open-source LLMs into gaslighters. In contrast, we advanced three safety alignment strategies to strengthen (by 12.05%) the safety guardrail of LLMs. Our safety alignment strategies have minimal impacts on the utility of LLMs. Empirical studies indicate that an LLM may be a potential gaslighter, even if it passed the harmfulness test on general dangerous queries.

cs.CR

Optimizing Class-Level Probability Reweighting Coefficients for Equitable Prompting Accuracy

Even as we engineer LLMs for alignment and safety, they often uncover biases from pre-training data's statistical regularities (from disproportionate co-occurrences to stereotypical associations mirroring human cognitive biases). This leads to persistent, uneven class accuracy in classification and QA. Such per-class accuracy disparities are not inherently resolved by architectural/training evolutions or data scaling, making post-hoc correction essential for equitable performance. To mitigate LLM class accuracy imbalance, we develop a post-hoc probability reweighting method that directly optimizes for non-differentiable performance-driven and fairness-aligned metrics, through a novel COBias metric that highlights disparities in class accuracies. This post-hoc bias mitigation method is grounded in discrete optimization with nonlinear integer programming (NIP) objectives and an efficient metaheuristic solution framework with theoretical convergence guarantees. Operating model-agnostically, it learns reweighting coefficients from output class probabilities to adjust LLM inference outputs without internal weight updates. Evaluations demonstrate its effectiveness: reducing COBias (61% relative reduction), increasing overall accuracy (18% relative increase), and achieving robust within-task generalization across diverse prompt configurations.

cs.CL

System Combination for Grammatical Error Correction Based on Integer Programming

In this paper, we propose a system combination method for grammatical error correction (GEC), based on nonlinear integer programming (IP). Our method optimizes a novel F score objective based on error types, and combines multiple end-to-end GEC systems. The proposed IP approach optimizes the selection of a single best system for each grammatical error type present in the data. Experiments of the IP approach on combining state-of-the-art standalone GEC systems show that the combined system outperforms all standalone systems. It improves F0.5 score by 3.61% when combining the two best participating systems in the BEA 2019 shared task, and achieves F0.5 score of 73.08%. We also perform experiments to compare our IP approach with another state-of-the-art system combination method for GEC, demonstrating IP's competitive combination capability.

cs.CL

The black hole of logistics costs of digitizing commodity money

In this paper, we reveal the depreciation mechanism of representative money (banknotes) from the perspective of logistics warehousing costs. Although it has long been the dream of economists to stabilize the buying power of the monetary units, the goal we have honest money always broken since the central bank depreciate the currency without limit. From the point of view of modern logistics, the key functions of money are the store of value and low logistics (circulation and warehouse) cost. Although commodity money (such as gold and silver) has the advantages of a wealth store, its disadvantage is the high logistics cost. In comparison to commodity money, credit currency and digital currency cannot protect wealth from loss over a long period while their logistics costs are negligible. We proved that there is not such honest money from the perspective of logistics costs, which is both the store of value like precious metal and without logistics costs in circulation like digital currency. The reason hidden in the back of the depreciation of banknotes is the black hole of storage charge of the anchor overtime after digitizing commodity money. Accordingly, it is not difficult to infer the inevitable collapse of the Bretton woods system. Therefore, we introduce a brand-new currency named honest devalued stable-coin and built a attenuation model of intrinsic value of the honest money based on the change mechanism of storage cost of anchor assets, like gold, which will lay the theoretical foundation for a stable monetary system.

q-fin.GN

Improved Robust ASR for Social Robots in Public Spaces

Social robots deployed in public spaces present a challenging task for ASR because of a variety of factors, including noise SNR of 20 to 5 dB. Existing ASR models perform well for higher SNRs in this range, but degrade considerably with more noise. This work explores methods for providing improved ASR performance in such conditions. We use the AiShell-1 Chinese speech corpus and the Kaldi ASR toolkit for evaluations. We were able to exceed state-of-the-art ASR performance with SNR lower than 20 dB, demonstrating the feasibility of achieving relatively high performing ASR with open-source toolkits and hundreds of hours of training data, which is commonly available.

eess.AS

Multi-Layer Ensembling Techniques for Multilingual Intent Classification

In this paper we determine how multi-layer ensembling improves performance on multilingual intent classification. We develop a novel multi-layer ensembling approach that ensembles both different model initializations and different model architectures. We also introduce a new banking domain dataset and compare results against the standard ATIS dataset and the Chinese SMP2017 dataset to determine ensembling performance in multilingual and multi-domain contexts. We run ensemble experiments across all three datasets, and conclude that ensembling provides significant performance increases, and that multi-layer ensembling is a no-risk way to improve performance on intent classification. We also find that a diverse ensemble of simple models can reach perform comparable to much more sophisticated state-of-the-art models. Our best F 1 scores on ATIS, Banking, and SMP are 97.54%, 91.79%, and 93.55% respectively, which compare well with the state-of-the-art on ATIS and best submission to the SMP2017 competition. The total ensembling performance increases we achieve are 0.23%, 1.96%, and 4.04% F 1 respectively.

cs.CL

Combining Word Feature Vector Method with the Convolutional Neural Network for Slot Filling in Spoken Language Understanding

Slot filling is an important problem in Spoken Language Understanding (SLU) and Natural Language Processing (NLP), which involves identifying a user's intent and assigning a semantic concept to each word in a sentence. This paper presents a word feature vector method and combines it into the convolutional neural network (CNN). We consider 18 word features and each word feature is constructed by merging similar word labels. By introducing the concept of external library, we propose a feature set approach that is beneficial for building the relationship between a word from the training dataset and the feature. Computational results are reported using the ATIS dataset and comparisons with traditional CNN as well as bi-directional sequential CNN are also presented.

cs.CL

Enhancing Chinese Intent Classification by Dynamically Integrating Character Features into Word Embeddings with Ensemble Techniques

Intent classification has been widely researched on English data with deep learning approaches that are based on neural networks and word embeddings. The challenge for Chinese intent classification stems from the fact that, unlike English where most words are made up of 26 phonologic alphabet letters, Chinese is logographic, where a Chinese character is a more basic semantic unit that can be informative and its meaning does not vary too much in contexts. Chinese word embeddings alone can be inadequate for representing words, and pre-trained embeddings can suffer from not aligning well with the task at hand. To account for the inadequacy and leverage Chinese character information, we propose a low-effort and generic way to dynamically integrate character embedding based feature maps with word embedding based inputs, whose resulting word-character embeddings are stacked with a contextual information extraction module to further incorporate context information for predictions. On top of the proposed model, we employ an ensemble method to combine single models and obtain the final result. The approach is data-independent without relying on external sources like pre-trained word embeddings. The proposed model outperforms baseline models and existing methods.

cs.CL

An Approach to the High-level Maintenance Planning for EMU Trains Based on Simulated Annealing

A high-speed train needs high-level maintenance when its accumulated running mileage or time reaches predefined threshold. The date of delivering an Electric Multiple Unit (EMU) train to maintenance ranges within a time window rather than be a fixed date. Obviously, changing the delivering date always means a different impact on the supply of EMU trains and operation cost. Therefore, the delivering plan has the potential to be optimized. This paper formulates the EMU train high-level maintenance planning problem as a non-linear 0-1 programming model. The model aims at minimizing the mileage loss of all EMU trains with the consideration of the maintenance capacity of the workshop and maintenance ratio at different times. The number of trains under maintenance not only depends on the current maintenance plan, but also influenced by the trains whose maintenance time span from the last planning horizon to current horizon. A state function is established to describe whether a train is under maintenance. By using this function the constraint of restricting the total number of trains that are under maintenance can be formulated reasonably well. Finally, a simulated annealing algorithm is proposed for solving the problem.

math.OC

Sentiment Classification of Food Reviews

Sentiment analysis of reviews is a popular task in natural language processing. In this work, the goal is to predict the score of food reviews on a scale of 1 to 5 with two recurrent neural networks that are carefully tuned. As for baseline, we train a simple RNN for classification. Then we extend the baseline to GRU. In addition, we present two different methods to deal with highly skewed data, which is a common problem for reviews. Models are evaluated using accuracies.

cs.CL

A New Currency of the Future: The Novel Commodity Money with Attenuation Coefficient Based on the Logistics Cost of Anchor

In this paper, we reveal the attenuation mechanism of anchor of the commodity money from the perspective of logistics warehousing costs, and propose a novel Decayed Commodity Money (DCM) for the store of value across time and space. Considering the logistics cost of commodity warehousing by the third financial institution such as London Metal Exchange, we can award the difference between the original and the residual value of the anchor to the financial institution. This type of currency has the characteristic of self-decaying value over time. Therefore DCM has the advantages of both the commodity money which has the function of preserving wealth and credit currency without the logistics cost. In addition, DCM can also avoid the defects that precious metal money is hoarded by market and credit currency often leads to excessive liquidity. DCM is also different from virtual currency, such as bitcoin, which does not have a corresponding commodity anchor. As a conclusion, DCM can provide a new way of storing wealth for nations, corporations and individuals effectively.

q-fin.GN