arXiv ScienceSearch

arXiv subjects

Han Ying Lim

Publications and source records attributed to Han Ying Lim.

3 recordsLinked to original sources

Variable-Projection Sparse Functional Principal Component Analysis: Interpretable Functional Dimensionality Reduction with Applications to Raman Spectral Data

Functional principal component analysis (FPCA) provides low-rank representations of functional data but generally produces dense components, making it difficult to identify the localised regions contributing to dominant modes of variation. This limitation is particularly relevant in Raman spectroscopy, where spectra are observed over an ordered domain and interpretation often focuses on chemically meaningful spectral regions. This study proposes variable-projection sparse FPCA (VP-SFPCA) for interpretable functional dimensionality reduction through sparse weight functions that promote localisation. The method formulates sparse FPCA as a regularised matrix-factorisation problem that incorporates the functional inner-product geometry and distinguishes sparse weight functions used to generate component scores from orthonormal loading functions used for reconstruction. Variable projection conditionally minimises over the loading functions, reducing the optimisation problem to the sparse weights. Performance was evaluated through simulation studies and empirical analyses of surface-enhanced Raman scattering (SERS) spectra, with conventional FPCA and SCAD-SFPCA serving as dense and sparse functional benchmarks, respectively. In the simulations, VP-SFPCA recovered localised functional structure while requiring substantially less computation than SCAD-SFPCA. In the empirical analysis, VP-SFPCA retained this computational advantage while yielding held-out reconstruction error close to that of conventional FPCA. Several prominent features of the estimated weight functions also coincided with established adenine SERS bands. Overall, VP-SFPCA provides a computationally practical approach to improving the interpretability of dominant functional modes through localisation.

stat.ME

Hidden Consensus:Preference-Validity Compression in Human Feedback

Standard RLHF pipelines often reduce heterogeneous human judgments into a single scalar reward target. We argue that this reduction can mis-measure alignment in structurally plural societies, where disagreement may reflect culturally, historically, linguistically, regionally, or normatively grounded interpretations rather than annotation noise. We call this failure Preference-Validity Compression, the collapse of multiple plural-valid response options into a single optimization target. Using Malaysia as a diagnostic setting, we analyze RLHF-style feedback aggregation through preference events linking prompts, responses, and acceptability judgments across interpretive frames. Across 321 preference events from 20 participants and 107 trio-annotated prompts, 79% of prompts contain more than one majority-supported response that single-winner aggregation would discard, and apparent dominance gaps between top responses diminish when all majority-supported options are considered. Participants frequently select multiple acceptable responses, and discarded responses demonstrably reflect coherent local, practical, or cultural frames. These findings show that majority aggregation in this corpus measures argmax acceptability rather than plural alignment. We treat this as a measurement-validity issue and argue that future alignment methods should satisfy Validity-Preserving Consistency, remaining stable across plural-valid interpretive frames rather than collapsing them into a single reward target.

cs.CL

Compositional data analysis for modelling and forecasting mortality using the α-transformation

Mortality forecasting is crucial for demographic planning and actuarial studies, especially for projecting population ageing and longevity risk. Classical approaches largely rely on extrapolative methods, such as the Lee-Carter (LC) model, which use mortality rates as the mortality measure. In recent years, compositional data analysis (CoDA), which respects summability and non-negativity constraints, has gained increasing attention for mortality forecasting. While the centred log-ratio (CLR) transformation is commonly used to map compositional data to real space, the α-transformation, a generalisation of log-ratio transformations, offers greater flexibility and adaptability. This study contributes to mortality forecasting by introducing the α-transformation as an alternative to the CLR transformation within a non-functional CoDA model that has not been previously investigated in existing literature. To fairly compare the impact of transformation choices on forecast accuracy, zero values in the data are imputed, although the α-transformation can inherently handle them. Using age-specific life table death counts for males and females in 31 selected European countries/regions from 1983 to 2018, the proposed method demonstrates comparable performance to the CLR transformation in most cases, with improved forecast accuracy in some instances. These findings highlight the potential of the α-transformation for enhancing mortality forecasting within the non-functional CoDA framework.

stat.AP