arXiv ScienceSearch

arXiv subjects

Aman Verma

Publications and source records attributed to Aman Verma.

17 recordsLinked to original sources

Comparative Analysis of Low-Rank Adaptation in Large Language Models versus Dense Embedding Regression for Headline Click-Through Rate Prediction

Optimizing digital content headlines for click-through rate (CTR) is an important problem in online media and recommendation systems. While large language models (LLMs) have demonstrated strong generative capabilities, their effectiveness for discriminative ranking tasks, such as selecting the highest-performing headline from a set of candidates, remains less well understood. In this work, we compare a LoRA-fine-tuned causal language model, LOLAQwen (0.6B), with a dense embedding regression model for headline selection. We formulate headline selection as a winner-take-all classification problem and evaluate both approaches using a dataset of 3,263 A/B-tested headline groups. Performance is measured using Top-1 accuracy, defined as the proportion of groups for which the model correctly identifies the highest-performing headline. The embedding regression model achieves a Top-1 accuracy of 42.79%, compared with 35.70% for the LoRA-fine-tuned language model. These results indicate that, for this headline selection task, a lightweight discriminative approach can outperform a small generative language model fine-tuned using parameter-efficient adaptation. The findings highlight the potential of embedding-based regression models as efficient alternatives to generative models for high-throughput content ranking applications.

cs.CE

Dynamic Windowing in Transformers via Regime Incorporation for Financial Time Series

Financial time series exhibit non-stationary behavior, where the strength and extent of temporal dependencies vary across market regimes. Trending, low-volatility phases typically require long-range contextual information, whereas mean-reverting, high-volatility periods rely more heavily on short-term dynamics. Standard Transformer architectures, with fixed attention windows and static positional encodings, are therefore unable to adapt to such variations. In this work, we propose a regime-aware dynamic windowing framework that incorporates market regime information directly into the Transformer. We construct four generic regime signals from price series: volatility ratio, trend strength, local predictability ratio (LPR), and rolling autocorrelation. We incorporate these signals into the model through two mechanisms: (i) regime-augmented inputs to a standard Transformer architecture, and (ii) a modified attention layer that modulates attention weights using regime embeddings. Experiments on five S&P 500 stocks show consistent improvements across five evaluation metrics, demonstrating that regime-aware dynamic windowing enhances both interpretability and predictive performance in financial forecasting tasks.

cs.CE

GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence

Remote-sensing vision-language models (RS-VLMs) have advanced Earth-observation analysis toward visual interpretation and instruction-following, yet fall short of operational geo-intelligence, which demands tool-grounded spatial reasoning and structured, evidence-backed decisions. We introduce GeoDisaster, an operational geospatial disaster reasoning benchmark with 2,921 verified instances across 43 question types and five task families: deforestation monitoring, multi-hazard analysis, building-damage assessment, flood-safe routing, and Sentinel-1 SAR flood monitoring. Instances integrate heterogeneous EO/GIS evidence-optical and SAR imagery, raster masks, vector geometries, road networks, and exposure layers-spanning hazard detection, damage assessment, exposure estimation, and diagnostic report generation. Ground-truth answers are grounded in executable geospatial workflows and deterministic consistency checks, removing the need for language-model annotation. We further propose an orchestrated multi-agent framework with 18 disaster-oriented tools, where role-specialized agents coordinate through explicit execution contracts, aligned via Role-Contract Expectation Alignment (RCEA): failure-aware supervised fine-tuning combined with contract-grounded reinforcement learning over dense step-level signals. Experiments show that GeoDisaster challenges existing RS-VLMs and agentic systems, while RCEA improves tool use, evidence grounding, state consistency, and decision generation.

cs.CV

Advanced Acceptance Score: A Holistic Measure for Biometric Quantification

Quantifying biometric characteristics within hand gestures involve derivation of fitness scores from a gesture and identity aware feature space. However, evaluating the quality of these scores remains an open question. Existing biometric capacity estimation literature relies upon error rates. But these rates do not indicate goodness of scores. Thus, in this manuscript we present an exhaustive set of evaluation measures. We firstly identify ranking order and relevance of output scores as the primary basis for evaluation. In particular, we consider both rank deviation as well as rewards for: (i) higher scores of high ranked gestures and (ii) lower scores of low ranked gestures. We also compensate for correspondence between trends of output and ground truth scores. Finally, we account for disentanglement between identity features of gestures as a discounting factor. Integrating these elements with adequate weighting, we formulate advanced acceptance score as a holistic evaluation measure. To assess effectivity of the proposed we perform in-depth experimentation over three datasets with five state-of-the-art (SOTA) models. Results show that the optimal score selected with our measure is more appropriate than existing other measures. Also, our proposed measure depicts correlation with existing measures. This further validates its reliability. We have made our \href{https://github.com/AmanVerma2307/MeasureSuite}{code} public.

cs.CV

CPEP: Contrastive Pose-EMG Pre-training Enhances Gesture Generalization on EMG Signals

Hand gesture classification using high-quality structured data such as videos, images, and hand skeletons is a well-explored problem in computer vision. Leveraging low-power, cost-effective biosignals, e.g. surface electromyography (sEMG), allows for continuous gesture prediction on wearables. In this paper, we demonstrate that learning representations from weak-modality data that are aligned with those from structured, high-quality data can improve representation quality and enables zero-shot classification. Specifically, we propose a Contrastive Pose-EMG Pre-training (CPEP) framework to align EMG and pose representations, where we learn an EMG encoder that produces high-quality and pose-informative representations. We assess the gesture classification performance of our model through linear probing and zero-shot setups. Our model outperforms emg2pose benchmark models by up to 21% on in-distribution gesture classification and 72% on unseen (out-of-distribution) gesture classification.

cs.LG

Multimodal Real-Time Anomaly Detection and Industrial Applications

This paper presents the design, implementation, and evolution of a comprehensive multimodal room-monitoring system that integrates synchronized video and audio processing for real-time activity recognition and anomaly detection. We describe two iterations of the system: an initial lightweight implementation using YOLOv8, ByteTrack, and the Audio Spectrogram Transformer (AST), and an advanced version that incorporates multi-model audio ensembles, hybrid object detection, bidirectional cross-modal attention, and multi-method anomaly detection. The evolution demonstrates significant improvements in accuracy, robustness, and industrial applicability. The advanced system combines three audio models (AST, Wav2Vec2, and HuBERT) for comprehensive audio understanding, dual object detectors (YOLO and DETR) for improved accuracy, and sophisticated fusion mechanisms for enhanced cross-modal learning. Experimental evaluation shows the system's effectiveness in general monitoring scenarios as well as specialized industrial safety applications, achieving real-time performance on standard hardware while maintaining high accuracy.

cs.SD

Role of inefficient measurement in realizing post-selection-based non-Hermitian qubits

Post-selecting against quantum jumps into the ground state confines the evolution of the three-level system to the excited states manifold, effectively realizing a PT-symmetric non-Hermitian qubit. In this work, by introducing post-selection efficiencies for both decay channels, the second-excited to first-excited and the first-excited to ground-state transitions, we formulate a hybrid-Liouvillian framework that captures the unmonitored dynamics of the non-Hermitian qubit. We find that the decoherence effects arising from quantum jumps within the second-excited and first-excited manifold also manifest under inefficient post-selection of the second-excited to first-excited transitions, thereby modifying the spectral properties of the Liouvillian and leading to a splitting of the exceptional points. A comparative analysis shows that the trajectory-based approach, obtained by ensemble-averaging stochastic measurement trajectories generated via the Bayesian state update rule, and the Lindblad evolution remain consistent. Our results highlight the fundamental role of measurement inefficiency in realizing post-selection-based non-Hermitian qubits and in shaping the structure of Liouvillian exceptional points. These findings provide new insights into how inefficient measurement processes influence non-Hermitian behavior in open quantum systems.

quant-ph

MFConvTr: Multi-Frequency Convolutional Transformer for Fetal Arrhythmia Detection in Non-Invasive fECG

NI-fECG have emerged as alternative for fetal arrhythmia monitoring. But due to multi-signal waveform they are tough to understand and due to highly varying and complex nature traditional fiducial methods cannot be applied. Further, it has also been observed that the fetal arrhythmia can be differentiated from the normal signals in both spectral and temporal scales. To this end, we propose Multi-Frequency Convolutional Transformer, a novel deep learning architecture that learns information in contexts with multiple-frequency and can model long-term dependencies. The proposed model utilizes a convolutional-backbone consisting of model Multi-Frequency Convolutions (MF-Conv) and residual connections. MF-Conv in-turn captures multi-frequency contexts in an efficient manner by splitting the input channel and then convoluting each of the splits individually with different kernel size. Accredited to these properties, the proposed model attains state-of-the-art results and that too utilizing very low number of parameters. To evaluate the proposed we also perform extensive ablation studies.

eess.SP

Towards reconstruction of Pulsed-wave Doppler signals from Non-invasive fetal ECG

Fetal cardiac health monitoring with invasive methods have a limited viability because they can only be utilized during labor and are uncomfortable. On the other hand non-invasive fECG are adulterated with maternal ECG, and hence resulting in poor analysis. In contrast, Pulsed-wave Doppler (PwD) echocardiography generates high-quality signals representing fetal blood volume inflow-outflow. It also follows non-invasive signal acquisition. The only drawback is that it requires highly expensive setup. To address this aspect, we put forward a challenging research question - can we reconstruct PwD signals using non-invasive fetal ECG?

eess.SP

EEG-Based Mental Imagery Task Adaptation via Ensemble of Weight-Decomposed Low-Rank Adapters

Electroencephalography (EEG) is widely researched for neural decoding in Brain Computer Interfaces (BCIs) as it is non-invasive, portable, and economical. However, EEG signals suffer from inter- and intra-subject variability, leading to poor performance. Recent technological advancements have led to deep learning (DL) models that have achieved high performance in various fields. However, such large models are compute- and resource-intensive and are a bottleneck for real-time neural decoding. Data distribution shift can be handled with the help of domain adaptation techniques of transfer learning (fine-tuning) and adversarial training that requires model parameter updates according to the target domain. One such recent technique is Parameter-efficient fine-tuning (PEFT), which requires only a small fraction of the total trainable parameters compared to fine-tuning the whole model. Therefore, we explored PEFT methods for adapting EEG-based mental imagery tasks. We considered two mental imagery tasks: speech imagery and motor imagery, as both of these tasks are instrumental in post-stroke neuro-rehabilitation. We proposed a novel ensemble of weight-decomposed low-rank adaptation methods, EDoRA, for parameter-efficient mental imagery task adaptation through EEG signal classification. The performance of the proposed PEFT method is validated on two publicly available datasets, one speech imagery, and the other motor imagery dataset. In extensive experiments and analysis, the proposed method has performed better than full fine-tune and state-of-the-art PEFT methods for mental imagery EEG classification.

eess.SP

A Mini Review on The Applications of Nanomaterials in Forensic Science

Herein, we report a minireview to give a brief introduction of applications of nanomaterials in the field of forensic science. The materials that have their size in nanoscale (1 - 100 nm) comes under the category of nanomaterials. Nanomaterials possess various applications in different fields like cosmetic production, medical, photoconductivity etc. because of their physio-chemical, electrical and magnetic properties. Due to the different characteristic property that nanomaterials have, they are widely employed in diverse domains. In various fields of forensic science such as fingerprints, toxicology, medicine, serology, nanomaterials are being used extensively. Large surface area to volume ratio of the materials in nano-regime makes the nanomaterials suitable for all these application with high efficiency. This review article briefs about the nanomaterials, their advantages and their novel applications in various fields, focusing especially in the field of forensic science. The basic idea of different areas of forensic science such as development of fingerprints, detection of drugs, estimating the time since death, analysis of GSR, detection of various explosives and for the extraction of DNA etc. has also been provided.

cond-mat.mtrl-sci

Modeling electronic health record data using a knowledge-graph-embedded topic model

The rapid growth of electronic health record (EHR) datasets opens up promising opportunities to understand human diseases in a systematic way. However, effective extraction of clinical knowledge from the EHR data has been hindered by its sparsity and noisy information. We present KG-ETM, an end-to-end knowledge graph-based multimodal embedded topic model. KG-ETM distills latent disease topics from EHR data by learning the embedding from the medical knowledge graphs. We applied KG-ETM to a large-scale EHR dataset consisting of over 1 million patients. We evaluated its performance based on EHR reconstruction and drug imputation. KG-ETM demonstrated superior performance over the alternative methods on both tasks. Moreover, our model learned clinically meaningful graph-informed embedding of the EHR codes. In additional, our model is also able to discover interpretable and accurate patient representations for patient stratification and drug recommendations.

cs.LG

HSADML: Hyper-Sphere Angular Deep Metric based Learning for Brain Tumor Classification

Brain Tumors are abnormal mass of clustered cells penetrating regions of brain. Their timely identification and classification help doctors to provide appropriate treatment. However, Classifi-cation of Brain Tumors is quite intricate because of high-intra class similarity and low-inter class variability. Due to morphological similarity amongst various MRI-Slices of different classes the challenge deepens more. This all leads to hampering generalizability of classification models. To this end, this paper proposes HSADML, a novel framework which enables deep metric learning (DML) using SphereFace Loss. SphereFace loss embeds the features into a hyperspheric-manifold and then imposes margin on the embeddings to enhance differentiability between the classes. With utilization of SphereFace loss based deep metric learning it is ensured that samples from class clustered together while the different ones are pushed apart. Results reflects the promi-nence in the approach, the proposed framework achieved state-of-the-art 98.69% validation accu-racy using k-NN (k=1) and this is significantly higher than normal SoftMax Loss training which though obtains 98.47% validation accuracy but that too with limited inter-class separability and intra-class closeness. Experimental analysis done over various classifiers and loss function set-tings suggests potential in the approach.

cs.CV

Design, Manufacturing, and Controls of a Prismatic Quadruped Robot: PRISMA

Most of the quadrupeds developed are highly actuated, and their control is hence quite cumbersome. They need advanced electronics equipment to solve convoluted inverse kinematic equations continuously. In addition, they demand special and costly sensors to autonomously navigate through the environment as traditional distance sensors usually fail because of the continuous perturbation due to the motion of the robot. Another challenge is maintaining the continuous dynamic stability of the robot while walking, which requires complicated and state-of-the-art control algorithms. This paper presents a thorough description of the hardware design and control architecture of our in-house prismatic joint quadruped robot called the PRISMA. We aim to forge a robust and kinematically stable quadruped robot that can use elementary control algorithms and utilize conventional sensors to navigate an unknown environment. We discuss the benefits and limitations of the robot in terms of its motion, different foot trajectories, manufacturability, and controls.

cs.RO

DFCANet: Dense Feature Calibration-Attention Guided Network for Cross Domain Iris Presentation Attack Detection

An iris presentation attack detection (IPAD) is essential for securing personal identity is widely used iris recognition systems. However, the existing IPAD algorithms do not generalize well to unseen and cross-domain scenarios because of capture in unconstrained environments and high visual correlation amongst bonafide and attack samples. These similarities in intricate textural and morphological patterns of iris ocular images contribute further to performance degradation. To alleviate these shortcomings, this paper proposes DFCANet: Dense Feature Calibration and Attention Guided Network which calibrates the locally spread iris patterns with the globally located ones. Uplifting advantages from feature calibration convolution and residual learning, DFCANet generates domain-specific iris feature representations. Since some channels in the calibrated feature maps contain more prominent information, we capitalize discriminative feature learning across the channels through the channel attention mechanism. In order to intensify the challenge for our proposed model, we make DFCANet operate over nonsegmented and non-normalized ocular iris images. Extensive experimentation conducted over challenging cross-domain and intra-domain scenarios highlights consistent outperforming results. Compared to state-of-the-art methods, DFCANet achieves significant gains in performance for the benchmark IIITD CLI, IIIT CSD and NDCLD13 databases respectively. Further, a novel incremental learning-based methodology has been introduced so as to overcome disentangled iris-data characteristics and data scarcity. This paper also pursues the challenging scenario that considers soft-lens under the attack category with evaluation performed under various cross-domain protocols. The code will be made publicly available.

cs.CV

Supervised multi-specialist topic model with applications on large-scale electronic health record data

Motivation: Electronic health record (EHR) data provides a new venue to elucidate disease comorbidities and latent phenotypes for precision medicine. To fully exploit its potential, a realistic data generative process of the EHR data needs to be modelled. We present MixEHR-S to jointly infer specialist-disease topics from the EHR data. As the key contribution, we model the specialist assignments and ICD-coded diagnoses as the latent topics based on patient's underlying disease topic mixture in a novel unified supervised hierarchical Bayesian topic model. For efficient inference, we developed a closed-form collapsed variational inference algorithm to learn the model distributions of MixEHR-S. We applied MixEHR-S to two independent large-scale EHR databases in Quebec with three targeted applications: (1) Congenital Heart Disease (CHD) diagnostic prediction among 154,775 patients; (2) Chronic obstructive pulmonary disease (COPD) diagnostic prediction among 73,791 patients; (3) future insulin treatment prediction among 78,712 patients diagnosed with diabetes as a mean to assess the disease exacerbation. In all three applications, MixEHR-S conferred clinically meaningful latent topics among the most predictive latent topics and achieved superior target prediction accuracy compared to the existing methods, providing opportunities for prioritizing high-risk patients for healthcare services. MixEHR-S source code and scripts of the experiments are freely available at https://github.com/li-lab-mcgill/mixehrS

cs.LG

Modeling disease progression in longitudinal EHR data using continuous-time hidden Markov models

Modeling disease progression in healthcare administrative databases is complicated by the fact that patients are observed only at irregular intervals when they seek healthcare services. In a longitudinal cohort of 76,888 patients with chronic obstructive pulmonary disease (COPD), we used a continuous-time hidden Markov model with a generalized linear model to model healthcare utilization events. We found that the fitted model provides interpretable results suitable for summarization and hypothesis generation.

cs.LG