arXiv ScienceSearch

arXiv subjects

Siddharth

Publications and source records attributed to Siddharth.

8 recordsLinked to original sources

SherpaAI: A Multi-modal Solution for Delivering Personalized and Adaptive Fitness Interventions

Personalization of exercise routines is a crucial factor in helping people achieve their fitness goals. Despite this, many contemporary solutions fail to offer real-time, adaptive feedback tailored to an individual's physiological states. Contemporary solutions often rely only on static, pre-set plans and rarely adjust in real time to factors such as a user's pain thresholds, fatigue levels, or form during a workout. This work introduces SherpaAI, a multi-modal system that unifies computer vision, physiological sensing (heart rate and voice), and the reasoning capabilities of Large Language Models (LLMs)---modalities that prior systems have largely explored in isolation---to deliver real-time and individually-adaptive guidance across a set of strength, balance, and flexibility exercises. SherpaAI continuously monitors a user's physical form and level of exertion, among other parameters, to provide dynamic interventions focused on exercise intensity, rest periods, and motivation. To validate our system, we performed a technical evaluation confirming our models' accuracy and quantifying pipeline latency, alongside an expert review where certified trainers validated the correctness of the LLM's interventions. Furthermore, in a controlled within-subject study with 25 participants, SherpaAI demonstrated significant improvements over a non-adaptive baseline modeled on typical self-guided workouts (e.g., following online videos or general-purpose AI chat tools for guidance). With SherpaAI, users reported significantly greater enjoyment, a stronger sense of achievement, and significantly lower levels of boredom and frustration. These results indicate that by integrating multi-modal sensing with LLM-driven reasoning, adaptive systems like SherpaAI can create a more engaging and emotionally satisfying workout experience.

cs.HC

Financial Instruction Following Evaluation (FIFE)

Language Models (LMs) struggle with complex, interdependent instructions, particularly in high-stakes domains like finance where precision is critical. We introduce FIFE, a novel, high-difficulty benchmark designed to assess LM instruction-following capabilities for financial analysis tasks. FIFE comprises 88 human-authored prompts and employs a verification system with chainable, verifiable constraints for fine-grained reward signals. We evaluate 53 models (proprietary, open-weight, open-source) in a zero-shot setting. Our key findings reveal a clear performance hierarchy: the top open-weight model (76.1 strict / 79.5 loose) surpasses the leading proprietary system (65.9 strict / 70.5 loose), while the best open-source models lag significantly (45.5 strict / 48.9 loose). However, even top-performing models struggle with FIFE's complex requirements, failing to achieve perfect compliance. We release our dataset and code as an open-source resource to promote research in Reinforcement Learning for the financial domain.

cs.LG

First-principles many-body study for electronic, optical, and excitonic properties of RbTlCl3 perovskite for solar cells

We present a detailed many-body ab initio study of the valence-skipper RbTlCl$_{3}$ perovskite compound for photovoltaic (PV) applications. The electronic and optical properties, both with and without spin-orbit coupling, have been calculated using density functional theory (DFT) and many-body excited-state calculations. The band gap, which is indirect in nature, is found to be 0.95 eV and 0.89 eV from PBE and PBEsol, respectively. The optical properties have been computed using four different approximations: independent particle approximation (IPA), IPA with scissor correction (IQPA), random phase approximation for local-field effects (LFEs), and the Bethe-Salpeter equation (BSE). The estimated highest value of the imaginary part of the dielectric function using IQPA is 7 at 2 eV, which slightly decreases to 5.7 due to LFEs. Within BSE, the peak value is obtained to be maximum at 1.6 eV with a magnitude of 10.8, which indicates the strong excitonic effect below the optical gap. Large number of bright and dark bound excitons are found, where the binding energies of four main bound bright excitons are found in the range of 299-350 meV. The exciton amplitude in both reciprocal and real space is analyzed. The main bound bright exciton is localized in the reciprocal space, while this exhibits a delocalized nature in real space. The BSE predicts a highest absorption coefficient of 3.6 $\times$ $10^{6}$ cm$^{-1}$ at 1.7 eV, while a minimum reflectivity in the active region of the solar energy spectrum is obtained to be around 2.7\%. Finally, the solar efficiency has been estimated using the spectroscopic limited maximum efficiency approach and obtained highest value is 15.5% at a thickness of 0.5 $\mu$m. These findings reveal a significant excitonic effect in the absorption spectra of RbTlCl$_{3}$ and highlight its potential as a promising material for single-junction thin-film solar cells.

cond-mat.mtrl-sci

Saliency-guided Emotion Modeling: Predicting Viewer Reactions from Video Stimuli

Understanding the emotional impact of videos is crucial for applications in content creation, advertising, and Human-Computer Interaction (HCI). Traditional affective computing methods rely on self-reported emotions, facial expression analysis, and biosensing data, yet they often overlook the role of visual saliency -- the naturally attention-grabbing regions within a video. In this study, we utilize deep learning to introduce a novel saliency-based approach to emotion prediction by extracting two key features: saliency area and number of salient regions. Using the HD2S saliency model and OpenFace facial action unit analysis, we examine the relationship between video saliency and viewer emotions. Our findings reveal three key insights: (1) Videos with multiple salient regions tend to elicit high-valence, low-arousal emotions, (2) Videos with a single dominant salient region are more likely to induce low-valence, high-arousal responses, and (3) Self-reported emotions often misalign with facial expression-based emotion detection, suggesting limitations in subjective reporting. By leveraging saliency-driven insights, this work provides a computationally efficient and interpretable alternative for emotion modeling, with implications for content creation, personalized media experiences, and affective computing research.

cs.CV

PosePilot: An Edge-AI Solution for Posture Correction in Physical Exercises

Automated pose correction remains a significant challenge in AI-driven fitness systems, despite extensive research in activity recognition. This work presents PosePilot, a novel system that integrates pose recognition with real-time personalized corrective feedback, overcoming the limitations of traditional fitness solutions. Using Yoga, a discipline requiring precise spatio-temporal alignment as a case study, we demonstrate PosePilot's ability to analyze complex physical movements. Designed for deployment on edge devices, PosePilot can be extended to various at-home and outdoor exercises. We employ a Vanilla LSTM, allowing the system to capture temporal dependencies for pose recognition. Additionally, a BiLSTM with multi-head Attention enhances the model's ability to process motion contexts, selectively focusing on key limb angles for accurate error detection while maintaining computational efficiency. As part of this work, we introduce a high-quality video dataset used for evaluating our models. Most importantly, PosePilot provides instant corrective feedback at every stage of a movement, ensuring precise posture adjustments throughout the exercise routine. The proposed approach 1) performs automatic human posture recognition, 2) provides personalized posture correction feedback at each instant which is crucial in Yoga, and 3) offers a lightweight and robust posture correction model feasible for deploying on edge devices in real-world environments.

cs.CV

Attention Monitoring and Hazard Assessment with Bio-Sensing and Vision: Empirical Analysis Utilizing CNNs on the KITTI Dataset

Assessing the driver's attention and detecting various hazardous and non-hazardous events during a drive are critical for driver's safety. Attention monitoring in driving scenarios has mostly been carried out using vision (camera-based) modality by tracking the driver's gaze and facial expressions. It is only recently that bio-sensing modalities such as Electroencephalogram (EEG) are being explored. But, there is another open problem which has not been explored sufficiently yet in this paradigm. This is the detection of specific events, hazardous and non-hazardous, during driving that affects the driver's mental and physiological states. The other challenge in evaluating multi-modal sensory applications is the absence of very large scale EEG data because of the various limitations of using EEG in the real world. In this paper, we use both of the above sensor modalities and compare them against the two tasks of assessing the driver's attention and detecting hazardous vs. non-hazardous driving events. We collect user data on twelve subjects and show how in the absence of very large-scale datasets, we can still use pre-trained deep learning convolution networks to extract meaningful features from both of the above modalities. We used the publicly available KITTI dataset for evaluating our platform and to compare it with previous studies. Finally, we show that the results presented in this paper surpass the previous benchmark set up in the above driver awareness-related applications.

cs.HC

An Affordable Bio-Sensing and Activity Tagging Platform for HCI Research

We present a novel multi-modal bio-sensing platform capable of integrating multiple data streams for use in real-time applications. The system is composed of a central compute module and a companion headset. The compute node collects, time-stamps and transmits the data while also providing an interface for a wide range of sensors including electroencephalogram, photoplethysmogram, electrocardiogram, and eye gaze among others. The companion headset contains the gaze tracking cameras. By integrating many of the measurements systems into an accessible package, we are able to explore previously unanswerable questions ranging from open-environment interactions to emotional response studies. Though some of the integrated sensors are designed from the ground-up to fit into a compact form factor, we validate the accuracy of the sensors and find that they perform similarly to, and in some cases better than, alternatives.

cs.HC

Driver Hand Localization and Grasp Analysis: A Vision-based Real-time Approach

Extracting hand regions and their grasp information from images robustly in real-time is critical for occupants' safety and in-vehicular infotainment applications. It must however, be noted that naturalistic driving scenes suffer from rapidly changing illumination and occlusion. This is aggravated by the fact that hands are highly deformable objects, and change in appearance frequently. This work addresses the task of accurately localizing driver hands and classifying the grasp state of each hand. We use a fast ConvNet to first detect likely hand regions. Next, a pixel-based skin classifier that takes into account the global illumination changes is used to refine the hand detections and remove false positives. This step generates a pixel-level mask for each hand. Finally, we study each such masked regions and detect if the driver is grasping the wheel, or in some cases a mobile phone. Through evaluation we demonstrate that our method can outperform state-of-the-art pixel based hand detectors, while running faster (at 35 fps) than other deep ConvNet based frameworks even for grasp analysis. Hand mask cues are shown to be crucial when analyzing a set of driver hand gestures (wheel/mobile phone grasp and no-grasp) in naturalistic driving settings. The proposed detection and localization pipeline hence can act as a general framework for real-time hand detection and gesture classification.

cs.CV