arXiv ScienceSearch

arXiv subjects

Ajay Gupta

Publications and source records attributed to Ajay Gupta.

At least 19 recordsLinked to original sources

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultimately leading to more precise predictions and interactions in real-world applications. This paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. We systematically review approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. We also outline current limitations related to dataset bias, knowledge incompleteness, and integration challenges. Finally, we highlight prospective research trajectories in cross-modal reasoning, scalable commonsense knowledge injection, and neuro-symbolic hybrid architectures to develop truly intelligent visual systems.

cs.AI

Rubrics as Privileged Information for Open-Ended Generation

On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structurally constrains valid continuations. We extend OPSD to open-ended generation using soft PI in the form of rubrics that guide preferences but admit many valid responses. Rubrics have served as scalar rewards for reinforcement learning (RL); we show that they provide substantially richer signal as dense PI for distillation, and contrary to intuition, soft rubric PI provides a larger and more effective training signal on student roll-outs than hard reference completion PI in this regime. A reference completion is one point in a set of valid responses, so distilling towards it over-constrains the student, while rubrics specify the preference structure shared across the set of valid responses. We show the effectiveness of using rubrics as PI for open-ended generation across Qwen and Llama model families and show that it outperforms rubric-as-reward (RaR) RL using HealthBench, a benchmark that grades open-ended health responses against physician-created rubrics, providing dense token-level supervision for open-ended tasks; RuPI beats RaR RL by up to +0.10 absolute score and, under matched recipe and KL direction, beats reference-PI by +0.034 to +0.079 absolute score across three models. We further show that these findings generalize to training on the RubricHub Science corpus and evaluating on ResearchQA: soft rubric PI outperforms both reference-PI distillation and RaR RL (66.6% vs. 64.2% and 57.6%).

cs.LG

Influence of oxygen ion implantation on magnetic microstructure in Pt/Co/Pt multilayers with perpendicular magnetic anisotropy

The interaction of oxygen with cobalt and cobalt-based alloys has been a very important topic in the field of spintronics as it leads to enhanced orbital anisotropy and interfacial Dzyaloshinskii-Moriya interaction (DMI), which are crucial in the context of applications such as magnetic tunnel junctions (MTJs) based data storage and domain wall (DW) motion. To understand the complex and interesting relationship between oxygen and ferromagnetic (FM)/heavy metal (HM) interfaces, we studied controlled oxygen ion implantation in a cobalt layer located in a Pt/Co 1.2 /Pt (nm) multilayer with a specific structure. At high implantation fluence, the perpendicular anisotropy was lost, as verified by in-plane hysteresis measurements. Under low magnetic field conditions, the DW dynamics of Co/Pt multilayers were analyzed, highlighting key parameters such as DW velocity, roughness amplitude, and roughness exponent. After O+-ion implantation, the DW velocity increased by more than 50 times, rising from 5 um/s to 300 um/s compared with the as-deposited multilayer. The fundamental cause of this improvement is the structural and magnetic changes brought by the implantation, which successfully lower the energy barriers preventing DW movements. The results show how oxygen implantation can be used to precisely tailor the ferromagnetic interfaces, leading to promised improvements in the functionality of next-generation spintronic devices.

cond-mat.mtrl-sci

Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse

Detecting phishing, spam, fake accounts, data scraping, and other malicious activity in online social networks (OSNs) is a problem that has been studied for well over a decade, with a number of important results. Nearly all existing works on abuse detection have as their goal producing the best possible binary classifier; i.e., one that labels unseen examples as "benign" or "malicious" with high precision and recall. However, no prior published work considers what comes next: what does the service actually do after it detects abuse? In this paper, we argue that detection as described in previous work is not the goal of those who are fighting OSN abuse. Rather, we believe the goal to be selecting actions (e.g., ban the user, block the request, show a CAPTCHA, or "collect more evidence") that optimize a tradeoff between harm caused by abuse and impact on benign users. With this framing, we see that enlarging the set of possible actions allows us to move the Pareto frontier in a way that is unattainable by simply tuning the threshold of a binary classifier. To demonstrate the potential of our approach, we present Predictive Response Optimization (PRO), a system based on reinforcement learning that utilizes available contextual information to predict future abuse and user-experience metrics conditioned on each possible action, and select actions that optimize a multi-dimensional tradeoff between abuse/harm and impact on user experience. We deployed versions of PRO targeted at stopping automated activity on Instagram and Facebook. In both cases our experiments showed that PRO outperforms a baseline classification system, reducing abuse volume by 59% and 4.5% (respectively) with no negative impact to users. We also present several case studies that demonstrate how PRO can quickly and automatically adapt to changes in business constraints, system behavior, and/or adversarial tactics.

cs.LG

Investigation of Fe-Ag and Ag-Fe Interfaces in Ag-57Fe-Ag trilayer Using Nuclear Resonance Scattering under X-ray Standing Wave Conditions

Understanding the interfaces of layered nanostructures is key to optimizing their structural and magnetic properties for the desired functionality. In the present work, the two interfaces of a few nm thick Fe layer in Ag-57Fe-Ag trilayers are studied with a depth resolution of a fraction of a nanometer using x-ray standing waves (XSWs) generated by an underlying [W-Si]x10 multilayer (MLT) at an x-ray incident angle around the Bragg peak of the MLT. Interface selectivity in Ag-57Fe-Ag trilayers was achieved by moving XSW antinodes across the interfaces by optimizing suitable incident angles and performing depth-resolved nuclear resonance scattering (NRS) and X-ray fluorescence (XRF) measurements for magnetic and structural properties. The combined analysis revealed that the rms roughness of 57Fe-on-Ag and Ag-on-57Fe interfaces are not equal. The roughness of the 57Fe-on-Ag interface is 10 Angstrom, while that of the Ag-on-57Fe interface is 6 Angstrom. 57Fe isotope sensitive NRS revealed that hyperfine field (HFF) at both interfaces of 57Fe-on-Ag and Ag-on-57Fe interfaces are distinct, which is consistent with the difference in interface roughnesses measured as root mean square (RMS) roughness. Thermal annealing induces 57Fe diffusion into the Ag layer, and annealing at 325 C transforms the sample into a paramagnetic state. This behavior is attributed to forming 57Fe nanoparticles within the Ag matrix, exhibiting a paramagnetic nature. These findings provide deep insights into interface properties crucial for developing advanced nanostructures and spintronic devices.

cond-mat.mtrl-sci

Leveraging Explainable AI to Analyze Researchers' Aspect-Based Sentiment about ChatGPT

The groundbreaking invention of ChatGPT has triggered enormous discussion among users across all fields and domains. Among celebration around its various advantages, questions have been raised with regards to its correctness and ethics of its use. Efforts are already underway towards capturing user sentiments around it. But it begs the question as to how the research community is analyzing ChatGPT with regards to various aspects of its usage. It is this sentiment of the researchers that we analyze in our work. Since Aspect-Based Sentiment Analysis has usually only been applied on a few datasets, it gives limited success and that too only on short text data. We propose a methodology that uses Explainable AI to facilitate such analysis on research data. Our technique presents valuable insights into extending the state of the art of Aspect-Based Sentiment Analysis on newer datasets, where such analysis is not hampered by the length of the text data.

cs.CL

CovidMis20: COVID-19 Misinformation Detection System on Twitter Tweets using Deep Learning Models

Online news and information sources are convenient and accessible ways to learn about current issues. For instance, more than 300 million people engage with posts on Twitter globally, which provides the possibility to disseminate misleading information. There are numerous cases where violent crimes have been committed due to fake news. This research presents the CovidMis20 dataset (COVID-19 Misinformation 2020 dataset), which consists of 1,375,592 tweets collected from February to July 2020. CovidMis20 can be automatically updated to fetch the latest news and is publicly available at: https://github.com/everythingguy/CovidMis20. This research was conducted using Bi-LSTM deep learning and an ensemble CNN+Bi-GRU for fake news detection. The results showed that, with testing accuracy of 92.23% and 90.56%, respectively, the ensemble CNN+Bi-GRU model consistently provided higher accuracy than the Bi-LSTM model.

cs.LG

Combining Compressions for Multiplicative Size Scaling on Natural Language Tasks

Quantization, knowledge distillation, and magnitude pruning are among the most popular methods for neural network compression in NLP. Independently, these methods reduce model size and can accelerate inference, but their relative benefit and combinatorial interactions have not been rigorously studied. For each of the eight possible subsets of these techniques, we compare accuracy vs. model size tradeoffs across six BERT architecture sizes and eight GLUE tasks. We find that quantization and distillation consistently provide greater benefit than pruning. Surprisingly, except for the pair of pruning and quantization, using multiple methods together rarely yields diminishing returns. Instead, we observe complementary and super-multiplicative reductions to model size. Our work quantitatively demonstrates that combining compression methods can synergistically reduce model size, and that practitioners should prioritize (1) quantization, (2) knowledge distillation, and (3) pruning to maximize accuracy vs. model size tradeoffs.

cs.CL

Offensive Language and Hate Speech Detection with Deep Learning and Transfer Learning

Toxic online speech has become a crucial problem nowadays due to an exponential increase in the use of internet by people from different cultures and educational backgrounds. Differentiating if a text message belongs to hate speech and offensive language is a key challenge in automatic detection of toxic text content. In this paper, we propose an approach to automatically classify tweets into three classes: Hate, offensive and Neither. Using public tweet data set, we first perform experiments to build BI-LSTM models from empty embedding and then we also try the same neural network architecture with pre-trained Glove embedding. Next, we introduce a transfer learning approach for hate speech detection using an existing pre-trained language model BERT (Bidirectional Encoder Representations from Transformers), DistilBert (Distilled version of BERT) and GPT-2 (Generative Pre-Training). We perform hyper parameters tuning analysis of our best model (BI-LSTM) considering different neural network architectures, learn-ratings and normalization methods etc. After tuning the model and with the best combination of parameters, we achieve over 92 percent accuracy upon evaluating it on test data. We also create a class module which contains main functionality including text classification, sentiment checking and text data augmentation. This model could serve as an intermediate module between user and Twitter.

cs.CL

Picking Pearl From Seabed: Extracting Artefacts from Noisy Issue Triaging Collaborative Conversations for Hybrid Cloud Services

Site Reliability Engineers (SREs) play a key role in issue identification and resolution. After an issue is reported, SREs come together in a virtual room (collaboration platform) to triage the issue. While doing so, they leave behind a wealth of information which can be used later for triaging similar issues. However, usability of the conversations offer challenges due to them being i) noisy and ii) unlabelled. This paper presents a novel approach for issue artefact extraction from the noisy conversations with minimal labelled data. We propose a combination of unsupervised and supervised model with minimum human intervention that leverages domain knowledge to predict artefacts for a small amount of conversation data and use that for fine-tuning an already pretrained language model for artefact prediction on a large amount of conversation data. Experimental results on our dataset show that the proposed ensemble of unsupervised and supervised model is better than using either one of them individually.

cs.AI

Ensembling Low Precision Models for Binary Biomedical Image Segmentation

Segmentation of anatomical regions of interest such as vessels or small lesions in medical images is still a difficult problem that is often tackled with manual input by an expert. One of the major challenges for this task is that the appearance of foreground (positive) regions can be similar to background (negative) regions. As a result, many automatic segmentation algorithms tend to exhibit asymmetric errors, typically producing more false positives than false negatives. In this paper, we aim to leverage this asymmetry and train a diverse ensemble of models with very high recall, while sacrificing their precision. Our core idea is straightforward: A diverse ensemble of low precision and high recall models are likely to make different false positive errors (classifying background as foreground in different parts of the image), but the true positives will tend to be consistent. Thus, in aggregate the false positive errors will cancel out, yielding high performance for the ensemble. Our strategy is general and can be applied with any segmentation model. In three different applications (carotid artery segmentation in a neck CT angiography, myocardium segmentation in a cardiovascular MRI and multiple sclerosis lesion segmentation in a brain MRI), we show how the proposed approach can significantly boost the performance of a baseline segmentation method.

eess.IV

Carbon to Diamond: An Incident Remediation Assistant System From Site Reliability Engineers' Conversations in Hybrid Cloud Operations

Conversational channels are changing the landscape of hybrid cloud service management. These channels are becoming important avenues for Site Reliability Engineers (SREs) %Subject Matter Experts (SME) to collaboratively work together to resolve an incident or issue. Identifying segmented conversations and extracting key insights or artefacts from them can help engineers to improve the efficiency of the incident remediation process by using information retrieval mechanisms for similar incidents. However, it has been empirically observed that due to the semi-formal behavior of such conversations (human language) they are very unique in nature and also contain lot of domain-specific terms. This makes it difficult to use the standard natural language processing frameworks directly, which are popularly used in standard NLP tasks. %It is important to identify the correct keywords and artefacts like symptoms, issue etc., present in the conversation chats. In this paper, we build a framework that taps into the conversational channels and uses various learning methods to (a) understand and extract key artefacts from conversations like diagnostic steps and resolution actions taken, and (b) present an approach to identify past conversations about similar issues. Experimental results on our dataset show the efficacy of our proposed method.

cs.CL

Volumetric landmark detection with a multi-scale shift equivariant neural network

Deep neural networks yield promising results in a wide range of computer vision applications, including landmark detection. A major challenge for accurate anatomical landmark detection in volumetric images such as clinical CT scans is that large-scale data often constrain the capacity of the employed neural network architecture due to GPU memory limitations, which in turn can limit the precision of the output. We propose a multi-scale, end-to-end deep learning method that achieves fast and memory-efficient landmark detection in 3D images. Our architecture consists of blocks of shift-equivariant networks, each of which performs landmark detection at a different spatial scale. These blocks are connected from coarse to fine-scale, with differentiable resampling layers, so that all levels can be trained together. We also present a noise injection strategy that increases the robustness of the model and allows us to quantify uncertainty at test time. We evaluate our method for carotid artery bifurcations detection on 263 CT volumes and achieve a better than state-of-the-art accuracy with mean Euclidean distance error of 2.81mm.

cs.CV

Influence of interface and microstructure on magnetization of epitaxial Fe4N thin film

Epitaxial Fe4N thin films grown on lattice-matched LaAlO3 (LAO) substrate using sputtering and molecular beam epitaxy techniques have been studied in this work. Within the sputtering process, films were grown with conventional direct current magnetron sputtering (dcMS) and for the first time, using a high power impulse magnetron sputtering (HiPIMS) process. Surface morphology and depth profile reveal that HiPIMS deposited film has the lowest roughness, the highest packing density and the sharpest interface. La from the LAO substrate and Fe from the film interdiffuse and forms an undesired interface spreading to an extent of about 10-20 nm. In the HiPIMS process, layer by layer type growth leads to a globular microstructure which restricts the extent of the interdiffused interface. Such substrate-film interactions and microstructure play a vital role in affecting the electronic hybridization and magnetic properties of Fe4N films. The magnetic moment (Ms) was compared using bulk, element-specific and magnetic depth profiling techniques. We found that Ms was the highest when the thickness of the interdiffused layer was lowest and only be achieved in the HiPIMS grown samples. Presence of small moment at the N site was also evidenced by element-specific x-ray circular dichroism measurement in HiPIMS grown sample. A large variation in the Ms values of Fe4N found in the experimental works carried out so far could be due to such interdiffused layer which is generally not expected to form in otherwise stable oxide substrate. In addition, a consequence of substrate-film interdiffusion and microstructure results in different kinds of different kind of magnetic anisotropies in films grown using different techniques.

cond-mat.mtrl-sci

Leveraging Machine Learning and Big Data for Smart Buildings: A Comprehensive Survey

Future buildings will offer new convenience, comfort, and efficiency possibilities to their residents. Changes will occur to the way people live as technology involves into people's lives and information processing is fully integrated into their daily living activities and objects. The future expectation of smart buildings includes making the residents' experience as easy and comfortable as possible. The massive streaming data generated and captured by smart building appliances and devices contains valuable information that needs to be mined to facilitate timely actions and better decision making. Machine learning and big data analytics will undoubtedly play a critical role to enable the delivery of such smart services. In this paper, we survey the area of smart building with a special focus on the role of techniques from machine learning and big data analytics. This survey also reviews the current trends and challenges faced in the development of smart building services.

cs.CY

Magnetism and structure of in-situ grown FeN films studied using N K-edge XAS and nuclear resonance scattering

We studied the structural and magnetic properties of \textit{in-situ} grown iron mononitride (FeN) thin films. Initial stages of film growth were trapped utilizing synchrotron based soft x-ray absorption near edge spectroscopy (XANES) at the N $K$-edge and nuclear resonant scattering (NRS). Films were grown using dc-magnetron sputtering, separately at the experimental stations of SXAS beamline (BL01, Indus 2) and NRS beamline (P01, Petra III). It was found that the initial stages of film growth differs from the bulk of it. Ultrathin FeN films, exhibited larger energy separation between the t$_{2g}$ and e$_g$ features and an intense e$_g$ feature in the N $K$-edge pattern. This indicates that a structural transition is taking place from the rock-slat (RS)-type FeN to zinc-blende(ZB)-type FeN when the thickness of films increases beyond 5\,nm. The behavior of such N $K$-edge features correlates very well with the emergence of a magnetic component appearing in the NRS pattern at 100\,K in ultrathin FeN films. Combining the \textit{in-situ} XANES and NRS measurements, it appears that initial FeN layers grow in RS-type structure having a magnetic ground state. Subsequently, the structure changes to ZB-type which is known to be non-magnetic. Observed results help in resolving the long standing debate about the structure and the magnetic ground state of FeN.

cond-mat.mtrl-sci

Evolution of magnetic anisotropy in cobalt film on nanopatterned silicon substrate studied in situ using MOKE

Evolution of magnetization behaviour of cobalt film on nano patterned silicon substrate, with film thickness, has been studied. In situ magneto-optical Kerr effect measurements during film deposition allowed us to study genuine thickness dependence of magnetization behaviour, all other parameters like surface topology, deposition conditions remaining invariant. The film exhibits uniaxial magnetic anisotropy, with its magnitude decreasing with increasing film thickness. Analysis shows that anisotropy has contributions from both, i) exchange energy which is volume dependent and, ii) stray dipolar fields at the surface/interface. This suggests that local magnetization follows only partially the topology of the rippled surface. As expected from energy considerations, for small film thickness, the local magnetization closely follows the surface contour of the ripples making the volume term as the dominant contribution. With increasing film thickness, the local magnetization gradually deviates from the local slope and approaches towards a uniform magnetization along the macroscopic film plane making the surface term as the dominant contribution. Significant deviation from the anisotropy energy expected on the basis of theoretical considerations can be attributed to several factors like, deviation of surface topology from an ideal sinusoidal wave, breaks of continuity along the ripple direction, defects like pattern dislocations, and possible decrease in surface modulation depth with increasing film thickness.

cond-mat.mtrl-sci

Interface Sharpening in Miscible and Isotopic Multilayers: Role of Short-Circuit Diffusion

Atomic diffusion at nanometer length scale may differ significantly from bulk diffusion, and may sometimes even exhibit counterintuitive behavior. In the present work, taking Cu/Ni as a model system, a general phenomenon is reported which results in sharpening of interfaces upon thermal annealing, even in miscible systems. Anomalous x-ray reflectivity from a Cu/Ni multilayer has been used to study the evolution of interfaces with thermal annealing. Annealing at 423 K results in sharpening of interfaces by about 38%. This is the temperature at which no asymmetry exists in the inter-diffusivities of Ni and Cu. Thus, the effect is very general in nature, and is different from the one reported in the literature, which requires a large asymmetry in the diffusivities of the two constituents [Z. Erd\'elyi et al., Science 306, 1913 (2004).]. General nature of the effect is conclusively demonstrated using isotopic multilayers of 57Fe/naturalFe, in which evolution of isotopic interfaces has been observed using nuclear resonance reflectivity. It is found that annealing at suitably low temperature (e.g. 523 K) results in sharpening of the isotopic interfaces. Since chemically it is a single Fe layer, any effect associated with concentration dependent diffusivity can be ruled out. The results can be understood in terms of fast diffusion along short-circuit paths like triple junctions, which results in an effective sharpening of the interfaces at relatively low temperatures.

cond-mat.mtrl-sci